Skip to content

promote injection-phrase tier to default-block (gated on field FP data) #12

Description

@ernestprovo23

The fuzzy prompt-injection phrase tier (MR1 / WRD-RES-INJECT-PHRASE) stays opt-in/monitor-only in v0.3 — there is no field false-positive data yet, and default-blocking a fuzzy tier risks alert fatigue.

Blocked on: collecting field FP telemetry on the curated phrase denylist.

Scope (once unblocked)

  • Gather FP rate data from real guard/inspect runs
  • If FP rate is acceptable, promote deterministic-enough subset to default-block
  • Otherwise keep opt-in and document the decision

Source: docs/THREAT_MODEL_V2.md MR1 / §4.2.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    blockedGated on a prerequisite (e.g. field data)enhancementNew feature or requestroadmapPlanned work from the threat model / roadmapsecuritySecurity-relevant work or review

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions