Skip to content

chore: port prompt injection rules - #10

Merged
sarahxsanders merged 2 commits into
mainfrom
port-prompt-injection-rules
May 1, 2026
Merged

sarahxsanders merged 2 commits into
mainfrom
port-prompt-injection-rules

Conversation

@sarahxsanders

Copy link
Copy Markdown
Collaborator

ports existing wizard prompt injection rules into 8 focused sub-rules in the Warlock, each covering a
specific attack class

Changes

  • 8 new prompt injection rules (category prompt_injection,
    action block):
    • instruction_override, role_hijack, jailbreak_persona,
      chat_markup (all critical)
    • system_prompt_leak (high)
    • posthog_integration_attack, posthog_feature_attack (medium)
    • base64_in_comment (critical)
  • Test layout restructure: moved from a single rules.test.ts to
    one .test.ts per rule under __tests__/rules/, plus a shared
    helpers.ts. Mirrors the one-rule-per-file pattern that
    src/scanner/rules/ already uses — makes finding a rule's tests
    trivial and keeps files from sprawling as we add more rules. The 3
    PR-1 tests were migrated to the new layout (same assertions).

Test plan

  • pnpm test passes (169 tests)
  • pnpm build passes, all 11 .yar files ship to dist/
  • Every new rule has positive + negative match tests per
    CONTRIBUTING.md
  • YARA compile is clean (rules load without syntax errors)

@sarahxsanders sarahxsanders changed the title port prompt injection rules chore: port prompt injection rules Apr 24, 2026
@sarahxsanders
sarahxsanders requested a review from a team April 24, 2026 21:28
@sarahxsanders sarahxsanders mentioned this pull request Apr 24, 2026
3 tasks done

@gewenyu99 gewenyu99 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will for sure be useful. There are two things I want to discuss:

  • How will these be run? I certainly think these are fragile if anyone can run random strings against these, even if it's a black box. I feel like the run and the results of these cannot be public. The ability to observe inputs and patterns of what passes and doesn't is a big enough attack surface for regex matches like these.

  • How strict should these regex's be? I think all of these are great, but very easy to skirt around. And I'm not sure if that's a problem or not. (we cannot create all powerful regexes xD)

Comment thread src/scanner/rules/prompt_injection_base64_in_comment.yar
Comment thread src/scanner/rules/prompt_injection_instruction_override.yar
Comment thread src/scanner/rules/prompt_injection_jailbreak_persona.yar
Comment thread src/scanner/rules/prompt_injection_posthog_feature_attack.yar
Comment thread src/scanner/__tests__/rules/prompt_injection_base64_in_comment.test.ts Outdated
Comment thread src/scanner/rules/prompt_injection_base64_in_comment.yar Outdated
@sarahxsanders

Copy link
Copy Markdown
Collaborator Author

@gewenyu99 > This will for sure be useful. There are two things I want to discuss:

  • How will these be run? I certainly think these are fragile if anyone can run random strings against these, even if it's a black box. I feel like the run and the results of these cannot be public. The ability to observe inputs and patterns of what passes and doesn't is a big enough attack surface for regex matches like these.
  • How strict should these regex's be? I think all of these are great, but very easy to skirt around. And I'm not sure if that's a problem or not. (we cannot create all powerful regexes xD)

all very good questions!!

How will these be run?

TL;DR of this thread: it runs in process inside the consumer as a dependency (wizard: in runs, context mill: in release builds)

rule secrecy isn't really the strategy, defense in depth is. v2 of the security hardening introduces agentsh + an agent layer to make that defense in depth model deeper.

How strict should these regex's be?

the regex rules are designed to be the prevention layer of the defense in depth design, not all powerful. these rules are part of the v1 foundation, and v2 will layer agentsh. with v2 in mind, my thought was:

  • regex catches the lazy 80% before execution
  • agentsh + the agent layer (v2) handle the multi-step, semantic, encoded stuff that regex can't

@sarahxsanders
sarahxsanders requested a review from gewenyu99 April 30, 2026 21:04
@sarahxsanders
sarahxsanders merged commit e080bc2 into main May 1, 2026
9 checks passed
@sarahxsanders
sarahxsanders deleted the port-prompt-injection-rules branch May 1, 2026 21:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants