feat(loop): prompt-injection regression suite (pure mechanism) - #216
raymondginger2018-sudo wants to merge 1 commit into
Conversation
The #1 threat for agent systems is prompt injection: the model cannot reliably distinguish a malicious instruction from benign data, so the harness must separate data from instructions and keep untrusted content out of the privileged system-prompt region. This module provides: * ATTACK_SAMPLES — a structured regression corpus across four injection surfaces (spawn prompt, tool output, memory note, MCP remote content), each tagged with the guard it must satisfy. * render_data_block — the canonical data-boundary wrapper: untrusted content is injected inside delimiters with an explicit reference-only clause. * has_data_boundary — a pure check for tests to assert a surface got isolated. No LLM, no subprocess — the suite is a static contract that makes injection hardening a regression, not a one-off red-team exercise.
设计说明问题:agent 系统的三个关键注入面(system prompt 指令冲突、tool output 伪指令、memory note 越狱)缺乏统一的检测机制。现有做法是每个 profile 各自硬编码。 解法:定义 4 种标准攻击面常量 + 两个核心函数 关键设计决策:
与 PR #204 的关系:#204 在上游加入了 测试建议:每个 surface 写一个"注入成功"和一个"边界防护成功"的测试用例 |
|
Thank you for the submission, @raymondginger2018-sudo. Closing this one because it is already covered: the file is byte-identical to the |
…ta boundary `_escape_data_block` replaced four literal strings, so only the exact lower-case, space-free spellings were escaped. A note containing `</UNTRUSTED-DATA>` or `</untrusted-data >` therefore stayed in plain text, and the framed block carried a second closing tag: the remainder of the note read as if it sat outside the untrusted-data boundary (#216). Escape the tags case-insensitively, tolerating whitespace inside the delimiters, and add a regression test next to the existing boundary test.
Summary
The #1 threat for agent systems is prompt injection: the model cannot reliably distinguish a malicious instruction from benign data, so the harness must separate data from instructions and keep untrusted content out of the privileged system-prompt region.
This PR introduces a static, mechanism-only injection regression suite — no LLM calls, no subprocesses.
What it provides
4 injection surfaces, each with a guard:
spawn_prompttool_outputmemory_notemcp_contentKey functions
render_data_block(source, content)— the canonical data-boundary wrapper with reference-only clausehas_data_boundary(text, surface)— pure check for regression testsSURFACE_*constants — typed injection surface identifiersDesign principles
File
core/loop/injection_regression.py(new, 189 lines)Part of GenAI lessons 13/15 security module family.