feat(agent_runtime): advisory repeated-evidence ledger alongside repeat_guard - #247
Conversation
repeat_guard only sees *consecutive* identical calls. The other classic stall interleaves different calls (A, B, A, B, ...) and gets the same evidence back every time: no call is ever repeated consecutively, so nothing fires while the run learns nothing. EvidenceLedger keys on the call *and* its result. For each canonical call signature it counts the result fingerprints already seen; when one pair comes back `evidence_ledger_threshold` times (default 3) the call provably is not making progress and one reminder is emitted. Same discipline as repeat_guard, which is the point: - advisory only: nothing is delayed, rewritten, or blocked, and the decision stays with the model; the ledger sits in front of any hard stop (`max_iterations`, `should_stop_callback`); - it rides the existing reminder channel (`AgentRunSpec.evidence_ledger_threshold`, mirroring `repeat_call_thresholds`, and the same single user message and `repeat_guard` note source), so a call the tracker already flagged is left to the tracker — one iteration injects at most one reminder per call; - the reminder never quotes tool output, and its state is bounded (256 calls, LRU); - results are a dynamic boundary, so a non-text result (the Goal tools return dicts) is canonicalized instead of assuming `str`. `None` disables the layer. (cherry picked from commit f1b9b2a2ba16299db64c0ad42ae5fa7b0208b060)
|
Merged into For the record, since it now affects every run: an interleaved poll such as |
Description
Adds an advisory count of repeated evidence next to
repeat_guard, in the shape suggested when #178 was closed: an advisory layer feeding the guard's existing injection channel, with a real entry point and no hard blocks.repeat_guardonly counts consecutive identical calls. A stall that alternates calls (A, B, A, B, …) never trips it even though the model is learning nothing. This ledger keys on (canonical call signature, result fingerprint) and counts how often the same evidence comes back, regardless of interleaving.Related Issues
Refs #178 (closed) — reworked per review.
Changes Made
core/agent_runtime/evidence_ledger.py(+102):EvidenceLedger.observe(tool_name, arguments, result) -> str | None. Default thresholdDEFAULT_NO_PROGRESS_THRESHOLD = 3, matchingrepeat_guard's default-on posture; bounded at 256 tracked calls with LRU eviction; fires exactly once per distinct result; invalid thresholds fail loudly (boolrejected, ints ≥ 2). Non-text results are canonicalized — the Goal tools return dicts — so a reminder can always be computed and a run never fails because a reminder could not be.core/agent_runtime/runner.py(+54/−7): the ledger rides the existing reminder channel rather than adding one.AgentRunSpec.evidence_ledger_thresholdmirrorsrepeat_call_thresholds; the reminder is a single injected user message plus the samerepeat_guardnote source used today. Calls already flagged by the tracker are left to the tracker, so the two never double-message.Nonedisables it.tests/test_evidence_ledger.py(+369, 12 tests): threshold semantics and once-only firing; the interleaved stall the guard cannot see; argument-order insensitivity; new evidence restarting the count; bounded tracking with eviction; invalid thresholds; structured results; and four runner-level integration tests (injection exactly where the guard is blind; the guard still owns consecutive repeats; a structured result counts as evidence; disabled means silent).Checklist
ruff check+ruff format --checkclean)Additional Notes
max_iterations,should_stop_callback) rather than replacing one.tests/test_hooks.pyandtests/test_session_end_lifecycle.pywere triaged as unrelated (identical results withrunner.pystashed) before this was submitted.