Skip to content

issue25 completion failure recovery timing in task report - #26

Open
lukemaj wants to merge 8 commits into
mainfrom
feat/25-completion-recovery-timing
Open

lukemaj wants to merge 8 commits into
mainfrom
feat/25-completion-recovery-timing

Conversation

@lukemaj

@lukemaj lukemaj commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Implements #25 on the approved timing/recovery base, plus the Router87 contract vocabulary for planner-directed routing.

Exact candidate: 7f3537f. Base: 6d096af (main). Branch: feat/25-completion-recovery-timing. PR remains unmerged; no merge, release, install, deploy, or publish.

What this candidate changes (narrow, additive only):

  • Per-attempt timing rows (agent_observer/timing.py) now carry the already-imported attempt evidence fields rc, proof_class, and meta_seq, which the human formatter (agent_observer/cli.py) reads from those rows. Previously every proof attempt line read proof_class=unknown and never showed rc/seq, while the top-level JSON attempts list had the values, so human and JSON reports disagreed. Copy uses .get (old records stay compatible, nothing inferred); JSON schema is additive only.
  • Strengthened the existing Router-shaped CLI regression (tests/test_router_contract.py): it now asserts the structured attempt_timing row for router:p5 (proof_class failed, rc 1, meta_seq 5) and asserts those values on the actual human formatter attempt line (attempt router:p5 ... proof_class=failed rc=1 seq=5), so a separate verification_attempt event line can no longer satisfy the check.
  • Preserved approved behavior unchanged: submission-to-accepted-completion vs elapsed-so-far vs partial span; per-attempt/session/role wall time with waiting only from explicit compatible timestamps; Router/native reconciliation without double counting; failures by class (verification counts, timeout/infrastructure rules, provider context pressure, launch reason never classifies); stage-specific recovery with same-stage completion; planner_directed vocabulary and four planner harnesses; privacy/event/state/seq compatibility. Classification and counts are not affected by this fix.

Proof on this exact candidate:

  • Failing-before: the strengthened test run against the pre-fix code fails with KeyError: 'proof_class' on the timing row. Post-fix it passes.
  • python3 -m unittest discover -s tests: 421 tests OK.
  • python3 -m compileall agent_observer tests: OK (no errors).

Private baselines (under /tmp/agent-recovery-spec/observer-25-review, disposable, no user DB mutation):

  • Controlled Router injection regenerated with this candidate: live-final-7f3537f.json + note live-final-7f3537f.md, from router-87-live-proof/live-final/state/jobs.db (read-only; sha256 be21deff...b83aa5 verified identical before and after) into a throwaway Observer DB (single task model-router-87). The per-attempt row for the failed proof now reads proof_class failed, rc 1, meta_seq 1, and the human line matches; raw 11 reconciled 11, verification counted once, recovery paired same-request/same-stage, planner harness codex on both jobs. Controlled injection, not organic evidence.
  • Bounded T3 t3code-1-4af5983.json + t3-baseline-4af5983.md stand unchanged: this fix touches no classification or counts, so the scoped 3/7 figures, d1/f1 recoveries, and the two-missing-v2-jobs gap are unaffected.

Router87 status and gaps (coordinated via issue87, Router workspace not edited):

  • Router #87 remains OPEN; combined acceptance is not claimed.
  • Remaining gaps: live planner-directed trace, live exhaustion/planner-wake path, live PR path, and organic verification/cancel-intent events remain unobserved; rc143/125 alignment stays Router-owned. No cancellation intent or verification events beyond the source are inferred; job-scoped intent limit documented.
  • No raw prompts, transcripts, or secrets leave the machine.

Version: single 0.4.0 transition (VERSION, mirrors, CHANGELOG unchanged by this narrow fix); no second bump or release entry. Independent exact-candidate final review is coordinator-owned and is not claimed here.

Correct review findings: quota exhaustion counts as provider failure
inside failed/production with separate visibility; failure class uses
terminal and stage only, never Router reason; recovery pairs only
compatible request/session identities with progress at completed end;
Router/native attempt rows deduplicate on explicit evidence; union span
is merged covered duration and waiting needs same session/request;
first accepted-completion timestamp never moves; snapshot covers all
report inputs; usage coverage uses attributable responses.
@lukemaj
lukemaj force-pushed the feat/25-completion-recovery-timing branch from 2f203e3 to a00eda4 Compare September 24, 2026 21:10
… implementation work

Preserve first subsequent progress as any compatible completion with
explicit first_progress_stage, but require failed and candidate stages
both known and equal for recovered, active and repeated counts.
Unknown stage never matches. Add same-stage progress turn and timing
so recovery is measured at recovered work, not a dispatcher decision.
Update fixtures to explicit stages, expose stage labels in human and
JSON reports, and cover inputs in snapshot identity.
…pletion

Same-stage failures after a successful recovery start a new chain;
f1 recovered by r1 no longer counts later f5 as repeated failed
recovery. f5 keeps its own recovery row. Correct T3 fixture to the
real same-job request t3-fleet-pane-1-complete so d1->c1 dispatch
recovery and f1->r1 implementation recovery lock real gaps.
@lukemaj
lukemaj marked this pull request as ready for review September 24, 2026 23:49
@lukemaj

lukemaj commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

issue25 contract correction candidate 17f2165 (base 6d096af, branch feat/25-completion-recovery-timing, version 0.4.0 unchanged).

Implements report-contract.md findings 1 and 2 only: removed blanket rc124 to infrastructure in timing.py so explicit terminal timeout stays timeout including historical rc124 rows; added explicit terminal infrastructure in adapters/router.py TERMINAL_CLASSES so it imports as failed/infrastructure and classifies as infrastructure. Documented job-scoped cancel intent limit in docs/contracts.md. Updated GLOSSARY, README, contracts wording.

Regression: tests/test_router_contract.py fixture freshly imported via router.sync now holds historical worker timeout s7 with rc124 and 1801.0 s asserting stored timeout, attempt failed, class timeout and wall 1801, plus distinct explicit infrastructure row n11 with rc124 asserting stored infrastructure, attempt failed, class infrastructure, with denominators timeout 2 and infrastructure 1 and raw equals reconciled.

Proof on exact candidate: python3 -m unittest tests.test_router_contract.RouterContractImportTest.test_classification_contract pass; tests.test_router_contract plus tests.test_timing_recovery plus tests.test_router 73 pass; full suite python3 -m unittest discover -s tests 419 pass; python3 -m compileall -q agent_observer tests ok; public CLI proof via test_public_cli_carries_same_measurements pass.

Private bounded baseline refreshed from read-only copy with candidate CLI: /tmp/agent-recovery-spec/observer-25-review/t3-baseline-17f2165.md, export t3code-1-17f2165.json, copy live-ro-17f2165.db, stderr t3-17f2165-stderr.log. Cutoff 1790308143.683525, snapshot e869d1b3373c173f47051e89979b498ecd41fb47162518ff0eff4b4c64a7aa10, exit 3, live observer.db mtime 1790308148 unchanged. Scope stays 7 Router attempts from 2 completion jobs with 3/7 scoped failures (implementation 1, stall 1, timeout 1); 2 earlier v2 jobs unattributed; real 1801 s timeout stays timeout; job outcomes succeeded/blocked with cancel intent unknown; router_recovery empty.

Existing-data readiness: ready on old ledgers; legacy rows without rc keep terminal-only behavior. Pending Router87 gaps: Router87 still open, no real proof, recovery or cancel-intent events in live ledger, live copy lacks rc and cancel_requested columns so rc124 path is proven by fixture not by stale import, exit 143 without explicit job intent stays unknown and no jobwide cancel is inferred. No merge, install, release or publication. Independent final review remains coordinator-owned.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant