From 7a17d31be62339443671b8db480d20136e23da28 Mon Sep 17 00:00:00 2001 From: kjgbot Date: Wed, 16 Sep 2026 07:06:29 -0700 Subject: [PATCH 1/2] fix(drive): trust an escalation only when this tick wrote it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ops/NEEDS_HUMAN.md is committed on main (082c62aa, 2026-09-06) and nothing has ever deleted it. assess-gate escalated on the file's mere existence, so every drive tick since 2026-09-12 exited 75 before doing any work: PRs #417, #420, #422, #424, #426, #427 and #428 are seven consecutive cloud runs whose entire diff is that file and ops/NEXT.md. None merged. The escalation was also correct about a real defect nobody had traced: ops/autodrive.sh launched every run with `launch-gate.sh 3` while passing it the gate-2 brief, so launch-gate.sh synthesised "TARGET — gate 3" wrapped around gate-2 work on every tick. The assessors were reporting a launcher bug, once per run, for four days. Operator decision: the next gate is Gate 2, not Gate 3. - assess-gate now requires two independent signals before trusting an escalation — the file exists AND this tick wrote it. Freshness reuses the `git log --oneline main..HEAD -- ` idiom already used for ops/NEXT.md a few lines below, widened by the uncommitted case because per-step propagation is lossy and losing a live escalation is the worse error. A stale file is ignored loudly; a live one still exits 75. - ops/drive-assess-gate.test.mjs pins both directions, extracting the gate script from workflows/drive.yaml so the test cannot drift from it. Verified by mutation: against unmodified main the stale case fails with exit 75, reproducing the wedge. - ops/NEEDS_HUMAN.md deleted; its durable content preserved in a dated ops/STATE.md block, including the still-open question that ops/TARGET.md is synthesised into a throwaway worktree and never reaches the diff. - ops/autodrive.sh launches gate 2, matching the brief it passes. - ops/AUTODRIVE_BRIEF.md retargeted off the hn-monitor runner PR #120 already shipped, onto the open half of RFC-0001 deviation D1. - ops/NEXT.md rewritten as that Gate 2 package. - ops/STATE.md gate-2 clause 1 corrected: it claimed trigger-plane liveness was unimplemented, but PR #122 shipped it two weeks ago. That entry would have sent the next run to rebuild working code. - workflows/drive-cloud.yaml regenerated with ops/gen-drive-cloud.py; only assess-gate-1 differs semantically, the rest is pre-existing PyYAML reflow. Co-Authored-By: Claude Opus 5 (1M context) --- ops/AUTODRIVE_BRIEF.md | 92 ++++----------- ops/NEEDS_HUMAN.md | 83 -------------- ops/NEXT.md | 187 +++++++++++++++---------------- ops/STATE.md | 76 +++++++++++-- ops/autodrive.sh | 8 +- ops/drive-assess-gate.test.mjs | 130 +++++++++++++++++++++ workflows/drive-cloud.yaml | 199 ++++++++++++++++++--------------- workflows/drive.yaml | 43 ++++++- 8 files changed, 464 insertions(+), 354 deletions(-) delete mode 100644 ops/NEEDS_HUMAN.md create mode 100644 ops/drive-assess-gate.test.mjs diff --git a/ops/AUTODRIVE_BRIEF.md b/ops/AUTODRIVE_BRIEF.md index 81494eb3..3560064d 100644 --- a/ops/AUTODRIVE_BRIEF.md +++ b/ops/AUTODRIVE_BRIEF.md @@ -1,82 +1,36 @@ -Build sub-PR A of the Gate 2 push: a real `hn-monitor` polling runner in the SDK. CODE task, `sdk/src/`-side. This is a scaffolding PR — proof that the workload EXECUTES end-to-end is deliberately deferred to sub-PR B (integration test). Do not conflate the two. +Close the rest of RFC-0001 deviation D1: classify a wake-context resolution failure by cause and journal it. GATE 2. CODE task, `kernel/`-side (Rust). -**Context:** RFC-0001 §3 gate 2 is done when "hn-monitor runs as a relayflow in production, triggered by its real events, with zero bespoke persistence." Every primitive already exists in this repo — event triggers (PR #14, `kernel/relayflowd/tests/event_wake.rs`), the flow spec (`testdata/hn-monitor.flow.yaml`), the poller (`sdk/src/hn-poller.ts`), the agent worker (`sdk/src/worker.ts` from PR #53), a one-shot demo (`sdk/src/demo-hn-monitor.ts`) — but nothing has ever run them together as a continuous workload. This PR fixes that. +**Retargeted 2026-09-16.** This brief previously asked for a new hn-monitor runner module in the SDK, as a gate-2 scaffolding PR. PR #120 shipped that work on 2026-09-01 as `packages/sdk/src/cli/hn-monitor.ts`. The brief outlived its work package by two weeks, and because `ops/autodrive.sh` also launched it under a hardcoded "gate 3", every assessor that read both correctly reported a contradiction and escalated instead of building. Seven consecutive runs (#417, #420, #422, #424, #426, #427, #428) produced nothing but escalation notes. **Do not rebuild the hn-monitor runner. It exists.** -**Prior attempt (PR #83, closed):** produced a functional runner but was rejected by the swarm on five real findings. Address them in this attempt: +**Context:** RFC-0001 states that gate 2 stays RED until deviations D1 and D2 are both closed. D2 is blocked upstream on engine-side epoch rollover — the RFC sequences it rollover → D2 → gate 2 and says D2 is not implementable in isolation, because every construction of `EpochSummaryPayload` in the kernel is inside a test. That leaves D1 as the actionable half, and it is already half done: PR #252 stopped `resolve_wake_context` in `kernel/relayflowd/src/engine/drive.rs` swallowing a journal scan error into `wake_context: None`. What remains is the classification that function's own doc comment says is missing — "transient → retry the attempt under its budget; permanent → park `needs_human`, and none of that classification exists yet". -1. **Fail-closed on journal errors.** #83's `catch (err) { onPollError(err) }` swallowed EVERY error including `eventSubmit` journal failures — violates covenant 2 (fail-closed) and RFC-0001 §1. Only fetch-level errors (network flakiness, HN API rate limits) may be swallowed; a journal write failure MUST throw and terminate the runner. Split: `try { fetch } catch { onFetchError }` around the network call, `try { eventSubmit } catch { rethrow }` around the journal call. +**The task.** Implement RFC rule 10a: -2. **AgentWorker.close() must release the worker (or explicitly document it does not).** #83 added `await worker.close()` to shutdown but the current `close()` only drains local promises — it does NOT tell the kernel to release the worker registration. Either: - - Add a `workerRelease` verb to `sdk/src/protocol.ts` and call it from `close()` (preferred — completes the shutdown contract), OR - - Add a one-line comment on `close()` naming exactly what shutdown intentionally does NOT do + - a **transient** failure (I/O error, lock contention, a truncated tail still being written) fails the attempt and stays **retryable** under the step's ordinary retry budget, journalling `wake_context_unresolved` with `reason: "transient"` and the underlying error + - a **permanent** failure (the run is open on a wake, the current segment reads cleanly, and no wake context is present) fails the step and parks as **`needs_human`**, not retryable, journalling `wake_context_unresolved` with `reason: "absent"` + - neither may fall back to `wake_context: None` — dispatching with no context is reserved for runs that were never woken -3. **Class field declaration order.** #83 declared `private readonly fetcher` AFTER the constructor. Works today because of ES2022 hoisting semantics but breaks silently if someone adds `= someDefault` to a declaration. Declare ALL fields at the top of the class body, before the constructor. +`wake_context_unresolved` does not exist anywhere in the kernel today, so it needs a typed entry beside `SubscriptionMatched` in `kernel/relayflowd-core/src/entry.rs`. -4. **Signal handlers must be opt-in via AbortSignal.** #83 registered `SIGTERM`/`SIGINT` handlers on the process directly with no opt-out. A library user embedding this can't cancel one runner without affecting others. Accept `signal?: AbortSignal` in options; the CLI wrapper (sub-PR C) can create + wire a process-signal-driven AbortController. +**The gate is a journal shape, not a log line.** RFC rule 11c: *never woken* is the ABSENCE of any `wake_context_unresolved` entry alongside a dispatch carrying no wake context; *failed to resolve* is the PRESENCE of that entry naming its `reason`, with no dispatch for that attempt. Assert on those two shapes. `kernel/relayflowd/tests/event_wake.rs` already builds a woken run and reads entries back out of the journal — model the new test on it. -5. **Test coverage for pollError branch.** #83's tests never asserted the loop survives a fetcher throw AND the loop TERMINATES on a journal throw. Add both cases; without them, someone regresses `onPollError` to a no-op and every test still passes. +**Definition of done, all of it** -## Do not re-do these - -Merged and closed; a PR redoing any will be closed: - - picker actionability (#42), unterminated backticks (#45) - - deterministic-command preflight refusal (#47) — do not touch preflight - - gate-1 race regression test (#48) — do not touch `kernel/relayflowd/src/server/tests.rs` or `server.rs` - - ops/NEXT.md validation (#50) — do not touch `sdk/src/work-package-validator.ts` - - SDK agent worker (#53) — `sdk/src/worker.ts` is done; you MAY modify `close()` per finding #2 above, but do NOT rewrite the attach/dispatch/complete flow - - SDK pretest hook (#69) — `sdk/package.json` builds the kernel before `npm test`; don't touch - -## The task - -Add `sdk/src/hn-monitor-runner.ts`. It composes the existing pieces into a continuous runner: - - - constructs a `JournalClient` connected to the running `relayflowd` socket - - constructs an `AgentWorker` (from `sdk/src/worker.ts`) and calls `workerAttach()` for `agent` steps — attach BEFORE first poll (a run parked because no worker attached is only revived by `run.resume`; the live-kernel suite pins this) - - loops: `pollHackerNewsOnce(spec, sink)` → sleep `POLL_INTERVAL_MS` (env-configurable, default 60000 = 60s) → repeat - - exit cleanly on `AbortSignal.abort` (drain in-flight steps, close client, release worker per finding #2) - - exported from `sdk/src/index.ts` - -Keep it small and honest: - - the worker must attach BEFORE the first poll - - the poller layer handles single-fetch failures with a typed error; the loop just moves to the next tick — but journal errors MUST fail the runner (finding #1) - - no scheduling logic beyond the sleep (the kernel owns retry and dedupe policy) - - no LLM calls; the runner is glue, not a reviewer - -## Explicit non-goals for THIS PR (belongs to later sub-PRs) - - - Proving the workload actually executes end-to-end (dispatch → step complete). That is sub-PR B (integration test with real relayflowd + fake HN fetch + assert step reaches `done`). This PR ONLY proves the runner assembles and its unit tests hold. - - CLI wrapper (`flows hn-monitor start`). That is sub-PR C. - - Ops/STATE.md gate-2 GREEN declaration. That is sub-PR D. - -Say all three explicitly in the PR body so the history lens doesn't reject on "runner doesn't prove workload runs." - -## Definition of done, all of it - - - `sdk/src/hn-monitor-runner.ts` exists, exports `HnMonitorRunner` from `sdk/src/index.ts` - - `sdk/src/worker.ts` — either `close()` calls `workerRelease` (add to protocol.ts if missing), OR a one-line comment names what close() intentionally does NOT do - - `sdk/src/protocol.ts` — if you added `workerRelease`, matching request/response definitions - - `sdk/tests/hn-monitor-runner.test.ts` covers ALL of these: - - fake fetch + mock journal client → runner submits an event on each tick - - abort signal triggers clean shutdown within one tick (worker released or documented) - - worker attach happens before first poll - - **fetch throw → loop survives** (onPollError called, next tick still runs) - - **journal throw → loop TERMINATES** (runner.run() rejects with the error) - - `cd sdk && npm test` green (pretest hook builds the kernel automatically) - - EVERY new test confirmed to FAIL against current code (comment out the source; the test fails), with the literal failing output pasted in your summary - - PR body explicitly names the non-goals (test-actually-runs is sub-PR B; CLI is sub-PR C; gate-2 declaration is sub-PR D) - - ops/NEXT.md correctly says Gate 2, not Gate 3 (the assessor on #83 confused itself) + - `wake_context_unresolved` is a typed journal entry carrying `reason`, written by the kernel at the classification site rather than only by a test + - transient fails the attempt and retries under the ordinary budget; permanent parks `needs_human` and does not retry + - a new kernel test asserts both journal shapes from rule 11c + - the `resolve_wake_context` doc comment no longer describes the classification as absent — it currently does, at length, and a stale comment that contradicts the code is worse than none + - EVERY new test confirmed to FAIL against current code (revert the source change; watch it fail), with the literal failing output pasted in your summary + - `cd kernel && sh ../ops/cargo.sh test --workspace` must be green, with the output pasted - as your LAST action, run `git status --porcelain` and paste it -## Out of scope for THIS tick — DO NOT TOUCH +**Out of scope for THIS tick — DO NOT TOUCH** + - D2, `EpochSummaryPayload`, and engine-side epoch rollover — blocked upstream, and the RFC says adding the field ahead of its producer is "the appearance of a fix, not one" + - trigger-plane liveness — shipped in PR #122 (`kernel/relayflowd/src/server/liveness.rs`) + - the hn-monitor runner or CLI — shipped in PR #120 + - flipping the gate-2 verdict in `ops/STATE.md` — Khaliq's read, never a run's - `.github/workflows/*` — no GHA changes - - `kernel/*` — the kernel side of gate 2 already works via PR #14 - - `workflows/*.yaml` — those are for later sub-PRs - - `ops/AUTODRIVE_BRIEF.md` — chief owns this file, not the drive loop - - CLI wrapper — sub-PR C, separate PR - - end-to-end integration test with real relayflowd — sub-PR B, separate PR - - ops/STATE.md gate-2 declaration — sub-PR D, separate PR - -## If you cannot finish + - `packages/sdk/*` — this is a kernel task -Say so and file what you learned. A minimal runner with an honest gap description beats a complete-looking one that doesn't shut down cleanly or leaks journal errors. +**If you cannot finish.** Say so and file what you learned. A partial classification with an honest gap description beats a complete-looking one that cannot tell transient from permanent — conflating those two is the exact defect D1 exists to name. diff --git a/ops/NEEDS_HUMAN.md b/ops/NEEDS_HUMAN.md deleted file mode 100644 index 8bfae4c9..00000000 --- a/ops/NEEDS_HUMAN.md +++ /dev/null @@ -1,83 +0,0 @@ -# NEEDS_HUMAN — Conflicting Work Package Context - -**Situation:** This run has conflicting scope context that requires human clarification. - -## The Conflict - -1. **ops/TARGET.md says:** Gate 3, build hn-monitor runner (sub-PR A), `sdk/src/` code task -2. **ops/NEXT.md says:** Gate 3, cloud review-swarm preflight validation, `.github/workflows/` task -3. **These are completely different tasks** — one is SDK code (track A per TARGET), one is GitHub Actions (track D per NEXT) - -## Evidence - -**ops/TARGET.md line 1-5:** -``` -# TARGET — gate 3 - -This run is pinned to **gate 3** and must not work on any other gate. - -**Scope:** Build sub-PR A of the Gate 2 push: a real `hn-monitor` polling runner in the SDK. CODE task, `sdk/src/`-side. -``` - -**ops/NEXT.md line 1-3:** -``` -# NEXT — gate 3: complete cloud review-swarm preflight validation and documentation - -**Scope:** Track D: Cloud review-swarm redesign — build `.github/workflows/review-swarm.yml` correctly this time -``` - -## The Charter Says - -Per charter/LEAD.md (the instruction I received): -- "Read ops/TARGET.md if it exists" — it does, says hn-monitor -- "Then read ops/STATE.md, ops/DIRECTIVES.md" — done -- "Then write ops/NEXT.md: the SINGLE highest-priority work package toward the current gate" - -But ops/NEXT.md ALREADY EXISTS with different work. - -## Additional Context Found - -**ops/STATE.md gate 2 block (lines 39-81)** says: -- PR #120 merged 2026-09-01 — `flows hn-monitor start` CLI runner -- Gate 2 is AMBER, not GREEN -- Two clauses remain: trigger-plane liveness, analyze-agent execution - -**Actual file check:** -- `packages/sdk/src/cli/hn-monitor.ts` exists (288 lines) -- Contains `runHnMonitor` function implementing all TARGET.md requirements -- Addresses all five findings from closed PR #83 - -## The Question - -**Which work package should this run execute?** - -**Option A: ops/TARGET.md wins (hn-monitor, sdk/ territory)** -- Overwrite ops/NEXT.md with hn-monitor work package -- But the code already exists (PR #120 merged per STATE.md) -- Task would be verification/testing, not building from scratch - -**Option B: ops/NEXT.md wins (review-swarm, .github/ territory)** -- Ignore ops/TARGET.md -- Execute the review-swarm validation work -- But TARGET.md explicitly pins this run to gate 3 hn-monitor work - -**Option C: TARGET.md is stale** -- The launcher wrote an outdated TARGET.md referencing closed PR #83 -- Real work is in ops/NEXT.md (review-swarm) -- Proceed with review-swarm, update TARGET understanding - -**Option D: Both are stale** -- Neither accurately reflects current gate 3 needs -- Assessor should read RFC-0001 §3 gate 3 definition -- Write fresh work package from RFC requirements - -## Recommendation - -**Option C** — ops/TARGET.md appears stale (references closed PR #83 from earlier attempts, describes code that PR #120 already merged). The active work package is ops/NEXT.md (review-swarm). But I need human confirmation before overwriting NEXT.md or executing potentially wrong work. - -## What I Need - -**Clear answer:** Which work package is correct for this run? -- If hn-monitor: shall I overwrite the review-swarm NEXT.md, or is there a different file I should write? -- If review-swarm: shall I proceed with ops/NEXT.md as-is and ignore TARGET.md? -- If neither: what is the actual gate 3 work I should assess? diff --git a/ops/NEXT.md b/ops/NEXT.md index cb9d1847..b5d0bc64 100644 --- a/ops/NEXT.md +++ b/ops/NEXT.md @@ -1,123 +1,112 @@ -# NEXT — gate 3 work package: document review-swarm secrets in README +# NEXT — gate 2: split wake-context resolution failure by cause (the rest of D1) -**Scope (from TARGET.md):** - -Track D: Cloud review-swarm redesign — build `.github/workflows/review-swarm.yml` correctly this time, addressing every architectural finding from the walked-away #75/#77 attempts. +**Scope.** The next gate is **gate 2**, by Khaliq's decision on 2026-09-16. +Not gate 3. If you are reading a target that says gate 3, it is stale — the +launcher used to synthesise "gate 3" over a gate-2 brief, which is the +contradiction that wedged the loop for four days. ## Objective -Complete the final missing piece of gate 3's Definition of Done: document `RELAY_WORKSPACE_KEY` and `CLOUD_API_KEY` secrets in README.md with instructions on how to obtain them. - -## Current state assessment - -All 9 architectural requirements from TARGET.md are SATISFIED in the existing code: +Close the remainder of **deviation D1** in +`docs/RFC-0001-everything-is-a-relayflow.md`. That document names it a +binding obligation, in these words: -1. ✅ Immutable gate — two checkout steps (`.github/workflows/review-swarm.yml:32-53`) -2. ✅ Unified verdict logic — `swarm-verdict.sh` sourced by both callers -3. ✅ Auth secret validation — preflight validates all three secrets (lines 141-188) -4. ✅ Sticky marker + transcripts — HTML anchors with upsert_comment -5. ✅ No author whitelist — verified absent -6. ✅ Cloud sandbox fetch on GHA runner — `swarm-prepare.sh` with GH_TOKEN -7. ✅ Timeout ordering — 60m < 65m < 75m with comments -8. ✅ Wait step records status — swarm_status output, always() post step -9. ✅ Transcript freshness — run-start marker with stale detection - -Verification commands all pass: ``` -bash -n .github/workflows/scripts/swarm-post.sh && \ -bash -n .github/workflows/scripts/swarm-prepare.sh && \ -bash -n .github/workflows/scripts/swarm-verdict.sh && \ -echo "All bash scripts parse OK" -# Output: All bash scripts parse OK - -python3 -c "import yaml; yaml.safe_load(open('.github/workflows/review-swarm.yml'))" && \ -python3 -c "import yaml; yaml.safe_load(open('workflows/review-swarm.yaml'))" && \ -echo "YAML files parse OK" -# Output: YAML files parse OK - -grep -i "whitelist\|github.event.pull_request.user.login" .github/workflows/review-swarm.yml || echo "No author whitelist found (GOOD)" -# Output: No author whitelist found (GOOD) - -grep -c "actions/checkout@v4" .github/workflows/review-swarm.yml -# Output: 2 +gate 2 cannot go green until D1 and D2 are closed ``` -**The gap:** TARGET.md Definition of Done item 6 requires: -> README.md — document `RELAY_WORKSPACE_KEY` secret + how to obtain - -Current reality: -``` -grep -c "RELAY_WORKSPACE_KEY\|CLOUD_API_KEY" README.md -# Output: 0 -``` - -README.md does NOT document these secrets. The workflow comment (`.github/workflows/review-swarm.yml:21-24`) references a runbook in the `AgentWorkforce/cloud` repo, but README has no such documentation. - -From `ops/NEEDS_HUMAN.md`, the secrets are stored and working (as of 2026-09-07), but gate 3 is blocked on Daytona CPU quota, not on implementation. The workflow WORKS; the documentation is missing. +D1 is half done. PR #252 stopped the resume path swallowing a journal scan +error into `wake_context: None`, so the error now propagates out of +`resolve_wake_context` in `kernel/relayflowd/src/engine/drive.rs`. That +function's own doc comment says what is still missing, in its own words: -## Files in scope +> That is deliberately narrower than the eventual contract. The intended +> end state classifies the failure and journals it (transient → retry the +> attempt under its budget; permanent → park `needs_human`), and none of +> that classification exists yet. -- `README.md` — add section documenting GitHub Actions secrets required for review-swarm +The classification still does not exist. `wake_context_unresolved` — the +journal entry RFC rule 10a requires, and the entry rule 11c makes the gate +assert on — appears nowhere in the kernel. Today an unreadable segment +aborts the dispatch before any attempt is recorded, and recovery only +arrives when the lease expires and a later drive abandons the attempt as +`Crashed`. Nothing in the journal says the wake context was the reason. -## Work package +## What RFC rule 10a requires -Add a "GitHub Actions Secrets" section to README.md documenting: +Resolution failure splits by cause, because transient I/O and permanent +corruption deserve opposite handling: -1. `RELAY_WORKSPACE_KEY` — Agent Relay workspace key for review swarm communication - - How to obtain: Contact repository administrator or see ops/NEEDS_HUMAN.md for historical context - - Why required: Enables agent coordination within review swarm workflow +- **Transient** (segment unreadable right now: I/O error, lock contention, a + truncated tail still being written) — fail the attempt, **retryable**, + under the step's ordinary retry budget. Journal `wake_context_unresolved` + with `reason: "transient"` and the underlying error. +- **Permanent** (the run is open on a wake, the current segment reads + cleanly, and no wake context is present) — fail the step and park as + **`needs_human`**, not retryable. Journal `wake_context_unresolved` with + `reason: "absent"`. -2. `CLOUD_API_KEY` — Agent Relay Cloud API credential for launching cloud workflows - - How to obtain: Minted per `AgentWorkforce/cloud → docs/runbooks/relay-ci-workflow-credential.md` - - Profile: `workflow-invoke` - - Scopes: `workflow:invoke:read` and `workflow:invoke:write` - - How to store: Repository Settings → Secrets and variables → Actions → New repository secret +Neither case may fall back to `wake_context: None`. Dispatching with no +context is reserved for runs that were never woken. -3. `CLOUD_API_URL` — Cloud API endpoint (typically `https://agentrelay.com/cloud`) - - Usually set as repository variable, not secret - - Defaults to production endpoint if not set +## Files in scope -The section should be brief (10-15 lines) and reference the workflow files for implementation details. +- `kernel/relayflowd-core/src/entry.rs` — add the `WakeContextUnresolved` + entry type and its `wake_context_unresolved` wire string, beside the + existing `SubscriptionMatched` / `subscription.matched` pair. +- `kernel/relayflowd/src/engine/drive.rs` — classify the failure at the + `resolve_wake_context` call site and journal it, instead of aborting the + dispatch before anything is recorded. Update the doc comment: it currently + describes the classification as absent, and that must stop being true. +- A new kernel integration test under kernel/relayflowd/tests/ asserting on + the two journal shapes rule 11c names. Model it on + `kernel/relayflowd/tests/event_wake.rs`, which already builds a woken run + and reads `SubscriptionMatched` back out of the journal. + +**The permanent case is reachable today.** A run that was woken, whose +`subscription.matched` entry carries no `wake_context` key, currently +resolves to `Ok(None)` and is indistinguishable from never-woken. That is +the shape to detect: open on a wake, clean read, no context. ## Definition of done -1. README.md contains a section documenting the three secrets/variables -2. Each entry states what it is and how to obtain it -3. Parse checks continue to pass: - ``` - bash -n .github/workflows/scripts/swarm-*.sh - python3 -c "import yaml; yaml.safe_load(open('.github/workflows/review-swarm.yml'))" - python3 -c "import yaml; yaml.safe_load(open('workflows/review-swarm.yaml'))" - ``` -4. Verification remains true: - ``` - grep -c "RELAY_WORKSPACE_KEY\|CLOUD_API_KEY" README.md - # Should return > 0 - grep -i "whitelist\|github.event.pull_request.user.login" .github/workflows/review-swarm.yml || echo "GOOD" - # Should return "GOOD" or nothing (no whitelist) +1. `wake_context_unresolved` exists as a typed journal entry with a `reason` + of `transient` or `absent`, and the kernel writes it at the classification + site — not only in a test. +2. A transient failure fails the attempt and stays retryable under the step's + ordinary budget. A permanent failure parks the run as `needs_human` and is + not retried. +3. A new test asserts on the two journal shapes from RFC rule 11c: *never + woken* is the ABSENCE of any `wake_context_unresolved` entry alongside a + dispatch with no wake context; *failed to resolve* is the PRESENCE of that + entry naming its reason, with no dispatch for that attempt. Assert on + journal entries, not on log text. +4. Every new test must be confirmed to FAIL against current code before the + fix — revert the source change, watch it fail, and paste the literal + failing output into your summary. A test that has never been observed + failing pins nothing. +5. `cargo test --workspace` must be green, run from the kernel directory: ``` -5. As final action: - ``` - git status --porcelain + cd kernel && sh ../ops/cargo.sh test --workspace ``` +6. As your last action, run `git status --porcelain` and paste it. ## Explicitly OUT of scope -- `.github/workflows/review-swarm.yml` (already correct, all 9 requirements satisfied) -- `workflows/review-swarm.yaml` (already correct) -- `.github/workflows/scripts/swarm-*.sh` (all already correct) -- `.gitignore` (no .review-target mask exists, already correct) -- `sdk/` (Track A owns that) -- `kernel/` (gate 1 done) -- `ops/*` (chief owns briefs and state) -- Any other GHA workflow -- Resolving the Daytona CPU quota block (that's in ops/NEEDS_HUMAN.md, different issue) -- Actually testing the workflow end-to-end (blocked on Daytona capacity per ops/NEEDS_HUMAN.md) - -## Why this is the work package - -TARGET.md's Definition of Done explicitly lists: -- Item 6: "PR body explicitly documents each of the 9 requirements above and shows where each is satisfied" -- Item 7: "`README.md` — document `RELAY_WORKSPACE_KEY` secret + how to obtain" - -The 9 requirements are satisfied in code. Item 7 is not satisfied. This is the remaining gap between current state and TARGET.md's done-when. +- **D2 and epoch rollover.** The RFC sequences this as engine-side epoch + rollover → D2 → gate 2 and states that D2 is not implementable in + isolation: every construction of `EpochSummaryPayload` in the kernel is + inside a test, so there is no site at which to add the carry-forward. + Adding the field first would be "the appearance of a fix, not one". Do not + start it here. +- **Trigger-plane liveness.** Already shipped in PR #122 — + `kernel/relayflowd/src/server/liveness.rs` with + `kernel/relayflowd/tests/subscription_liveness.rs`. Do not rebuild it. +- **The hn-monitor runner.** Already shipped in PR #120 — + `packages/sdk/src/cli/hn-monitor.ts`. Three separate escalations have now + proposed rebuilding it. Do not. +- **Flipping the gate-2 verdict.** That is Khaliq's read, not a run's. See + `ops/STATE.md`. +- **The analyze-agent clause.** Whether it is gate-2 or gate-4 scope is an + open question for Khaliq, recorded in `ops/STATE.md`. +- `.github/workflows/*`, `packages/sdk/*`, and `ops/*` other than this file. diff --git a/ops/STATE.md b/ops/STATE.md index 1fd6bb0e..089ed84f 100644 --- a/ops/STATE.md +++ b/ops/STATE.md @@ -8,7 +8,11 @@ it is authoritative when history is unavailable. **Keep it current. A stale STATE.md is worse than none:** it does not merely fail to help, it actively misleads an assessor that cannot check it. -Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mini), on `main`. **STATE.md gate-2 block rewritten; verdict unchanged (still AMBER).** A new evidence file — `ops/reviews/20260901-1050-gate2-live-run.md` — is cited from the gate-2 block; AMBER→GREEN is Khaliq's read on the enclosed evidence. +Last updated: 2026-09-16, on `main`, unwedging the drive loop. **Gate-2 clause 1 +corrected — it had been stale for two weeks (PR #122 closed it); deviations D1/D2 +recorded as clause 3; the 2026-09-16 escalation resolved and its content preserved +below. Verdict unchanged (still AMBER).** Prior entry: 2026-09-01 10:16 UTC, by the +Relayflow Lead (flows-lead-1 on sf-mini). **STATE.md gate-2 block rewritten; verdict unchanged (still AMBER).** A new evidence file — `ops/reviews/20260901-1050-gate2-live-run.md` — is cited from the gate-2 block; AMBER→GREEN is Khaliq's read on the enclosed evidence. ## Where the program is @@ -56,15 +60,19 @@ Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mi rather than restating specific numbers here (STATE.md counts drift, an evidence transcript does not). - **Why AMBER, not GREEN.** RFC-0001 §3 gate 2 has two clauses this - evidence does NOT close: - 1. **Trigger plane liveness-checked** (RFC-0001 §3 gate 2, the paragraph - ending "Native's silent-death problem"). RelayCron's deterministic-id - single-winner claim + `stale_after` sweep is the pattern. Not - implemented inside `relayflowd`. The poller runs; the kernel does not - yet notice if it stops. The same section calls this "a requirement, - not an option" — it is a stated done-when clause, not follow-up - hardening. + **Why AMBER, not GREEN.** RFC-0001 §3 gate 2 had two clauses this + evidence did NOT close. **Clause 1 has since closed; clause 2 has not, + and a third was always open.** + 1. **Trigger plane liveness-checked — CLOSED by PR #122 (`a774d880`).** + This block used to say "Not implemented inside `relayflowd`", and that + was true when it was written on 2026-09-01. It stopped being true when + #122 landed `kernel/relayflowd/src/server/liveness.rs` (the background + sweep: `detect_stale` + `latch_stale`, single-winner claim bucketed by + sweep id, `subscription.stale` journalled and emitted) wired end to end + by `kernel/relayflowd/tests/subscription_liveness.rs`. The stale entry + stood for two weeks and is exactly the failure this file warns about in + its own header: a stale STATE.md does not merely fail to help, it sends + an assessor to build something that already exists. 2. **The analyze-agent step actually executing.** In the recorded run, every step ended in `worker_error` because `hn-monitor start`'s AgentWorker has no user-supplied step handler. The dispatch loop @@ -72,6 +80,17 @@ Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mi ("runs succeed") or gate-4 scope ("chief-as-relayflow supplies the runtime") is Khaliq's call. + 3. **RFC-0001's own deviations D1 and D2** ("gate 2 cannot go green until + D1 and D2 are closed"). D1 is half closed: PR #252 (`97c886d2`) stopped + the resume path swallowing a journal scan error into + `wake_context: None`, but rule 10a's split by cause is not implemented — + `wake_context_unresolved` appears nowhere in the kernel, so transient + and permanent resolution failures are still handled identically. D2 is + blocked upstream: the RFC sequences it **engine-side epoch rollover → + D2 → gate 2** and says plainly that rollover, not D2, is the next + action. **The remainder of D1 is the current work package** — see + ops/NEXT.md. + **AMBER → GREEN is Khaliq's read** on the enclosed evidence, per this block's prior wording ("that is a judgement, not a missing part") and per the charter's standing rule that the Lead never merges / never @@ -80,6 +99,43 @@ Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mi which is a different thing). This session prepared evidence and left the flip pending. +## 2026-09-16 — the drive loop was wedged for four days, and why + +Recorded here because the escalation file that carried it has been deleted, +and this is the file an assessor with no git history reads instead. + +**Khaliq's decision: the next gate is GATE 2, not gate 3.** That settles the +question seven consecutive runs escalated about (PRs #417, #420, #422, #424, +#426, #427, #428 — every one of them a diff touching only ops/NEEDS_HUMAN.md +and ops/NEXT.md; none merged). It confirms what ops/AUTODRIVE_BRIEF.md +already said and what the gate ladder above already implied. + +**What the escalation actually said**, and what was true in it: an assessor +found ops/TARGET.md pinned to "gate 3" while describing gate-2 hn-monitor +work that PR #120 had already merged, and ops/NEXT.md describing unrelated +gate-3 review-swarm documentation. It could not tell which was real. It was +right to be confused — both sources were wrong, and it had no way to know. + +**Root cause, now fixed:** `ops/autodrive.sh` launched every run with +`ops/launch-gate.sh 3` while passing it the gate-2 brief, so the synthesised +target contradicted the brief inside it on every single tick. The assessors +were reporting a real defect in the launcher, once per run, for four days. + +**Two things that made it unrecoverable, both now fixed:** +- The escalation was sticky. `assess-gate` escalated on the mere EXISTENCE of + ops/NEEDS_HUMAN.md, which is committed on main and which nothing ever + deletes, so every future tick died at the gate on a question already + answered. The gate now requires two independent signals — the file exists + AND this tick wrote it. Pinned by ops/drive-assess-gate.test.mjs. +- The escalation cited evidence no human could inspect. ops/TARGET.md is + synthesised per run into a throwaway worktree by ops/launch-gate.sh and is + never in the delivered diff, so a reviewer sees quotations from a file that + does not exist. Review lenses failed PRs #422 and #428 for exactly this. + **Still open as a design question** — the drive.yaml assess task already + instructs the Lead to quote the target rather than cite its path, and that + instruction is evidently not enough. Whoever next touches the launcher + should consider committing the target into the delivered diff instead. + ## HANDOFF (2026-08-30 03:05 UTC) — a lead is LIVE on sf-mini **`flows-lead-1` is running on node `sf-mini`.** Attach with: diff --git a/ops/autodrive.sh b/ops/autodrive.sh index 776f04b3..345b752d 100644 --- a/ops/autodrive.sh +++ b/ops/autodrive.sh @@ -112,7 +112,13 @@ while [ ! -f "$STOP_FILE" ]; do sleep "$INTERVAL" continue fi - out=$(sh ops/launch-gate.sh 3 "$brief" 2>&1) + # The gate number must match the brief. This read `3` while + # ops/AUTODRIVE_BRIEF.md described GATE 2 work, so ops/launch-gate.sh + # synthesised "TARGET — gate 3" over a gate-2 brief and every assessor + # that read both correctly reported a contradiction it could not resolve. + # That is the conflict seven consecutive runs escalated about. Khaliq's + # decision on 2026-09-16: the next gate is 2, not 3. + out=$(sh ops/launch-gate.sh 2 "$brief" 2>&1) rid=$(echo "$out" | sed -n 's/^Run created: //p' | head -1) if [ -n "$rid" ]; then say "launched ${rid%%-*}" diff --git a/ops/drive-assess-gate.test.mjs b/ops/drive-assess-gate.test.mjs new file mode 100644 index 00000000..5b780674 --- /dev/null +++ b/ops/drive-assess-gate.test.mjs @@ -0,0 +1,130 @@ +// The assess-gate escalation test, executed rather than asserted about. +// +// ops/NEEDS_HUMAN.md was committed to main on 2026-09-06 (082c62aa) and +// nothing has ever deleted it. The gate escalated on the file's mere +// EXISTENCE, so from 2026-09-12 every drive tick exited 75 before doing any +// work: PRs #417, #420, #422, #424, #426, #427 and #428 are seven consecutive +// cloud runs whose entire diff is ops/NEEDS_HUMAN.md and ops/NEXT.md. +// +// The fix requires two independent signals to agree — the file exists AND +// this tick wrote it — so these tests drive BOTH directions. A gate that +// cannot fire is as broken as one that always fires, so the stale case and +// the live case are equally load-bearing here. +// +// The gate script is EXTRACTED from workflows/drive.yaml rather than copied, +// so this test fails if someone edits the workflow and not the test. No YAML +// parser is available at the repo root (there is no root package.json), and +// the block is a literal scalar at a known indent, so it is read as text. +import assert from 'node:assert/strict'; +import { execFileSync, spawnSync } from 'node:child_process'; +import { mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join, resolve } from 'node:path'; +import test from 'node:test'; + +const WORKFLOW = resolve(import.meta.dirname, '../workflows/drive.yaml'); + +/** Pull the `command:` literal block of a named deterministic step out of the workflow. */ +function gateScript(stepName) { + const lines = readFileSync(WORKFLOW, 'utf8').split('\n'); + const start = lines.findIndex((line) => line.trim() === `- name: ${stepName}`); + assert.notEqual(start, -1, `no step named ${stepName} in workflows/drive.yaml`); + const commandAt = lines.findIndex((line, i) => i > start && line.trim() === 'command: |'); + assert.notEqual(commandAt, -1, `step ${stepName} has no literal command block`); + const indent = lines[commandAt].search(/\S/) + 2; + const body = []; + for (const line of lines.slice(commandAt + 1)) { + if (line.trim() !== '' && line.search(/\S/) < indent) break; + body.push(line.slice(indent)); + } + const script = body.join('\n'); + assert.match(script, /NEEDS_HUMAN/, 'extracted the wrong block'); + return script; +} + +const git = (cwd, ...args) => + execFileSync('git', args, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] }); + +// A work package that satisfies the gate's later checks, so the only thing +// under test is the escalation branch. +const NEXT_MD = `# NEXT — a work package\n\n## Definition of done\n\n- \`cargo test --workspace\` must be green.\n`; + +/** + * Build a repo shaped like a drive tick: a `main` baseline, then a working + * branch that is "this tick". `staleEscalation` puts NEEDS_HUMAN.md on the + * baseline — the wedged state that burned seven runs. + */ +function tick(t, { staleEscalation = false } = {}) { + const root = mkdtempSync(join(tmpdir(), 'assess-gate-')); + t.after(() => rmSync(root, { recursive: true, force: true })); + git(root, 'init', '-q', '-b', 'main'); + git(root, 'config', 'user.email', 'test@relayflows.local'); + git(root, 'config', 'user.name', 'Gate Test'); + git(root, 'config', 'commit.gpgsign', 'false'); + execFileSync('mkdir', ['-p', join(root, 'ops')]); + writeFileSync(join(root, 'ops/NEXT.md'), NEXT_MD); + if (staleEscalation) { + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — answered long ago\n'); + } + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'base'); + git(root, 'checkout', '-q', '-b', 'flow/drive-tick'); + return root; +} + +function runGate(root) { + const result = spawnSync('sh', ['-c', gateScript('assess-gate')], { + cwd: root, + encoding: 'utf8', + }); + return { code: result.status, out: `${result.stdout}${result.stderr}` }; +} + +test('a stale escalation committed on main does not wedge the tick', (t) => { + const root = tick(t, { staleEscalation: true }); + // This tick writes its own work package, as assess does. + writeFileSync(join(root, 'ops/NEXT.md'), `${NEXT_MD}\nthis tick's package\n`); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'assess: work package for this tick'); + + const { code, out } = runGate(root); + assert.notEqual(code, 75, `the gate escalated on a stale file:\n${out}`); + assert.match(out, /ASSESS_STALE_NEEDS_HUMAN_IGNORED/); + assert.doesNotMatch(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); + assert.match(out, /ASSESS_GATE_PASS/); +}); + +test('an escalation committed by this tick still parks the run', (t) => { + const root = tick(t); + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — a live question\n\nWhich gate?\n'); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'assess: escalate'); + + const { code, out } = runGate(root); + assert.equal(code, 75, `a genuine escalation did not park the run:\n${out}`); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); + assert.match(out, /a live question/, 'the escalation body must reach the operator'); +}); + +test('an uncommitted escalation from this tick still parks the run', (t) => { + // Propagation between per-step sandboxes is lossy, so assess may write the + // file and fail to commit it. An escalation must not be lost that way. + const root = tick(t); + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — uncommitted but live\n'); + + const { code, out } = runGate(root); + assert.equal(code, 75, `an uncommitted live escalation was ignored:\n${out}`); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); +}); + +test('a tick with no escalation at all passes the gate', (t) => { + const root = tick(t); + writeFileSync(join(root, 'ops/NEXT.md'), `${NEXT_MD}\nthis tick's package\n`); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'assess: work package for this tick'); + + const { code, out } = runGate(root); + assert.equal(code, 0, `a clean tick did not pass:\n${out}`); + assert.match(out, /ASSESS_GATE_PASS/); + assert.doesNotMatch(out, /NEEDS_HUMAN/); +}); diff --git a/workflows/drive-cloud.yaml b/workflows/drive-cloud.yaml index 71e43972..ea2ba290 100644 --- a/workflows/drive-cloud.yaml +++ b/workflows/drive-cloud.yaml @@ -46,8 +46,8 @@ workflows: \ that already has one\n# (laptop, fleet node). Both shapes are supported below; neither is\n# assumed,\ \ and an unmaterialized sandbox fails closed and typed\n# rather than failing later as an unexplained\ \ tool error.\nset -eu\necho \"SYNC_WORKDIR=$(pwd)\"\n\nmissing=\"\"\nfor required in AGENTS.md\ - \ docs/RFC-0001-everything-is-a-relayflow.md ops/DIRECTIVES.md kernel packages/sdk; do\n [ -e \"$required\"\ - \ ] || missing=\"$missing $required\"\ndone\nif [ -n \"$missing\" ]; then\n echo \"SYNC_FAIL_NOT_MATERIALIZED:\ + \ docs/RFC-0001-everything-is-a-relayflow.md ops/DIRECTIVES.md kernel packages/sdk; do\n [ -e \"\ + $required\" ] || missing=\"$missing $required\"\ndone\nif [ -n \"$missing\" ]; then\n echo \"SYNC_FAIL_NOT_MATERIALIZED:\ \ the repo is not present in this execution environment.\" >&2\n echo \" missing:$missing\" >&2\n\ \ echo \" cwd: $(pwd)\" >&2\n echo \" A cloud run must upload the working tree: \\`agent-relay\ \ cloud run\\`\" >&2\n echo \" syncs code by default; \\`--no-sync-code\\` produces exactly this\ @@ -101,14 +101,14 @@ workflows: \ the known\nsandbox faults that are NOT reasons to block. Then read\nops/DIRECTIVES.md \u2014 standing\ \ human directives outrank the backlog;\nif one is unsatisfied, it IS the work package.\nThen read\ \ docs/bootstrap-report.md and ops/DRIVE-LOG.md if they exist,\n`git log --oneline -15`, `gh pr\ - \ list --state open` and open PR review\nstate, kernel/ and packages/sdk/ test status. Then write ops/NEXT.md:\ - \ the\nSINGLE highest-priority work package toward the current gate\n(gate 1 until its done-when\ - \ in RFC-0001 \xA73 holds), with: objective,\nfiles in scope, definition of done (must include passing\ - \ commands),\nand what is explicitly OUT of scope for this tick. If an open PR is\nawaiting fixes\ - \ from review, the work package is fixing it \u2014 never\nstart new work over unfinished work.\ - \ If work is blocked on a human\ndecision, write ops/NEEDS_HUMAN.md stating the exact question and\ - \ the\noptions \u2014 and then STILL end with ASSESS_DONE.\n\nCOMMIT YOUR WORK PACKAGE BEFORE YOU\ - \ FINISH:\n git add -A && git commit -m \"assess: work package for this tick\"\nEach step runs\ + \ list --state open` and open PR review\nstate, kernel/ and packages/sdk/ test status. Then write\ + \ ops/NEXT.md: the\nSINGLE highest-priority work package toward the current gate\n(gate 1 until\ + \ its done-when in RFC-0001 \xA73 holds), with: objective,\nfiles in scope, definition of done (must\ + \ include passing commands),\nand what is explicitly OUT of scope for this tick. If an open PR is\n\ + awaiting fixes from review, the work package is fixing it \u2014 never\nstart new work over unfinished\ + \ work. If work is blocked on a human\ndecision, write ops/NEEDS_HUMAN.md stating the exact question\ + \ and the\noptions \u2014 and then STILL end with ASSESS_DONE.\n\nCOMMIT YOUR WORK PACKAGE BEFORE\ + \ YOU FINISH:\n git add -A && git commit -m \"assess: work package for this tick\"\nEach step runs\ \ in its OWN sandbox and files reach the next step only\nthrough the executor's propagation, which\ \ is lossy: on runs a2089144\nand 2560e02d your predecessor wrote ops/NEXT.md, said so truthfully,\n\ and the file never arrived \u2014 one of those runs finished with a\nzero-file patch. Committing\ @@ -129,24 +129,46 @@ workflows: - assess-1 command: "# A typed park, not a crash. The assess step cannot express \"blocked\"\n# in its final\ \ token (its gate only recognises ASSESS_DONE), so the\n# Lead writes ops/NEEDS_HUMAN.md instead\ - \ and this step reads it.\nset -u\nif [ -f ops/NEEDS_HUMAN.md ]; then\n echo \"ASSESS_BLOCKED_NEEDS_HUMAN:\ - \ the Lead escalated a decision it cannot make.\"\n echo \"--- ops/NEEDS_HUMAN.md ---\"\n cat\ - \ ops/NEEDS_HUMAN.md\n exit 75\nfi\nif [ ! -f ops/NEXT.md ]; then\n echo \"ASSESS_FAIL: no ops/NEXT.md\ - \ \u2014 an assessment that named no work package did not assess\"\n exit 1\nfi\n# The assessment\ - \ must have WRITTEN this tick's package, not merely\n# left the previous one in place. On run 457a6102\ - \ assess reported\n# \"The work package is written to ops/NEXT.md\" and the very next step\n# read\ - \ the OLD file \u2014 the logs carry the reason:\n# \"relayfile flush failed after the command\ - \ succeeded (exit 1);\n# a later agent step may see stale files\"\n# The builder then correctly\ - \ refused to invent scope, but only after\n# a whole build step had been spent. Catch it here instead:\ - \ if\n# ops/NEXT.md is identical to the base, the assessment did not land,\n# whoever is at fault.\n\ - # Look for the package in the working tree OR in a commit made this\n# tick. Propagation between\ - \ per-step sandboxes is lossy, so a package\n# that exists only as a loose file may not arrive;\ - \ one committed by\n# the assess step travels in git history instead.\nif git log --oneline main..HEAD\ - \ -- ops/NEXT.md 2>/dev/null | grep -q .; then\n echo \"ASSESS_PACKAGE_COMMITTED: found ops/NEXT.md\ - \ change in this tick's history\"\nelif git diff --quiet main -- ops/NEXT.md 2>/dev/null; then\n\ - \ # Warn, do not fail. This was fatal, and it killed four runs in six\n # while the loop produced\ - \ nothing \u2014 a worse outcome than the risk\n # it guarded against.\n #\n # The risk it guarded\ - \ was \"the builder gets scope nobody wrote this\n # tick\". But scope does not actually come from\ + \ and this step reads it.\nset -u\n# An escalation is trusted only when TWO INDEPENDENT SIGNALS\ + \ AGREE:\n# the file exists AND this tick is what wrote it.\n#\n# Existence alone is not a signal.\ + \ ops/NEEDS_HUMAN.md was committed to\n# main on 2026-09-06 (082c62aa) and nothing in this repo\ + \ has ever\n# deleted it \u2014 no `rm`, no `git rm`, and `git log --diff-filter=D`\n# over that\ + \ path is empty. ops/launch-gate.sh builds each run's\n# worktree from origin/main and does not\ + \ strip it, so every tick from\n# 2026-09-12 onward escalated here before doing any work, on a\n\ + # question a human had already answered. PRs #417, #420, #422, #424,\n# #426, #427 and #428 are\ + \ seven consecutive cloud runs whose entire\n# diff is this file and ops/NEXT.md, re-litigating\ + \ the same conflict.\n# None merged. A full cloud run was burned on each.\n#\n# A stale file must\ + \ never be able to masquerade as a live escalation.\n# \"This tick\" is the same test the ops/NEXT.md\ + \ freshness check below\n# uses \u2014 a commit in `main..HEAD` \u2014 widened by the uncommitted\ + \ case,\n# because propagation between per-step sandboxes is lossy and assess\n# may write the file\ + \ and fail to commit it. Losing a live escalation\n# is the worse error of the two, so an unproven-fresh\ + \ file that is\n# dirty in the working tree still parks the run.\nif [ -f ops/NEEDS_HUMAN.md ];\ + \ then\n escalation=stale\n if git log --oneline main..HEAD -- ops/NEEDS_HUMAN.md 2>/dev/null\ + \ | grep -q .; then\n escalation=this_tick_committed\n elif [ -n \"$(git status --porcelain\ + \ -- ops/NEEDS_HUMAN.md 2>/dev/null)\" ]; then\n escalation=this_tick_uncommitted\n fi\n if\ + \ [ \"$escalation\" = stale ]; then\n echo \"ASSESS_STALE_NEEDS_HUMAN_IGNORED: ops/NEEDS_HUMAN.md\ + \ exists but this tick did not write it.\"\n echo \" It is the committed record of an escalation\ + \ that has already been answered,\"\n echo \" not a live one, so it does not park this run.\ + \ Delete it from main once its\"\n echo \" question is resolved \u2014 a resolved escalation\ + \ left in the tree is a lie the\"\n echo \" next assessor has to spend a run disproving.\"\n\ + \ else\n echo \"ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make ($escalation).\"\ + \n echo \"--- ops/NEEDS_HUMAN.md ---\"\n cat ops/NEEDS_HUMAN.md\n exit 75\n fi\nfi\nif\ + \ [ ! -f ops/NEXT.md ]; then\n echo \"ASSESS_FAIL: no ops/NEXT.md \u2014 an assessment that named\ + \ no work package did not assess\"\n exit 1\nfi\n# The assessment must have WRITTEN this tick's\ + \ package, not merely\n# left the previous one in place. On run 457a6102 assess reported\n# \"The\ + \ work package is written to ops/NEXT.md\" and the very next step\n# read the OLD file \u2014 the\ + \ logs carry the reason:\n# \"relayfile flush failed after the command succeeded (exit 1);\n#\ + \ a later agent step may see stale files\"\n# The builder then correctly refused to invent scope,\ + \ but only after\n# a whole build step had been spent. Catch it here instead: if\n# ops/NEXT.md\ + \ is identical to the base, the assessment did not land,\n# whoever is at fault.\n# Look for the\ + \ package in the working tree OR in a commit made this\n# tick. Propagation between per-step sandboxes\ + \ is lossy, so a package\n# that exists only as a loose file may not arrive; one committed by\n\ + # the assess step travels in git history instead.\nif git log --oneline main..HEAD -- ops/NEXT.md\ + \ 2>/dev/null | grep -q .; then\n echo \"ASSESS_PACKAGE_COMMITTED: found ops/NEXT.md change in\ + \ this tick's history\"\nelif git diff --quiet main -- ops/NEXT.md 2>/dev/null; then\n # Warn,\ + \ do not fail. This was fatal, and it killed four runs in six\n # while the loop produced nothing\ + \ \u2014 a worse outcome than the risk\n # it guarded against.\n #\n # The risk it guarded was\ + \ \"the builder gets scope nobody wrote this\n # tick\". But scope does not actually come from\ \ ops/NEXT.md: it comes\n # from ops/TARGET.md, which the launcher COMMITS into the uploaded\n\ \ # tree, so it is present in every per-step sandbox and cannot be\n # lost to the propagation\ \ fault. NEXT.md refines the target; it does\n # not define it.\n echo \"ASSESS_WARN_STALE_NEXT:\ @@ -221,66 +243,67 @@ workflows: \ test --workspace 2>&1); rc=$?\n echo \"$out\" | tail -8; ran=1\n if [ $rc -eq 124 ]; then\n\ \ echo \"VERIFY_SUITE_TIMEOUT: the kernel suite exceeded ${VERIFY_SUITE_TIMEOUT:-900}s and was\ \ killed.\"\n echo \" A hanging test is a defect, not a pass \u2014 recording it as a failure.\"\ - \n fi\n [ $rc -eq 0 ] || ok=1\nfi\nif [ -f packages/sdk/package.json ]; then\n # node_modules is not in\ - \ `git ls-files`, so a sandbox has none.\n # Install before testing, and treat a failed install\ - \ as a failed\n # verify rather than letting `npm test` report a confusing error.\n if [ ! -d\ - \ packages/sdk/node_modules ]; then\n echo \"VERIFY_INSTALL: packages/sdk/node_modules absent \u2014 installing\"\ - \n out=$(cd packages/sdk && run_bounded \"npm ci\" npm ci 2>&1); rc=$?\n if [ $rc -ne 0 ]; then\n \ - \ echo \"$out\" | tail -12\n echo \"VERIFY_FAIL: npm ci failed \u2014 cannot test what\ - \ did not install\"\n exit 1\n fi\n fi\n # Restore exec bits. `npm ci` reported success\ - \ (96 packages) and\n # esbuild still failed with EACCES on run 909e18f6, so the mount\n # this\ - \ installs onto does not carry the executable bit. That is\n # the same fault that left ops/cargo.sh\ - \ non-executable. Which\n # layer drops it is NOT established; this repairs the symptom\n # where\ - \ it is observed and is a no-op where the bit survives.\n chmod -R +x packages/sdk/node_modules/.bin 2>/dev/null\ - \ || true\n find packages/sdk/node_modules -type d -name bin -path \"*esbuild*\" \\\n -exec chmod -R\ - \ +x {} + 2>/dev/null || true\n out=$(cd packages/sdk && run_bounded \"sdk suite\" npm test 2>&1); rc=$?\n\ - \ echo \"$out\" | tail -8; ran=1\n if [ $rc -eq 124 ]; then\n echo \"VERIFY_SUITE_TIMEOUT:\ - \ the sdk suite exceeded ${VERIFY_SUITE_TIMEOUT:-900}s and was killed.\"\n fi\n [ $rc -eq 0 ]\ - \ || ok=1\nfi\nif [ \"$ran\" -eq 0 ]; then\n echo \"VERIFY_FAIL: no suite found \u2014 a verify\ - \ that runs nothing cannot pass\"; exit 1\nfi\n# Enforce the NEXT.md contract instead of asking\ - \ for it.\n#\n# The assess prompt has told runs since PR #19 to quote their scope\n# rather than\ - \ cite ops/TARGET.md, because that file lives only in the\n# throwaway launch worktree and is not\ - \ in the delivered diff. Runs\n# kept citing it anyway \u2014 reviewers filed the same finding again\ - \ on\n# #35, #40 and #48. Four recurrences after the warning is enough\n# evidence that prose does\ - \ not hold and a check does.\n#\n# Same for unevidenced test claims: \"all three tests pass\" with\ - \ no\n# transcript. validateNextWorkPackage refuses both shapes.\n#\n# Runs BEFORE the node_modules\ - \ cleanup below, which needs packages/sdk/dist.\nif [ -f ops/NEXT.md ] && [ -f packages/sdk/dist/index.js ]; then\n\ - \ nextout=$(node -e '\n const fs = require(\"node:fs\");\n const { validateNextWorkPackage\ - \ } = require(\"./packages/sdk/dist/index.js\");\n const verdict = validateNextWorkPackage(\n fs.readFileSync(\"\ - ops/NEXT.md\", \"utf8\"),\n (p) => fs.existsSync(p),\n );\n if (!verdict.accepted) {\n\ - \ console.log(\"NEXT_REFUSED \" + verdict.reason);\n process.exit(1);\n }\n console.log(\"\ - NEXT_OK\");\n ' 2>&1) || true\n echo \"$nextout\" | tail -2\n case \"$nextout\" in\n *NEXT_REFUSED*)\n\ - \ echo \"VERIFY_FAIL: ops/NEXT.md was refused \u2014 quote your scope instead of citing\"\n\ - \ echo \" ops/TARGET.md (it is not in the delivered diff), and paste the literal\"\n \ - \ echo \" command output for anything you claim passes.\"\n ok=1\n ;;\n esac\nfi\n\n\ - # Shrink the tree before the flush that matters.\n#\n# The relayfile mount flush fails on a large\ - \ tree \u2014 http 413, or a\n# timeout waiting for the daemon to ack \u2014 and the failure is\n\ - # NON-FATAL, so the run reports success while later steps read stale\n# files and the delivered\ - \ patch silently loses the work. Five runs\n# were lost this way. PR #38 took kernel/target out\ - \ (4914 paths ->\n# 0), which was not enough on its own: run 5ecf7078 still flushed\n# 1932 files\ - \ and still hit 413.\n#\n# Dependency and compiler outputs are already gitignored, and nothing\n\ - # downstream of here reads them: the suites and NEXT validation have\n# finished by this line, while\ - \ commit/handoff touch tracked files.\n# Removing them after the gates run costs nothing and keeps\ - \ them out\n# of every flush from here on.\n#\n# Runs on pass AND fail. The first version only slimmed\ - \ on a green\n# verify, to keep a failing run's tree intact for diagnosis. That\n# defeated the\ - \ fix: run ae982aaa ended VERIFY_FAIL_NONFATAL, so the\n# cleanup was skipped, the tree stayed at\ - \ 1842 files and the flush\n# failed with three 413s \u2014 and the run still committed and delivered\n\ - # a PR. A run that delivers needs its flush to work whether or not\n# its gates passed, and node_modules\ - \ is reinstallable, so there was\n# never anything to preserve for diagnosis here.\n#\n# Guarded\ - \ with || true and placed AFTER the pass/fail decision is\n# computed, so a cleanup problem can\ - \ never change the verdict.\nrm -rf packages/sdk/node_modules packages/sdk/dist packages/surface/dist 2>/dev/null || true\n\ - echo \"VERIFY_TREE_SLIMMED: removed packages/sdk/node_modules and generated dist trees before flush (ok=$ok)\"\ - \n\n# Measure the tree instead of guessing at it.\n#\n# The relayfile flush keeps failing (http\ - \ 413) and three successive\n# hypotheses about WHY were wrong: kernel/target was removed and the\n\ - # count barely moved; node_modules was removed and it still failed;\n# HOME turned out to be /home/daytona,\ - \ outside the tree, so the\n# toolchain was never in it. Meanwhile the bootstrap reports ~1900\n\ - # changed files against a repo of only 257 tracked ones, and the\n# delivered patch is 12 files.\ - \ Nothing in the log says what the rest\n# are, so this prints it rather than inviting a fourth\ - \ guess.\necho \"TREE_CENSUS: top directories by file count\"\nfind . -type f -not -path './.git/*'\ - \ 2>/dev/null \\\n | sed -E 's|^\\./||; s|/[^/]*$||' \\\n | cut -d/ -f1-2 \\\n | sort | uniq\ - \ -c | sort -rn | head -12 || true\necho \"TREE_CENSUS_TOTAL: $(find . -type f -not -path './.git/*'\ - \ 2>/dev/null | wc -l | tr -d ' ')\"\n[ \"$ok\" -eq 0 ] && echo VERIFY_PASS || echo \"VERIFY_FAIL_NONFATAL:\ - \ recorded; the next cycle must address it\"\n" + \n fi\n [ $rc -eq 0 ] || ok=1\nfi\nif [ -f packages/sdk/package.json ]; then\n # node_modules\ + \ is not in `git ls-files`, so a sandbox has none.\n # Install before testing, and treat a failed\ + \ install as a failed\n # verify rather than letting `npm test` report a confusing error.\n if\ + \ [ ! -d packages/sdk/node_modules ]; then\n echo \"VERIFY_INSTALL: packages/sdk/node_modules\ + \ absent \u2014 installing\"\n out=$(cd packages/sdk && run_bounded \"npm ci\" npm ci 2>&1);\ + \ rc=$?\n if [ $rc -ne 0 ]; then\n echo \"$out\" | tail -12\n echo \"VERIFY_FAIL: npm\ + \ ci failed \u2014 cannot test what did not install\"\n exit 1\n fi\n fi\n # Restore exec\ + \ bits. `npm ci` reported success (96 packages) and\n # esbuild still failed with EACCES on run\ + \ 909e18f6, so the mount\n # this installs onto does not carry the executable bit. That is\n #\ + \ the same fault that left ops/cargo.sh non-executable. Which\n # layer drops it is NOT established;\ + \ this repairs the symptom\n # where it is observed and is a no-op where the bit survives.\n chmod\ + \ -R +x packages/sdk/node_modules/.bin 2>/dev/null || true\n find packages/sdk/node_modules -type\ + \ d -name bin -path \"*esbuild*\" \\\n -exec chmod -R +x {} + 2>/dev/null || true\n out=$(cd\ + \ packages/sdk && run_bounded \"sdk suite\" npm test 2>&1); rc=$?\n echo \"$out\" | tail -8; ran=1\n\ + \ if [ $rc -eq 124 ]; then\n echo \"VERIFY_SUITE_TIMEOUT: the sdk suite exceeded ${VERIFY_SUITE_TIMEOUT:-900}s\ + \ and was killed.\"\n fi\n [ $rc -eq 0 ] || ok=1\nfi\nif [ \"$ran\" -eq 0 ]; then\n echo \"VERIFY_FAIL:\ + \ no suite found \u2014 a verify that runs nothing cannot pass\"; exit 1\nfi\n# Enforce the NEXT.md\ + \ contract instead of asking for it.\n#\n# The assess prompt has told runs since PR #19 to quote\ + \ their scope\n# rather than cite ops/TARGET.md, because that file lives only in the\n# throwaway\ + \ launch worktree and is not in the delivered diff. Runs\n# kept citing it anyway \u2014 reviewers\ + \ filed the same finding again on\n# #35, #40 and #48. Four recurrences after the warning is enough\n\ + # evidence that prose does not hold and a check does.\n#\n# Same for unevidenced test claims: \"\ + all three tests pass\" with no\n# transcript. validateNextWorkPackage refuses both shapes.\n#\n\ + # Runs BEFORE the node_modules cleanup below, which needs packages/sdk/dist.\nif [ -f ops/NEXT.md\ + \ ] && [ -f packages/sdk/dist/index.js ]; then\n nextout=$(node -e '\n const fs = require(\"\ + node:fs\");\n const { validateNextWorkPackage } = require(\"./packages/sdk/dist/index.js\");\n\ + \ const verdict = validateNextWorkPackage(\n fs.readFileSync(\"ops/NEXT.md\", \"utf8\"),\n\ + \ (p) => fs.existsSync(p),\n );\n if (!verdict.accepted) {\n console.log(\"NEXT_REFUSED\ + \ \" + verdict.reason);\n process.exit(1);\n }\n console.log(\"NEXT_OK\");\n ' 2>&1)\ + \ || true\n echo \"$nextout\" | tail -2\n case \"$nextout\" in\n *NEXT_REFUSED*)\n echo\ + \ \"VERIFY_FAIL: ops/NEXT.md was refused \u2014 quote your scope instead of citing\"\n echo\ + \ \" ops/TARGET.md (it is not in the delivered diff), and paste the literal\"\n echo \" command\ + \ output for anything you claim passes.\"\n ok=1\n ;;\n esac\nfi\n\n# Shrink the tree\ + \ before the flush that matters.\n#\n# The relayfile mount flush fails on a large tree \u2014 http\ + \ 413, or a\n# timeout waiting for the daemon to ack \u2014 and the failure is\n# NON-FATAL, so\ + \ the run reports success while later steps read stale\n# files and the delivered patch silently\ + \ loses the work. Five runs\n# were lost this way. PR #38 took kernel/target out (4914 paths ->\n\ + # 0), which was not enough on its own: run 5ecf7078 still flushed\n# 1932 files and still hit 413.\n\ + #\n# Dependency and compiler outputs are already gitignored, and nothing\n# downstream of here reads\ + \ them: the suites and NEXT validation have\n# finished by this line, while commit/handoff touch\ + \ tracked files.\n# Removing them after the gates run costs nothing and keeps them out\n# of every\ + \ flush from here on.\n#\n# Runs on pass AND fail. The first version only slimmed on a green\n#\ + \ verify, to keep a failing run's tree intact for diagnosis. That\n# defeated the fix: run ae982aaa\ + \ ended VERIFY_FAIL_NONFATAL, so the\n# cleanup was skipped, the tree stayed at 1842 files and the\ + \ flush\n# failed with three 413s \u2014 and the run still committed and delivered\n# a PR. A run\ + \ that delivers needs its flush to work whether or not\n# its gates passed, and node_modules is\ + \ reinstallable, so there was\n# never anything to preserve for diagnosis here.\n#\n# Guarded with\ + \ || true and placed AFTER the pass/fail decision is\n# computed, so a cleanup problem can never\ + \ change the verdict.\nrm -rf packages/sdk/node_modules packages/sdk/dist packages/surface/dist\ + \ 2>/dev/null || true\necho \"VERIFY_TREE_SLIMMED: removed packages/sdk/node_modules and generated\ + \ dist trees before flush (ok=$ok)\"\n\n# Measure the tree instead of guessing at it.\n#\n# The\ + \ relayfile flush keeps failing (http 413) and three successive\n# hypotheses about WHY were wrong:\ + \ kernel/target was removed and the\n# count barely moved; node_modules was removed and it still\ + \ failed;\n# HOME turned out to be /home/daytona, outside the tree, so the\n# toolchain was never\ + \ in it. Meanwhile the bootstrap reports ~1900\n# changed files against a repo of only 257 tracked\ + \ ones, and the\n# delivered patch is 12 files. Nothing in the log says what the rest\n# are, so\ + \ this prints it rather than inviting a fourth guess.\necho \"TREE_CENSUS: top directories by file\ + \ count\"\nfind . -type f -not -path './.git/*' 2>/dev/null \\\n | sed -E 's|^\\./||; s|/[^/]*$||'\ + \ \\\n | cut -d/ -f1-2 \\\n | sort | uniq -c | sort -rn | head -12 || true\necho \"TREE_CENSUS_TOTAL:\ + \ $(find . -type f -not -path './.git/*' 2>/dev/null | wc -l | tr -d ' ')\"\n[ \"$ok\" -eq 0 ] &&\ + \ echo VERIFY_PASS || echo \"VERIFY_FAIL_NONFATAL: recorded; the next cycle must address it\"\n" timeoutMs: 1200000 - name: commit-1 type: deterministic diff --git a/workflows/drive.yaml b/workflows/drive.yaml index a52be0aa..93d31529 100644 --- a/workflows/drive.yaml +++ b/workflows/drive.yaml @@ -191,11 +191,46 @@ workflows: # in its final token (its gate only recognises ASSESS_DONE), so the # Lead writes ops/NEEDS_HUMAN.md instead and this step reads it. set -u + # An escalation is trusted only when TWO INDEPENDENT SIGNALS AGREE: + # the file exists AND this tick is what wrote it. + # + # Existence alone is not a signal. ops/NEEDS_HUMAN.md was committed to + # main on 2026-09-06 (082c62aa) and nothing in this repo has ever + # deleted it — no `rm`, no `git rm`, and `git log --diff-filter=D` + # over that path is empty. ops/launch-gate.sh builds each run's + # worktree from origin/main and does not strip it, so every tick from + # 2026-09-12 onward escalated here before doing any work, on a + # question a human had already answered. PRs #417, #420, #422, #424, + # #426, #427 and #428 are seven consecutive cloud runs whose entire + # diff is this file and ops/NEXT.md, re-litigating the same conflict. + # None merged. A full cloud run was burned on each. + # + # A stale file must never be able to masquerade as a live escalation. + # "This tick" is the same test the ops/NEXT.md freshness check below + # uses — a commit in `main..HEAD` — widened by the uncommitted case, + # because propagation between per-step sandboxes is lossy and assess + # may write the file and fail to commit it. Losing a live escalation + # is the worse error of the two, so an unproven-fresh file that is + # dirty in the working tree still parks the run. if [ -f ops/NEEDS_HUMAN.md ]; then - echo "ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make." - echo "--- ops/NEEDS_HUMAN.md ---" - cat ops/NEEDS_HUMAN.md - exit 75 + escalation=stale + if git log --oneline main..HEAD -- ops/NEEDS_HUMAN.md 2>/dev/null | grep -q .; then + escalation=this_tick_committed + elif [ -n "$(git status --porcelain -- ops/NEEDS_HUMAN.md 2>/dev/null)" ]; then + escalation=this_tick_uncommitted + fi + if [ "$escalation" = stale ]; then + echo "ASSESS_STALE_NEEDS_HUMAN_IGNORED: ops/NEEDS_HUMAN.md exists but this tick did not write it." + echo " It is the committed record of an escalation that has already been answered," + echo " not a live one, so it does not park this run. Delete it from main once its" + echo " question is resolved — a resolved escalation left in the tree is a lie the" + echo " next assessor has to spend a run disproving." + else + echo "ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make ($escalation)." + echo "--- ops/NEEDS_HUMAN.md ---" + cat ops/NEEDS_HUMAN.md + exit 75 + fi fi if [ ! -f ops/NEXT.md ]; then echo "ASSESS_FAIL: no ops/NEXT.md — an assessment that named no work package did not assess" From 89ff6ce3b0c250e6637847057733b5761b68f131 Mon Sep 17 00:00:00 2001 From: kjgbot Date: Wed, 16 Sep 2026 07:17:41 -0700 Subject: [PATCH 2/2] fix(drive): park the run when escalation freshness is unprovable MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Bugbot found a real defect in the first commit, and it was the one that mattered: the freshness check could not tell "git says this file is stale" from "git could not answer". Both printed nothing, so with no .git, with main absent, or on any git failure, a LIVE escalation classified as stale and the builder walked straight past a human decision — inverting the tradeoff the comment right above it claims to make. A sandbox is exactly where git cannot answer. SYNC_MODE=snapshot runs `git init` over an extracted tarball, so main does not exist until sync creates it, and the gate would have dropped live escalations there. The default is now to TRUST the escalation. Only a positive, SUCCESSFUL answer from git downgrades it to stale: the tree must be a repo, main must resolve, and both `git log` and `git status` must exit 0. Anything else prints ASSESS_ESCALATION_FRESHNESS_UNPROVABLE and exits 75, because ignoring a real escalation is the worse of the two errors. Exit codes are now checked rather than inferred from empty output, which also drops the `| grep -q .` that silently swallowed git's own exit status. Two tests cover the shapes Bugbot correctly noted were unexercised: no git repo at all, and a repo whose branch is not main with the escalation COMMITTED (the shape where a naive main..HEAD prints nothing and the file looks stale). Both fail against 7a17d31b and pass here. Co-Authored-By: Claude Opus 5 (1M context) --- ops/drive-assess-gate.test.mjs | 44 ++++++++++++++++++++++++++++++++++ workflows/drive-cloud.yaml | 33 +++++++++++++++++-------- workflows/drive.yaml | 38 +++++++++++++++++++++++++---- 3 files changed, 100 insertions(+), 15 deletions(-) diff --git a/ops/drive-assess-gate.test.mjs b/ops/drive-assess-gate.test.mjs index 5b780674..c3cc83c8 100644 --- a/ops/drive-assess-gate.test.mjs +++ b/ops/drive-assess-gate.test.mjs @@ -117,6 +117,50 @@ test('an uncommitted escalation from this tick still parks the run', (t) => { assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); }); +/** + * A tree carrying a live escalation that git cannot describe. `initGit: true` + * gives a real repo whose branch is not `main`; otherwise there is no repo at + * all. Both are shapes a sandbox actually produces — SYNC_MODE=snapshot runs + * `git init` over an extracted tarball, so `main` does not exist until sync + * creates it. + */ +function unprovableTick(t, { initGit = false } = {}) { + const root = mkdtempSync(join(tmpdir(), 'assess-gate-')); + t.after(() => rmSync(root, { recursive: true, force: true })); + execFileSync('mkdir', ['-p', join(root, 'ops')]); + writeFileSync(join(root, 'ops/NEXT.md'), NEXT_MD); + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — live, on a tree git cannot describe\n'); + if (initGit) { + git(root, 'init', '-q', '-b', 'work'); + git(root, 'config', 'user.email', 'test@relayflows.local'); + git(root, 'config', 'user.name', 'Gate Test'); + git(root, 'config', 'commit.gpgsign', 'false'); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'base on a branch that is not main'); + } + return root; +} + +test('a live escalation parks the run when there is no git repo at all', (t) => { + const root = unprovableTick(t); + const { code, out } = runGate(root); + assert.equal(code, 75, `freshness was unprovable but the run continued:\n${out}`); + assert.match(out, /ASSESS_ESCALATION_FRESHNESS_UNPROVABLE/); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); + assert.doesNotMatch(out, /ASSESS_STALE_NEEDS_HUMAN_IGNORED/); +}); + +test('a live escalation parks the run when main does not exist', (t) => { + // The escalation is COMMITTED here, so a naive `git log main..HEAD` prints + // nothing and the file looks stale — the precise shape that would drop a + // live escalation in a snapshot sandbox. + const root = unprovableTick(t, { initGit: true }); + const { code, out } = runGate(root); + assert.equal(code, 75, `main was missing but the run continued:\n${out}`); + assert.match(out, /ASSESS_ESCALATION_FRESHNESS_UNPROVABLE/); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); +}); + test('a tick with no escalation at all passes the gate', (t) => { const root = tick(t); writeFileSync(join(root, 'ops/NEXT.md'), `${NEXT_MD}\nthis tick's package\n`); diff --git a/workflows/drive-cloud.yaml b/workflows/drive-cloud.yaml index ea2ba290..5e6c034e 100644 --- a/workflows/drive-cloud.yaml +++ b/workflows/drive-cloud.yaml @@ -142,16 +142,29 @@ workflows: \ freshness check below\n# uses \u2014 a commit in `main..HEAD` \u2014 widened by the uncommitted\ \ case,\n# because propagation between per-step sandboxes is lossy and assess\n# may write the file\ \ and fail to commit it. Losing a live escalation\n# is the worse error of the two, so an unproven-fresh\ - \ file that is\n# dirty in the working tree still parks the run.\nif [ -f ops/NEEDS_HUMAN.md ];\ - \ then\n escalation=stale\n if git log --oneline main..HEAD -- ops/NEEDS_HUMAN.md 2>/dev/null\ - \ | grep -q .; then\n escalation=this_tick_committed\n elif [ -n \"$(git status --porcelain\ - \ -- ops/NEEDS_HUMAN.md 2>/dev/null)\" ]; then\n escalation=this_tick_uncommitted\n fi\n if\ - \ [ \"$escalation\" = stale ]; then\n echo \"ASSESS_STALE_NEEDS_HUMAN_IGNORED: ops/NEEDS_HUMAN.md\ - \ exists but this tick did not write it.\"\n echo \" It is the committed record of an escalation\ - \ that has already been answered,\"\n echo \" not a live one, so it does not park this run.\ - \ Delete it from main once its\"\n echo \" question is resolved \u2014 a resolved escalation\ - \ left in the tree is a lie the\"\n echo \" next assessor has to spend a run disproving.\"\n\ - \ else\n echo \"ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make ($escalation).\"\ + \ file that is\n# dirty in the working tree still parks the run.\n#\n# The default is to TRUST the\ + \ escalation. Only a positive, SUCCESSFUL\n# answer from git may downgrade it to stale, because\ + \ \"git printed\n# nothing\" and \"git could not answer\" look identical otherwise \u2014 and\n\ + # a sandbox is exactly where git cannot answer. SYNC_MODE=snapshot\n# runs `git init` over an extracted\ + \ tarball, so a step that runs\n# before main exists, a missing .git, or any git failure would\n\ + # silently classify a LIVE escalation as stale and walk the builder\n# straight past a human decision.\ + \ That inverts the tradeoff above,\n# so an unprovable escalation parks the run.\nif [ -f ops/NEEDS_HUMAN.md\ + \ ]; then\n escalation=unprovable\n if git rev-parse --git-dir >/dev/null 2>&1 \\\n && git\ + \ rev-parse --verify --quiet main >/dev/null 2>&1; then\n tick_log=$(git log --oneline main..HEAD\ + \ -- ops/NEEDS_HUMAN.md 2>/dev/null)\n if [ $? -ne 0 ]; then\n escalation=unprovable\n \ + \ elif [ -n \"$tick_log\" ]; then\n escalation=this_tick_committed\n else\n tick_dirty=$(git\ + \ status --porcelain -- ops/NEEDS_HUMAN.md 2>/dev/null)\n if [ $? -ne 0 ]; then\n escalation=unprovable\n\ + \ elif [ -n \"$tick_dirty\" ]; then\n escalation=this_tick_uncommitted\n else\n\ + \ escalation=stale\n fi\n fi\n fi\n if [ \"$escalation\" = stale ]; then\n echo\ + \ \"ASSESS_STALE_NEEDS_HUMAN_IGNORED: ops/NEEDS_HUMAN.md exists but this tick did not write it.\"\ + \n echo \" It is the committed record of an escalation that has already been answered,\"\n \ + \ echo \" not a live one, so it does not park this run. Delete it from main once its\"\n echo\ + \ \" question is resolved \u2014 a resolved escalation left in the tree is a lie the\"\n echo\ + \ \" next assessor has to spend a run disproving.\"\n else\n if [ \"$escalation\" = unprovable\ + \ ]; then\n echo \"ASSESS_ESCALATION_FRESHNESS_UNPROVABLE: git could not say whether this tick\"\ + \n echo \" wrote ops/NEEDS_HUMAN.md (no repo, no main, or git failed). Failing safe and\"\n\ + \ echo \" treating it as live: ignoring a real escalation is the worse of the two errors.\"\ + \n fi\n echo \"ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make ($escalation).\"\ \n echo \"--- ops/NEEDS_HUMAN.md ---\"\n cat ops/NEEDS_HUMAN.md\n exit 75\n fi\nfi\nif\ \ [ ! -f ops/NEXT.md ]; then\n echo \"ASSESS_FAIL: no ops/NEXT.md \u2014 an assessment that named\ \ no work package did not assess\"\n exit 1\nfi\n# The assessment must have WRITTEN this tick's\ diff --git a/workflows/drive.yaml b/workflows/drive.yaml index 93d31529..ccf2b799 100644 --- a/workflows/drive.yaml +++ b/workflows/drive.yaml @@ -212,12 +212,35 @@ workflows: # may write the file and fail to commit it. Losing a live escalation # is the worse error of the two, so an unproven-fresh file that is # dirty in the working tree still parks the run. + # + # The default is to TRUST the escalation. Only a positive, SUCCESSFUL + # answer from git may downgrade it to stale, because "git printed + # nothing" and "git could not answer" look identical otherwise — and + # a sandbox is exactly where git cannot answer. SYNC_MODE=snapshot + # runs `git init` over an extracted tarball, so a step that runs + # before main exists, a missing .git, or any git failure would + # silently classify a LIVE escalation as stale and walk the builder + # straight past a human decision. That inverts the tradeoff above, + # so an unprovable escalation parks the run. if [ -f ops/NEEDS_HUMAN.md ]; then - escalation=stale - if git log --oneline main..HEAD -- ops/NEEDS_HUMAN.md 2>/dev/null | grep -q .; then - escalation=this_tick_committed - elif [ -n "$(git status --porcelain -- ops/NEEDS_HUMAN.md 2>/dev/null)" ]; then - escalation=this_tick_uncommitted + escalation=unprovable + if git rev-parse --git-dir >/dev/null 2>&1 \ + && git rev-parse --verify --quiet main >/dev/null 2>&1; then + tick_log=$(git log --oneline main..HEAD -- ops/NEEDS_HUMAN.md 2>/dev/null) + if [ $? -ne 0 ]; then + escalation=unprovable + elif [ -n "$tick_log" ]; then + escalation=this_tick_committed + else + tick_dirty=$(git status --porcelain -- ops/NEEDS_HUMAN.md 2>/dev/null) + if [ $? -ne 0 ]; then + escalation=unprovable + elif [ -n "$tick_dirty" ]; then + escalation=this_tick_uncommitted + else + escalation=stale + fi + fi fi if [ "$escalation" = stale ]; then echo "ASSESS_STALE_NEEDS_HUMAN_IGNORED: ops/NEEDS_HUMAN.md exists but this tick did not write it." @@ -226,6 +249,11 @@ workflows: echo " question is resolved — a resolved escalation left in the tree is a lie the" echo " next assessor has to spend a run disproving." else + if [ "$escalation" = unprovable ]; then + echo "ASSESS_ESCALATION_FRESHNESS_UNPROVABLE: git could not say whether this tick" + echo " wrote ops/NEEDS_HUMAN.md (no repo, no main, or git failed). Failing safe and" + echo " treating it as live: ignoring a real escalation is the worse of the two errors." + fi echo "ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make ($escalation)." echo "--- ops/NEEDS_HUMAN.md ---" cat ops/NEEDS_HUMAN.md