diff --git a/ops/AUTODRIVE_BRIEF.md b/ops/AUTODRIVE_BRIEF.md index 81494eb3f..3560064db 100644 --- a/ops/AUTODRIVE_BRIEF.md +++ b/ops/AUTODRIVE_BRIEF.md @@ -1,82 +1,36 @@ -Build sub-PR A of the Gate 2 push: a real `hn-monitor` polling runner in the SDK. CODE task, `sdk/src/`-side. This is a scaffolding PR — proof that the workload EXECUTES end-to-end is deliberately deferred to sub-PR B (integration test). Do not conflate the two. +Close the rest of RFC-0001 deviation D1: classify a wake-context resolution failure by cause and journal it. GATE 2. CODE task, `kernel/`-side (Rust). -**Context:** RFC-0001 §3 gate 2 is done when "hn-monitor runs as a relayflow in production, triggered by its real events, with zero bespoke persistence." Every primitive already exists in this repo — event triggers (PR #14, `kernel/relayflowd/tests/event_wake.rs`), the flow spec (`testdata/hn-monitor.flow.yaml`), the poller (`sdk/src/hn-poller.ts`), the agent worker (`sdk/src/worker.ts` from PR #53), a one-shot demo (`sdk/src/demo-hn-monitor.ts`) — but nothing has ever run them together as a continuous workload. This PR fixes that. +**Retargeted 2026-09-16.** This brief previously asked for a new hn-monitor runner module in the SDK, as a gate-2 scaffolding PR. PR #120 shipped that work on 2026-09-01 as `packages/sdk/src/cli/hn-monitor.ts`. The brief outlived its work package by two weeks, and because `ops/autodrive.sh` also launched it under a hardcoded "gate 3", every assessor that read both correctly reported a contradiction and escalated instead of building. Seven consecutive runs (#417, #420, #422, #424, #426, #427, #428) produced nothing but escalation notes. **Do not rebuild the hn-monitor runner. It exists.** -**Prior attempt (PR #83, closed):** produced a functional runner but was rejected by the swarm on five real findings. Address them in this attempt: +**Context:** RFC-0001 states that gate 2 stays RED until deviations D1 and D2 are both closed. D2 is blocked upstream on engine-side epoch rollover — the RFC sequences it rollover → D2 → gate 2 and says D2 is not implementable in isolation, because every construction of `EpochSummaryPayload` in the kernel is inside a test. That leaves D1 as the actionable half, and it is already half done: PR #252 stopped `resolve_wake_context` in `kernel/relayflowd/src/engine/drive.rs` swallowing a journal scan error into `wake_context: None`. What remains is the classification that function's own doc comment says is missing — "transient → retry the attempt under its budget; permanent → park `needs_human`, and none of that classification exists yet". -1. **Fail-closed on journal errors.** #83's `catch (err) { onPollError(err) }` swallowed EVERY error including `eventSubmit` journal failures — violates covenant 2 (fail-closed) and RFC-0001 §1. Only fetch-level errors (network flakiness, HN API rate limits) may be swallowed; a journal write failure MUST throw and terminate the runner. Split: `try { fetch } catch { onFetchError }` around the network call, `try { eventSubmit } catch { rethrow }` around the journal call. +**The task.** Implement RFC rule 10a: -2. **AgentWorker.close() must release the worker (or explicitly document it does not).** #83 added `await worker.close()` to shutdown but the current `close()` only drains local promises — it does NOT tell the kernel to release the worker registration. Either: - - Add a `workerRelease` verb to `sdk/src/protocol.ts` and call it from `close()` (preferred — completes the shutdown contract), OR - - Add a one-line comment on `close()` naming exactly what shutdown intentionally does NOT do + - a **transient** failure (I/O error, lock contention, a truncated tail still being written) fails the attempt and stays **retryable** under the step's ordinary retry budget, journalling `wake_context_unresolved` with `reason: "transient"` and the underlying error + - a **permanent** failure (the run is open on a wake, the current segment reads cleanly, and no wake context is present) fails the step and parks as **`needs_human`**, not retryable, journalling `wake_context_unresolved` with `reason: "absent"` + - neither may fall back to `wake_context: None` — dispatching with no context is reserved for runs that were never woken -3. **Class field declaration order.** #83 declared `private readonly fetcher` AFTER the constructor. Works today because of ES2022 hoisting semantics but breaks silently if someone adds `= someDefault` to a declaration. Declare ALL fields at the top of the class body, before the constructor. +`wake_context_unresolved` does not exist anywhere in the kernel today, so it needs a typed entry beside `SubscriptionMatched` in `kernel/relayflowd-core/src/entry.rs`. -4. **Signal handlers must be opt-in via AbortSignal.** #83 registered `SIGTERM`/`SIGINT` handlers on the process directly with no opt-out. A library user embedding this can't cancel one runner without affecting others. Accept `signal?: AbortSignal` in options; the CLI wrapper (sub-PR C) can create + wire a process-signal-driven AbortController. +**The gate is a journal shape, not a log line.** RFC rule 11c: *never woken* is the ABSENCE of any `wake_context_unresolved` entry alongside a dispatch carrying no wake context; *failed to resolve* is the PRESENCE of that entry naming its `reason`, with no dispatch for that attempt. Assert on those two shapes. `kernel/relayflowd/tests/event_wake.rs` already builds a woken run and reads entries back out of the journal — model the new test on it. -5. **Test coverage for pollError branch.** #83's tests never asserted the loop survives a fetcher throw AND the loop TERMINATES on a journal throw. Add both cases; without them, someone regresses `onPollError` to a no-op and every test still passes. +**Definition of done, all of it** -## Do not re-do these - -Merged and closed; a PR redoing any will be closed: - - picker actionability (#42), unterminated backticks (#45) - - deterministic-command preflight refusal (#47) — do not touch preflight - - gate-1 race regression test (#48) — do not touch `kernel/relayflowd/src/server/tests.rs` or `server.rs` - - ops/NEXT.md validation (#50) — do not touch `sdk/src/work-package-validator.ts` - - SDK agent worker (#53) — `sdk/src/worker.ts` is done; you MAY modify `close()` per finding #2 above, but do NOT rewrite the attach/dispatch/complete flow - - SDK pretest hook (#69) — `sdk/package.json` builds the kernel before `npm test`; don't touch - -## The task - -Add `sdk/src/hn-monitor-runner.ts`. It composes the existing pieces into a continuous runner: - - - constructs a `JournalClient` connected to the running `relayflowd` socket - - constructs an `AgentWorker` (from `sdk/src/worker.ts`) and calls `workerAttach()` for `agent` steps — attach BEFORE first poll (a run parked because no worker attached is only revived by `run.resume`; the live-kernel suite pins this) - - loops: `pollHackerNewsOnce(spec, sink)` → sleep `POLL_INTERVAL_MS` (env-configurable, default 60000 = 60s) → repeat - - exit cleanly on `AbortSignal.abort` (drain in-flight steps, close client, release worker per finding #2) - - exported from `sdk/src/index.ts` - -Keep it small and honest: - - the worker must attach BEFORE the first poll - - the poller layer handles single-fetch failures with a typed error; the loop just moves to the next tick — but journal errors MUST fail the runner (finding #1) - - no scheduling logic beyond the sleep (the kernel owns retry and dedupe policy) - - no LLM calls; the runner is glue, not a reviewer - -## Explicit non-goals for THIS PR (belongs to later sub-PRs) - - - Proving the workload actually executes end-to-end (dispatch → step complete). That is sub-PR B (integration test with real relayflowd + fake HN fetch + assert step reaches `done`). This PR ONLY proves the runner assembles and its unit tests hold. - - CLI wrapper (`flows hn-monitor start`). That is sub-PR C. - - Ops/STATE.md gate-2 GREEN declaration. That is sub-PR D. - -Say all three explicitly in the PR body so the history lens doesn't reject on "runner doesn't prove workload runs." - -## Definition of done, all of it - - - `sdk/src/hn-monitor-runner.ts` exists, exports `HnMonitorRunner` from `sdk/src/index.ts` - - `sdk/src/worker.ts` — either `close()` calls `workerRelease` (add to protocol.ts if missing), OR a one-line comment names what close() intentionally does NOT do - - `sdk/src/protocol.ts` — if you added `workerRelease`, matching request/response definitions - - `sdk/tests/hn-monitor-runner.test.ts` covers ALL of these: - - fake fetch + mock journal client → runner submits an event on each tick - - abort signal triggers clean shutdown within one tick (worker released or documented) - - worker attach happens before first poll - - **fetch throw → loop survives** (onPollError called, next tick still runs) - - **journal throw → loop TERMINATES** (runner.run() rejects with the error) - - `cd sdk && npm test` green (pretest hook builds the kernel automatically) - - EVERY new test confirmed to FAIL against current code (comment out the source; the test fails), with the literal failing output pasted in your summary - - PR body explicitly names the non-goals (test-actually-runs is sub-PR B; CLI is sub-PR C; gate-2 declaration is sub-PR D) - - ops/NEXT.md correctly says Gate 2, not Gate 3 (the assessor on #83 confused itself) + - `wake_context_unresolved` is a typed journal entry carrying `reason`, written by the kernel at the classification site rather than only by a test + - transient fails the attempt and retries under the ordinary budget; permanent parks `needs_human` and does not retry + - a new kernel test asserts both journal shapes from rule 11c + - the `resolve_wake_context` doc comment no longer describes the classification as absent — it currently does, at length, and a stale comment that contradicts the code is worse than none + - EVERY new test confirmed to FAIL against current code (revert the source change; watch it fail), with the literal failing output pasted in your summary + - `cd kernel && sh ../ops/cargo.sh test --workspace` must be green, with the output pasted - as your LAST action, run `git status --porcelain` and paste it -## Out of scope for THIS tick — DO NOT TOUCH +**Out of scope for THIS tick — DO NOT TOUCH** + - D2, `EpochSummaryPayload`, and engine-side epoch rollover — blocked upstream, and the RFC says adding the field ahead of its producer is "the appearance of a fix, not one" + - trigger-plane liveness — shipped in PR #122 (`kernel/relayflowd/src/server/liveness.rs`) + - the hn-monitor runner or CLI — shipped in PR #120 + - flipping the gate-2 verdict in `ops/STATE.md` — Khaliq's read, never a run's - `.github/workflows/*` — no GHA changes - - `kernel/*` — the kernel side of gate 2 already works via PR #14 - - `workflows/*.yaml` — those are for later sub-PRs - - `ops/AUTODRIVE_BRIEF.md` — chief owns this file, not the drive loop - - CLI wrapper — sub-PR C, separate PR - - end-to-end integration test with real relayflowd — sub-PR B, separate PR - - ops/STATE.md gate-2 declaration — sub-PR D, separate PR - -## If you cannot finish + - `packages/sdk/*` — this is a kernel task -Say so and file what you learned. A minimal runner with an honest gap description beats a complete-looking one that doesn't shut down cleanly or leaks journal errors. +**If you cannot finish.** Say so and file what you learned. A partial classification with an honest gap description beats a complete-looking one that cannot tell transient from permanent — conflating those two is the exact defect D1 exists to name. diff --git a/ops/NEEDS_HUMAN.md b/ops/NEEDS_HUMAN.md deleted file mode 100644 index 8bfae4c96..000000000 --- a/ops/NEEDS_HUMAN.md +++ /dev/null @@ -1,83 +0,0 @@ -# NEEDS_HUMAN — Conflicting Work Package Context - -**Situation:** This run has conflicting scope context that requires human clarification. - -## The Conflict - -1. **ops/TARGET.md says:** Gate 3, build hn-monitor runner (sub-PR A), `sdk/src/` code task -2. **ops/NEXT.md says:** Gate 3, cloud review-swarm preflight validation, `.github/workflows/` task -3. **These are completely different tasks** — one is SDK code (track A per TARGET), one is GitHub Actions (track D per NEXT) - -## Evidence - -**ops/TARGET.md line 1-5:** -``` -# TARGET — gate 3 - -This run is pinned to **gate 3** and must not work on any other gate. - -**Scope:** Build sub-PR A of the Gate 2 push: a real `hn-monitor` polling runner in the SDK. CODE task, `sdk/src/`-side. -``` - -**ops/NEXT.md line 1-3:** -``` -# NEXT — gate 3: complete cloud review-swarm preflight validation and documentation - -**Scope:** Track D: Cloud review-swarm redesign — build `.github/workflows/review-swarm.yml` correctly this time -``` - -## The Charter Says - -Per charter/LEAD.md (the instruction I received): -- "Read ops/TARGET.md if it exists" — it does, says hn-monitor -- "Then read ops/STATE.md, ops/DIRECTIVES.md" — done -- "Then write ops/NEXT.md: the SINGLE highest-priority work package toward the current gate" - -But ops/NEXT.md ALREADY EXISTS with different work. - -## Additional Context Found - -**ops/STATE.md gate 2 block (lines 39-81)** says: -- PR #120 merged 2026-09-01 — `flows hn-monitor start` CLI runner -- Gate 2 is AMBER, not GREEN -- Two clauses remain: trigger-plane liveness, analyze-agent execution - -**Actual file check:** -- `packages/sdk/src/cli/hn-monitor.ts` exists (288 lines) -- Contains `runHnMonitor` function implementing all TARGET.md requirements -- Addresses all five findings from closed PR #83 - -## The Question - -**Which work package should this run execute?** - -**Option A: ops/TARGET.md wins (hn-monitor, sdk/ territory)** -- Overwrite ops/NEXT.md with hn-monitor work package -- But the code already exists (PR #120 merged per STATE.md) -- Task would be verification/testing, not building from scratch - -**Option B: ops/NEXT.md wins (review-swarm, .github/ territory)** -- Ignore ops/TARGET.md -- Execute the review-swarm validation work -- But TARGET.md explicitly pins this run to gate 3 hn-monitor work - -**Option C: TARGET.md is stale** -- The launcher wrote an outdated TARGET.md referencing closed PR #83 -- Real work is in ops/NEXT.md (review-swarm) -- Proceed with review-swarm, update TARGET understanding - -**Option D: Both are stale** -- Neither accurately reflects current gate 3 needs -- Assessor should read RFC-0001 §3 gate 3 definition -- Write fresh work package from RFC requirements - -## Recommendation - -**Option C** — ops/TARGET.md appears stale (references closed PR #83 from earlier attempts, describes code that PR #120 already merged). The active work package is ops/NEXT.md (review-swarm). But I need human confirmation before overwriting NEXT.md or executing potentially wrong work. - -## What I Need - -**Clear answer:** Which work package is correct for this run? -- If hn-monitor: shall I overwrite the review-swarm NEXT.md, or is there a different file I should write? -- If review-swarm: shall I proceed with ops/NEXT.md as-is and ignore TARGET.md? -- If neither: what is the actual gate 3 work I should assess? diff --git a/ops/NEXT.md b/ops/NEXT.md index cb9d18473..b5d0bc64c 100644 --- a/ops/NEXT.md +++ b/ops/NEXT.md @@ -1,123 +1,112 @@ -# NEXT — gate 3 work package: document review-swarm secrets in README +# NEXT — gate 2: split wake-context resolution failure by cause (the rest of D1) -**Scope (from TARGET.md):** - -Track D: Cloud review-swarm redesign — build `.github/workflows/review-swarm.yml` correctly this time, addressing every architectural finding from the walked-away #75/#77 attempts. +**Scope.** The next gate is **gate 2**, by Khaliq's decision on 2026-09-16. +Not gate 3. If you are reading a target that says gate 3, it is stale — the +launcher used to synthesise "gate 3" over a gate-2 brief, which is the +contradiction that wedged the loop for four days. ## Objective -Complete the final missing piece of gate 3's Definition of Done: document `RELAY_WORKSPACE_KEY` and `CLOUD_API_KEY` secrets in README.md with instructions on how to obtain them. - -## Current state assessment - -All 9 architectural requirements from TARGET.md are SATISFIED in the existing code: +Close the remainder of **deviation D1** in +`docs/RFC-0001-everything-is-a-relayflow.md`. That document names it a +binding obligation, in these words: -1. ✅ Immutable gate — two checkout steps (`.github/workflows/review-swarm.yml:32-53`) -2. ✅ Unified verdict logic — `swarm-verdict.sh` sourced by both callers -3. ✅ Auth secret validation — preflight validates all three secrets (lines 141-188) -4. ✅ Sticky marker + transcripts — HTML anchors with upsert_comment -5. ✅ No author whitelist — verified absent -6. ✅ Cloud sandbox fetch on GHA runner — `swarm-prepare.sh` with GH_TOKEN -7. ✅ Timeout ordering — 60m < 65m < 75m with comments -8. ✅ Wait step records status — swarm_status output, always() post step -9. ✅ Transcript freshness — run-start marker with stale detection - -Verification commands all pass: ``` -bash -n .github/workflows/scripts/swarm-post.sh && \ -bash -n .github/workflows/scripts/swarm-prepare.sh && \ -bash -n .github/workflows/scripts/swarm-verdict.sh && \ -echo "All bash scripts parse OK" -# Output: All bash scripts parse OK - -python3 -c "import yaml; yaml.safe_load(open('.github/workflows/review-swarm.yml'))" && \ -python3 -c "import yaml; yaml.safe_load(open('workflows/review-swarm.yaml'))" && \ -echo "YAML files parse OK" -# Output: YAML files parse OK - -grep -i "whitelist\|github.event.pull_request.user.login" .github/workflows/review-swarm.yml || echo "No author whitelist found (GOOD)" -# Output: No author whitelist found (GOOD) - -grep -c "actions/checkout@v4" .github/workflows/review-swarm.yml -# Output: 2 +gate 2 cannot go green until D1 and D2 are closed ``` -**The gap:** TARGET.md Definition of Done item 6 requires: -> README.md — document `RELAY_WORKSPACE_KEY` secret + how to obtain - -Current reality: -``` -grep -c "RELAY_WORKSPACE_KEY\|CLOUD_API_KEY" README.md -# Output: 0 -``` - -README.md does NOT document these secrets. The workflow comment (`.github/workflows/review-swarm.yml:21-24`) references a runbook in the `AgentWorkforce/cloud` repo, but README has no such documentation. - -From `ops/NEEDS_HUMAN.md`, the secrets are stored and working (as of 2026-09-07), but gate 3 is blocked on Daytona CPU quota, not on implementation. The workflow WORKS; the documentation is missing. +D1 is half done. PR #252 stopped the resume path swallowing a journal scan +error into `wake_context: None`, so the error now propagates out of +`resolve_wake_context` in `kernel/relayflowd/src/engine/drive.rs`. That +function's own doc comment says what is still missing, in its own words: -## Files in scope +> That is deliberately narrower than the eventual contract. The intended +> end state classifies the failure and journals it (transient → retry the +> attempt under its budget; permanent → park `needs_human`), and none of +> that classification exists yet. -- `README.md` — add section documenting GitHub Actions secrets required for review-swarm +The classification still does not exist. `wake_context_unresolved` — the +journal entry RFC rule 10a requires, and the entry rule 11c makes the gate +assert on — appears nowhere in the kernel. Today an unreadable segment +aborts the dispatch before any attempt is recorded, and recovery only +arrives when the lease expires and a later drive abandons the attempt as +`Crashed`. Nothing in the journal says the wake context was the reason. -## Work package +## What RFC rule 10a requires -Add a "GitHub Actions Secrets" section to README.md documenting: +Resolution failure splits by cause, because transient I/O and permanent +corruption deserve opposite handling: -1. `RELAY_WORKSPACE_KEY` — Agent Relay workspace key for review swarm communication - - How to obtain: Contact repository administrator or see ops/NEEDS_HUMAN.md for historical context - - Why required: Enables agent coordination within review swarm workflow +- **Transient** (segment unreadable right now: I/O error, lock contention, a + truncated tail still being written) — fail the attempt, **retryable**, + under the step's ordinary retry budget. Journal `wake_context_unresolved` + with `reason: "transient"` and the underlying error. +- **Permanent** (the run is open on a wake, the current segment reads + cleanly, and no wake context is present) — fail the step and park as + **`needs_human`**, not retryable. Journal `wake_context_unresolved` with + `reason: "absent"`. -2. `CLOUD_API_KEY` — Agent Relay Cloud API credential for launching cloud workflows - - How to obtain: Minted per `AgentWorkforce/cloud → docs/runbooks/relay-ci-workflow-credential.md` - - Profile: `workflow-invoke` - - Scopes: `workflow:invoke:read` and `workflow:invoke:write` - - How to store: Repository Settings → Secrets and variables → Actions → New repository secret +Neither case may fall back to `wake_context: None`. Dispatching with no +context is reserved for runs that were never woken. -3. `CLOUD_API_URL` — Cloud API endpoint (typically `https://agentrelay.com/cloud`) - - Usually set as repository variable, not secret - - Defaults to production endpoint if not set +## Files in scope -The section should be brief (10-15 lines) and reference the workflow files for implementation details. +- `kernel/relayflowd-core/src/entry.rs` — add the `WakeContextUnresolved` + entry type and its `wake_context_unresolved` wire string, beside the + existing `SubscriptionMatched` / `subscription.matched` pair. +- `kernel/relayflowd/src/engine/drive.rs` — classify the failure at the + `resolve_wake_context` call site and journal it, instead of aborting the + dispatch before anything is recorded. Update the doc comment: it currently + describes the classification as absent, and that must stop being true. +- A new kernel integration test under kernel/relayflowd/tests/ asserting on + the two journal shapes rule 11c names. Model it on + `kernel/relayflowd/tests/event_wake.rs`, which already builds a woken run + and reads `SubscriptionMatched` back out of the journal. + +**The permanent case is reachable today.** A run that was woken, whose +`subscription.matched` entry carries no `wake_context` key, currently +resolves to `Ok(None)` and is indistinguishable from never-woken. That is +the shape to detect: open on a wake, clean read, no context. ## Definition of done -1. README.md contains a section documenting the three secrets/variables -2. Each entry states what it is and how to obtain it -3. Parse checks continue to pass: - ``` - bash -n .github/workflows/scripts/swarm-*.sh - python3 -c "import yaml; yaml.safe_load(open('.github/workflows/review-swarm.yml'))" - python3 -c "import yaml; yaml.safe_load(open('workflows/review-swarm.yaml'))" - ``` -4. Verification remains true: - ``` - grep -c "RELAY_WORKSPACE_KEY\|CLOUD_API_KEY" README.md - # Should return > 0 - grep -i "whitelist\|github.event.pull_request.user.login" .github/workflows/review-swarm.yml || echo "GOOD" - # Should return "GOOD" or nothing (no whitelist) +1. `wake_context_unresolved` exists as a typed journal entry with a `reason` + of `transient` or `absent`, and the kernel writes it at the classification + site — not only in a test. +2. A transient failure fails the attempt and stays retryable under the step's + ordinary budget. A permanent failure parks the run as `needs_human` and is + not retried. +3. A new test asserts on the two journal shapes from RFC rule 11c: *never + woken* is the ABSENCE of any `wake_context_unresolved` entry alongside a + dispatch with no wake context; *failed to resolve* is the PRESENCE of that + entry naming its reason, with no dispatch for that attempt. Assert on + journal entries, not on log text. +4. Every new test must be confirmed to FAIL against current code before the + fix — revert the source change, watch it fail, and paste the literal + failing output into your summary. A test that has never been observed + failing pins nothing. +5. `cargo test --workspace` must be green, run from the kernel directory: ``` -5. As final action: - ``` - git status --porcelain + cd kernel && sh ../ops/cargo.sh test --workspace ``` +6. As your last action, run `git status --porcelain` and paste it. ## Explicitly OUT of scope -- `.github/workflows/review-swarm.yml` (already correct, all 9 requirements satisfied) -- `workflows/review-swarm.yaml` (already correct) -- `.github/workflows/scripts/swarm-*.sh` (all already correct) -- `.gitignore` (no .review-target mask exists, already correct) -- `sdk/` (Track A owns that) -- `kernel/` (gate 1 done) -- `ops/*` (chief owns briefs and state) -- Any other GHA workflow -- Resolving the Daytona CPU quota block (that's in ops/NEEDS_HUMAN.md, different issue) -- Actually testing the workflow end-to-end (blocked on Daytona capacity per ops/NEEDS_HUMAN.md) - -## Why this is the work package - -TARGET.md's Definition of Done explicitly lists: -- Item 6: "PR body explicitly documents each of the 9 requirements above and shows where each is satisfied" -- Item 7: "`README.md` — document `RELAY_WORKSPACE_KEY` secret + how to obtain" - -The 9 requirements are satisfied in code. Item 7 is not satisfied. This is the remaining gap between current state and TARGET.md's done-when. +- **D2 and epoch rollover.** The RFC sequences this as engine-side epoch + rollover → D2 → gate 2 and states that D2 is not implementable in + isolation: every construction of `EpochSummaryPayload` in the kernel is + inside a test, so there is no site at which to add the carry-forward. + Adding the field first would be "the appearance of a fix, not one". Do not + start it here. +- **Trigger-plane liveness.** Already shipped in PR #122 — + `kernel/relayflowd/src/server/liveness.rs` with + `kernel/relayflowd/tests/subscription_liveness.rs`. Do not rebuild it. +- **The hn-monitor runner.** Already shipped in PR #120 — + `packages/sdk/src/cli/hn-monitor.ts`. Three separate escalations have now + proposed rebuilding it. Do not. +- **Flipping the gate-2 verdict.** That is Khaliq's read, not a run's. See + `ops/STATE.md`. +- **The analyze-agent clause.** Whether it is gate-2 or gate-4 scope is an + open question for Khaliq, recorded in `ops/STATE.md`. +- `.github/workflows/*`, `packages/sdk/*`, and `ops/*` other than this file. diff --git a/ops/STATE.md b/ops/STATE.md index 1fd6bb0e8..089ed84fa 100644 --- a/ops/STATE.md +++ b/ops/STATE.md @@ -8,7 +8,11 @@ it is authoritative when history is unavailable. **Keep it current. A stale STATE.md is worse than none:** it does not merely fail to help, it actively misleads an assessor that cannot check it. -Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mini), on `main`. **STATE.md gate-2 block rewritten; verdict unchanged (still AMBER).** A new evidence file — `ops/reviews/20260901-1050-gate2-live-run.md` — is cited from the gate-2 block; AMBER→GREEN is Khaliq's read on the enclosed evidence. +Last updated: 2026-09-16, on `main`, unwedging the drive loop. **Gate-2 clause 1 +corrected — it had been stale for two weeks (PR #122 closed it); deviations D1/D2 +recorded as clause 3; the 2026-09-16 escalation resolved and its content preserved +below. Verdict unchanged (still AMBER).** Prior entry: 2026-09-01 10:16 UTC, by the +Relayflow Lead (flows-lead-1 on sf-mini). **STATE.md gate-2 block rewritten; verdict unchanged (still AMBER).** A new evidence file — `ops/reviews/20260901-1050-gate2-live-run.md` — is cited from the gate-2 block; AMBER→GREEN is Khaliq's read on the enclosed evidence. ## Where the program is @@ -56,15 +60,19 @@ Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mi rather than restating specific numbers here (STATE.md counts drift, an evidence transcript does not). - **Why AMBER, not GREEN.** RFC-0001 §3 gate 2 has two clauses this - evidence does NOT close: - 1. **Trigger plane liveness-checked** (RFC-0001 §3 gate 2, the paragraph - ending "Native's silent-death problem"). RelayCron's deterministic-id - single-winner claim + `stale_after` sweep is the pattern. Not - implemented inside `relayflowd`. The poller runs; the kernel does not - yet notice if it stops. The same section calls this "a requirement, - not an option" — it is a stated done-when clause, not follow-up - hardening. + **Why AMBER, not GREEN.** RFC-0001 §3 gate 2 had two clauses this + evidence did NOT close. **Clause 1 has since closed; clause 2 has not, + and a third was always open.** + 1. **Trigger plane liveness-checked — CLOSED by PR #122 (`a774d880`).** + This block used to say "Not implemented inside `relayflowd`", and that + was true when it was written on 2026-09-01. It stopped being true when + #122 landed `kernel/relayflowd/src/server/liveness.rs` (the background + sweep: `detect_stale` + `latch_stale`, single-winner claim bucketed by + sweep id, `subscription.stale` journalled and emitted) wired end to end + by `kernel/relayflowd/tests/subscription_liveness.rs`. The stale entry + stood for two weeks and is exactly the failure this file warns about in + its own header: a stale STATE.md does not merely fail to help, it sends + an assessor to build something that already exists. 2. **The analyze-agent step actually executing.** In the recorded run, every step ended in `worker_error` because `hn-monitor start`'s AgentWorker has no user-supplied step handler. The dispatch loop @@ -72,6 +80,17 @@ Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mi ("runs succeed") or gate-4 scope ("chief-as-relayflow supplies the runtime") is Khaliq's call. + 3. **RFC-0001's own deviations D1 and D2** ("gate 2 cannot go green until + D1 and D2 are closed"). D1 is half closed: PR #252 (`97c886d2`) stopped + the resume path swallowing a journal scan error into + `wake_context: None`, but rule 10a's split by cause is not implemented — + `wake_context_unresolved` appears nowhere in the kernel, so transient + and permanent resolution failures are still handled identically. D2 is + blocked upstream: the RFC sequences it **engine-side epoch rollover → + D2 → gate 2** and says plainly that rollover, not D2, is the next + action. **The remainder of D1 is the current work package** — see + ops/NEXT.md. + **AMBER → GREEN is Khaliq's read** on the enclosed evidence, per this block's prior wording ("that is a judgement, not a missing part") and per the charter's standing rule that the Lead never merges / never @@ -80,6 +99,43 @@ Last updated: 2026-09-01 10:16 UTC, by the Relayflow Lead (flows-lead-1 on sf-mi which is a different thing). This session prepared evidence and left the flip pending. +## 2026-09-16 — the drive loop was wedged for four days, and why + +Recorded here because the escalation file that carried it has been deleted, +and this is the file an assessor with no git history reads instead. + +**Khaliq's decision: the next gate is GATE 2, not gate 3.** That settles the +question seven consecutive runs escalated about (PRs #417, #420, #422, #424, +#426, #427, #428 — every one of them a diff touching only ops/NEEDS_HUMAN.md +and ops/NEXT.md; none merged). It confirms what ops/AUTODRIVE_BRIEF.md +already said and what the gate ladder above already implied. + +**What the escalation actually said**, and what was true in it: an assessor +found ops/TARGET.md pinned to "gate 3" while describing gate-2 hn-monitor +work that PR #120 had already merged, and ops/NEXT.md describing unrelated +gate-3 review-swarm documentation. It could not tell which was real. It was +right to be confused — both sources were wrong, and it had no way to know. + +**Root cause, now fixed:** `ops/autodrive.sh` launched every run with +`ops/launch-gate.sh 3` while passing it the gate-2 brief, so the synthesised +target contradicted the brief inside it on every single tick. The assessors +were reporting a real defect in the launcher, once per run, for four days. + +**Two things that made it unrecoverable, both now fixed:** +- The escalation was sticky. `assess-gate` escalated on the mere EXISTENCE of + ops/NEEDS_HUMAN.md, which is committed on main and which nothing ever + deletes, so every future tick died at the gate on a question already + answered. The gate now requires two independent signals — the file exists + AND this tick wrote it. Pinned by ops/drive-assess-gate.test.mjs. +- The escalation cited evidence no human could inspect. ops/TARGET.md is + synthesised per run into a throwaway worktree by ops/launch-gate.sh and is + never in the delivered diff, so a reviewer sees quotations from a file that + does not exist. Review lenses failed PRs #422 and #428 for exactly this. + **Still open as a design question** — the drive.yaml assess task already + instructs the Lead to quote the target rather than cite its path, and that + instruction is evidently not enough. Whoever next touches the launcher + should consider committing the target into the delivered diff instead. + ## HANDOFF (2026-08-30 03:05 UTC) — a lead is LIVE on sf-mini **`flows-lead-1` is running on node `sf-mini`.** Attach with: diff --git a/ops/autodrive.sh b/ops/autodrive.sh index 776f04b35..345b752d1 100644 --- a/ops/autodrive.sh +++ b/ops/autodrive.sh @@ -112,7 +112,13 @@ while [ ! -f "$STOP_FILE" ]; do sleep "$INTERVAL" continue fi - out=$(sh ops/launch-gate.sh 3 "$brief" 2>&1) + # The gate number must match the brief. This read `3` while + # ops/AUTODRIVE_BRIEF.md described GATE 2 work, so ops/launch-gate.sh + # synthesised "TARGET — gate 3" over a gate-2 brief and every assessor + # that read both correctly reported a contradiction it could not resolve. + # That is the conflict seven consecutive runs escalated about. Khaliq's + # decision on 2026-09-16: the next gate is 2, not 3. + out=$(sh ops/launch-gate.sh 2 "$brief" 2>&1) rid=$(echo "$out" | sed -n 's/^Run created: //p' | head -1) if [ -n "$rid" ]; then say "launched ${rid%%-*}" diff --git a/ops/drive-assess-gate.test.mjs b/ops/drive-assess-gate.test.mjs new file mode 100644 index 000000000..c3cc83c8a --- /dev/null +++ b/ops/drive-assess-gate.test.mjs @@ -0,0 +1,174 @@ +// The assess-gate escalation test, executed rather than asserted about. +// +// ops/NEEDS_HUMAN.md was committed to main on 2026-09-06 (082c62aa) and +// nothing has ever deleted it. The gate escalated on the file's mere +// EXISTENCE, so from 2026-09-12 every drive tick exited 75 before doing any +// work: PRs #417, #420, #422, #424, #426, #427 and #428 are seven consecutive +// cloud runs whose entire diff is ops/NEEDS_HUMAN.md and ops/NEXT.md. +// +// The fix requires two independent signals to agree — the file exists AND +// this tick wrote it — so these tests drive BOTH directions. A gate that +// cannot fire is as broken as one that always fires, so the stale case and +// the live case are equally load-bearing here. +// +// The gate script is EXTRACTED from workflows/drive.yaml rather than copied, +// so this test fails if someone edits the workflow and not the test. No YAML +// parser is available at the repo root (there is no root package.json), and +// the block is a literal scalar at a known indent, so it is read as text. +import assert from 'node:assert/strict'; +import { execFileSync, spawnSync } from 'node:child_process'; +import { mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join, resolve } from 'node:path'; +import test from 'node:test'; + +const WORKFLOW = resolve(import.meta.dirname, '../workflows/drive.yaml'); + +/** Pull the `command:` literal block of a named deterministic step out of the workflow. */ +function gateScript(stepName) { + const lines = readFileSync(WORKFLOW, 'utf8').split('\n'); + const start = lines.findIndex((line) => line.trim() === `- name: ${stepName}`); + assert.notEqual(start, -1, `no step named ${stepName} in workflows/drive.yaml`); + const commandAt = lines.findIndex((line, i) => i > start && line.trim() === 'command: |'); + assert.notEqual(commandAt, -1, `step ${stepName} has no literal command block`); + const indent = lines[commandAt].search(/\S/) + 2; + const body = []; + for (const line of lines.slice(commandAt + 1)) { + if (line.trim() !== '' && line.search(/\S/) < indent) break; + body.push(line.slice(indent)); + } + const script = body.join('\n'); + assert.match(script, /NEEDS_HUMAN/, 'extracted the wrong block'); + return script; +} + +const git = (cwd, ...args) => + execFileSync('git', args, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] }); + +// A work package that satisfies the gate's later checks, so the only thing +// under test is the escalation branch. +const NEXT_MD = `# NEXT — a work package\n\n## Definition of done\n\n- \`cargo test --workspace\` must be green.\n`; + +/** + * Build a repo shaped like a drive tick: a `main` baseline, then a working + * branch that is "this tick". `staleEscalation` puts NEEDS_HUMAN.md on the + * baseline — the wedged state that burned seven runs. + */ +function tick(t, { staleEscalation = false } = {}) { + const root = mkdtempSync(join(tmpdir(), 'assess-gate-')); + t.after(() => rmSync(root, { recursive: true, force: true })); + git(root, 'init', '-q', '-b', 'main'); + git(root, 'config', 'user.email', 'test@relayflows.local'); + git(root, 'config', 'user.name', 'Gate Test'); + git(root, 'config', 'commit.gpgsign', 'false'); + execFileSync('mkdir', ['-p', join(root, 'ops')]); + writeFileSync(join(root, 'ops/NEXT.md'), NEXT_MD); + if (staleEscalation) { + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — answered long ago\n'); + } + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'base'); + git(root, 'checkout', '-q', '-b', 'flow/drive-tick'); + return root; +} + +function runGate(root) { + const result = spawnSync('sh', ['-c', gateScript('assess-gate')], { + cwd: root, + encoding: 'utf8', + }); + return { code: result.status, out: `${result.stdout}${result.stderr}` }; +} + +test('a stale escalation committed on main does not wedge the tick', (t) => { + const root = tick(t, { staleEscalation: true }); + // This tick writes its own work package, as assess does. + writeFileSync(join(root, 'ops/NEXT.md'), `${NEXT_MD}\nthis tick's package\n`); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'assess: work package for this tick'); + + const { code, out } = runGate(root); + assert.notEqual(code, 75, `the gate escalated on a stale file:\n${out}`); + assert.match(out, /ASSESS_STALE_NEEDS_HUMAN_IGNORED/); + assert.doesNotMatch(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); + assert.match(out, /ASSESS_GATE_PASS/); +}); + +test('an escalation committed by this tick still parks the run', (t) => { + const root = tick(t); + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — a live question\n\nWhich gate?\n'); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'assess: escalate'); + + const { code, out } = runGate(root); + assert.equal(code, 75, `a genuine escalation did not park the run:\n${out}`); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); + assert.match(out, /a live question/, 'the escalation body must reach the operator'); +}); + +test('an uncommitted escalation from this tick still parks the run', (t) => { + // Propagation between per-step sandboxes is lossy, so assess may write the + // file and fail to commit it. An escalation must not be lost that way. + const root = tick(t); + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — uncommitted but live\n'); + + const { code, out } = runGate(root); + assert.equal(code, 75, `an uncommitted live escalation was ignored:\n${out}`); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); +}); + +/** + * A tree carrying a live escalation that git cannot describe. `initGit: true` + * gives a real repo whose branch is not `main`; otherwise there is no repo at + * all. Both are shapes a sandbox actually produces — SYNC_MODE=snapshot runs + * `git init` over an extracted tarball, so `main` does not exist until sync + * creates it. + */ +function unprovableTick(t, { initGit = false } = {}) { + const root = mkdtempSync(join(tmpdir(), 'assess-gate-')); + t.after(() => rmSync(root, { recursive: true, force: true })); + execFileSync('mkdir', ['-p', join(root, 'ops')]); + writeFileSync(join(root, 'ops/NEXT.md'), NEXT_MD); + writeFileSync(join(root, 'ops/NEEDS_HUMAN.md'), '# NEEDS_HUMAN — live, on a tree git cannot describe\n'); + if (initGit) { + git(root, 'init', '-q', '-b', 'work'); + git(root, 'config', 'user.email', 'test@relayflows.local'); + git(root, 'config', 'user.name', 'Gate Test'); + git(root, 'config', 'commit.gpgsign', 'false'); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'base on a branch that is not main'); + } + return root; +} + +test('a live escalation parks the run when there is no git repo at all', (t) => { + const root = unprovableTick(t); + const { code, out } = runGate(root); + assert.equal(code, 75, `freshness was unprovable but the run continued:\n${out}`); + assert.match(out, /ASSESS_ESCALATION_FRESHNESS_UNPROVABLE/); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); + assert.doesNotMatch(out, /ASSESS_STALE_NEEDS_HUMAN_IGNORED/); +}); + +test('a live escalation parks the run when main does not exist', (t) => { + // The escalation is COMMITTED here, so a naive `git log main..HEAD` prints + // nothing and the file looks stale — the precise shape that would drop a + // live escalation in a snapshot sandbox. + const root = unprovableTick(t, { initGit: true }); + const { code, out } = runGate(root); + assert.equal(code, 75, `main was missing but the run continued:\n${out}`); + assert.match(out, /ASSESS_ESCALATION_FRESHNESS_UNPROVABLE/); + assert.match(out, /ASSESS_BLOCKED_NEEDS_HUMAN/); +}); + +test('a tick with no escalation at all passes the gate', (t) => { + const root = tick(t); + writeFileSync(join(root, 'ops/NEXT.md'), `${NEXT_MD}\nthis tick's package\n`); + git(root, 'add', '-A'); + git(root, 'commit', '-qm', 'assess: work package for this tick'); + + const { code, out } = runGate(root); + assert.equal(code, 0, `a clean tick did not pass:\n${out}`); + assert.match(out, /ASSESS_GATE_PASS/); + assert.doesNotMatch(out, /NEEDS_HUMAN/); +}); diff --git a/workflows/drive-cloud.yaml b/workflows/drive-cloud.yaml index 71e439720..5e6c034ee 100644 --- a/workflows/drive-cloud.yaml +++ b/workflows/drive-cloud.yaml @@ -46,8 +46,8 @@ workflows: \ that already has one\n# (laptop, fleet node). Both shapes are supported below; neither is\n# assumed,\ \ and an unmaterialized sandbox fails closed and typed\n# rather than failing later as an unexplained\ \ tool error.\nset -eu\necho \"SYNC_WORKDIR=$(pwd)\"\n\nmissing=\"\"\nfor required in AGENTS.md\ - \ docs/RFC-0001-everything-is-a-relayflow.md ops/DIRECTIVES.md kernel packages/sdk; do\n [ -e \"$required\"\ - \ ] || missing=\"$missing $required\"\ndone\nif [ -n \"$missing\" ]; then\n echo \"SYNC_FAIL_NOT_MATERIALIZED:\ + \ docs/RFC-0001-everything-is-a-relayflow.md ops/DIRECTIVES.md kernel packages/sdk; do\n [ -e \"\ + $required\" ] || missing=\"$missing $required\"\ndone\nif [ -n \"$missing\" ]; then\n echo \"SYNC_FAIL_NOT_MATERIALIZED:\ \ the repo is not present in this execution environment.\" >&2\n echo \" missing:$missing\" >&2\n\ \ echo \" cwd: $(pwd)\" >&2\n echo \" A cloud run must upload the working tree: \\`agent-relay\ \ cloud run\\`\" >&2\n echo \" syncs code by default; \\`--no-sync-code\\` produces exactly this\ @@ -101,14 +101,14 @@ workflows: \ the known\nsandbox faults that are NOT reasons to block. Then read\nops/DIRECTIVES.md \u2014 standing\ \ human directives outrank the backlog;\nif one is unsatisfied, it IS the work package.\nThen read\ \ docs/bootstrap-report.md and ops/DRIVE-LOG.md if they exist,\n`git log --oneline -15`, `gh pr\ - \ list --state open` and open PR review\nstate, kernel/ and packages/sdk/ test status. Then write ops/NEXT.md:\ - \ the\nSINGLE highest-priority work package toward the current gate\n(gate 1 until its done-when\ - \ in RFC-0001 \xA73 holds), with: objective,\nfiles in scope, definition of done (must include passing\ - \ commands),\nand what is explicitly OUT of scope for this tick. If an open PR is\nawaiting fixes\ - \ from review, the work package is fixing it \u2014 never\nstart new work over unfinished work.\ - \ If work is blocked on a human\ndecision, write ops/NEEDS_HUMAN.md stating the exact question and\ - \ the\noptions \u2014 and then STILL end with ASSESS_DONE.\n\nCOMMIT YOUR WORK PACKAGE BEFORE YOU\ - \ FINISH:\n git add -A && git commit -m \"assess: work package for this tick\"\nEach step runs\ + \ list --state open` and open PR review\nstate, kernel/ and packages/sdk/ test status. Then write\ + \ ops/NEXT.md: the\nSINGLE highest-priority work package toward the current gate\n(gate 1 until\ + \ its done-when in RFC-0001 \xA73 holds), with: objective,\nfiles in scope, definition of done (must\ + \ include passing commands),\nand what is explicitly OUT of scope for this tick. If an open PR is\n\ + awaiting fixes from review, the work package is fixing it \u2014 never\nstart new work over unfinished\ + \ work. If work is blocked on a human\ndecision, write ops/NEEDS_HUMAN.md stating the exact question\ + \ and the\noptions \u2014 and then STILL end with ASSESS_DONE.\n\nCOMMIT YOUR WORK PACKAGE BEFORE\ + \ YOU FINISH:\n git add -A && git commit -m \"assess: work package for this tick\"\nEach step runs\ \ in its OWN sandbox and files reach the next step only\nthrough the executor's propagation, which\ \ is lossy: on runs a2089144\nand 2560e02d your predecessor wrote ops/NEXT.md, said so truthfully,\n\ and the file never arrived \u2014 one of those runs finished with a\nzero-file patch. Committing\ @@ -129,24 +129,59 @@ workflows: - assess-1 command: "# A typed park, not a crash. The assess step cannot express \"blocked\"\n# in its final\ \ token (its gate only recognises ASSESS_DONE), so the\n# Lead writes ops/NEEDS_HUMAN.md instead\ - \ and this step reads it.\nset -u\nif [ -f ops/NEEDS_HUMAN.md ]; then\n echo \"ASSESS_BLOCKED_NEEDS_HUMAN:\ - \ the Lead escalated a decision it cannot make.\"\n echo \"--- ops/NEEDS_HUMAN.md ---\"\n cat\ - \ ops/NEEDS_HUMAN.md\n exit 75\nfi\nif [ ! -f ops/NEXT.md ]; then\n echo \"ASSESS_FAIL: no ops/NEXT.md\ - \ \u2014 an assessment that named no work package did not assess\"\n exit 1\nfi\n# The assessment\ - \ must have WRITTEN this tick's package, not merely\n# left the previous one in place. On run 457a6102\ - \ assess reported\n# \"The work package is written to ops/NEXT.md\" and the very next step\n# read\ - \ the OLD file \u2014 the logs carry the reason:\n# \"relayfile flush failed after the command\ - \ succeeded (exit 1);\n# a later agent step may see stale files\"\n# The builder then correctly\ - \ refused to invent scope, but only after\n# a whole build step had been spent. Catch it here instead:\ - \ if\n# ops/NEXT.md is identical to the base, the assessment did not land,\n# whoever is at fault.\n\ - # Look for the package in the working tree OR in a commit made this\n# tick. Propagation between\ - \ per-step sandboxes is lossy, so a package\n# that exists only as a loose file may not arrive;\ - \ one committed by\n# the assess step travels in git history instead.\nif git log --oneline main..HEAD\ - \ -- ops/NEXT.md 2>/dev/null | grep -q .; then\n echo \"ASSESS_PACKAGE_COMMITTED: found ops/NEXT.md\ - \ change in this tick's history\"\nelif git diff --quiet main -- ops/NEXT.md 2>/dev/null; then\n\ - \ # Warn, do not fail. This was fatal, and it killed four runs in six\n # while the loop produced\ - \ nothing \u2014 a worse outcome than the risk\n # it guarded against.\n #\n # The risk it guarded\ - \ was \"the builder gets scope nobody wrote this\n # tick\". But scope does not actually come from\ + \ and this step reads it.\nset -u\n# An escalation is trusted only when TWO INDEPENDENT SIGNALS\ + \ AGREE:\n# the file exists AND this tick is what wrote it.\n#\n# Existence alone is not a signal.\ + \ ops/NEEDS_HUMAN.md was committed to\n# main on 2026-09-06 (082c62aa) and nothing in this repo\ + \ has ever\n# deleted it \u2014 no `rm`, no `git rm`, and `git log --diff-filter=D`\n# over that\ + \ path is empty. ops/launch-gate.sh builds each run's\n# worktree from origin/main and does not\ + \ strip it, so every tick from\n# 2026-09-12 onward escalated here before doing any work, on a\n\ + # question a human had already answered. PRs #417, #420, #422, #424,\n# #426, #427 and #428 are\ + \ seven consecutive cloud runs whose entire\n# diff is this file and ops/NEXT.md, re-litigating\ + \ the same conflict.\n# None merged. A full cloud run was burned on each.\n#\n# A stale file must\ + \ never be able to masquerade as a live escalation.\n# \"This tick\" is the same test the ops/NEXT.md\ + \ freshness check below\n# uses \u2014 a commit in `main..HEAD` \u2014 widened by the uncommitted\ + \ case,\n# because propagation between per-step sandboxes is lossy and assess\n# may write the file\ + \ and fail to commit it. Losing a live escalation\n# is the worse error of the two, so an unproven-fresh\ + \ file that is\n# dirty in the working tree still parks the run.\n#\n# The default is to TRUST the\ + \ escalation. Only a positive, SUCCESSFUL\n# answer from git may downgrade it to stale, because\ + \ \"git printed\n# nothing\" and \"git could not answer\" look identical otherwise \u2014 and\n\ + # a sandbox is exactly where git cannot answer. SYNC_MODE=snapshot\n# runs `git init` over an extracted\ + \ tarball, so a step that runs\n# before main exists, a missing .git, or any git failure would\n\ + # silently classify a LIVE escalation as stale and walk the builder\n# straight past a human decision.\ + \ That inverts the tradeoff above,\n# so an unprovable escalation parks the run.\nif [ -f ops/NEEDS_HUMAN.md\ + \ ]; then\n escalation=unprovable\n if git rev-parse --git-dir >/dev/null 2>&1 \\\n && git\ + \ rev-parse --verify --quiet main >/dev/null 2>&1; then\n tick_log=$(git log --oneline main..HEAD\ + \ -- ops/NEEDS_HUMAN.md 2>/dev/null)\n if [ $? -ne 0 ]; then\n escalation=unprovable\n \ + \ elif [ -n \"$tick_log\" ]; then\n escalation=this_tick_committed\n else\n tick_dirty=$(git\ + \ status --porcelain -- ops/NEEDS_HUMAN.md 2>/dev/null)\n if [ $? -ne 0 ]; then\n escalation=unprovable\n\ + \ elif [ -n \"$tick_dirty\" ]; then\n escalation=this_tick_uncommitted\n else\n\ + \ escalation=stale\n fi\n fi\n fi\n if [ \"$escalation\" = stale ]; then\n echo\ + \ \"ASSESS_STALE_NEEDS_HUMAN_IGNORED: ops/NEEDS_HUMAN.md exists but this tick did not write it.\"\ + \n echo \" It is the committed record of an escalation that has already been answered,\"\n \ + \ echo \" not a live one, so it does not park this run. Delete it from main once its\"\n echo\ + \ \" question is resolved \u2014 a resolved escalation left in the tree is a lie the\"\n echo\ + \ \" next assessor has to spend a run disproving.\"\n else\n if [ \"$escalation\" = unprovable\ + \ ]; then\n echo \"ASSESS_ESCALATION_FRESHNESS_UNPROVABLE: git could not say whether this tick\"\ + \n echo \" wrote ops/NEEDS_HUMAN.md (no repo, no main, or git failed). Failing safe and\"\n\ + \ echo \" treating it as live: ignoring a real escalation is the worse of the two errors.\"\ + \n fi\n echo \"ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make ($escalation).\"\ + \n echo \"--- ops/NEEDS_HUMAN.md ---\"\n cat ops/NEEDS_HUMAN.md\n exit 75\n fi\nfi\nif\ + \ [ ! -f ops/NEXT.md ]; then\n echo \"ASSESS_FAIL: no ops/NEXT.md \u2014 an assessment that named\ + \ no work package did not assess\"\n exit 1\nfi\n# The assessment must have WRITTEN this tick's\ + \ package, not merely\n# left the previous one in place. On run 457a6102 assess reported\n# \"The\ + \ work package is written to ops/NEXT.md\" and the very next step\n# read the OLD file \u2014 the\ + \ logs carry the reason:\n# \"relayfile flush failed after the command succeeded (exit 1);\n#\ + \ a later agent step may see stale files\"\n# The builder then correctly refused to invent scope,\ + \ but only after\n# a whole build step had been spent. Catch it here instead: if\n# ops/NEXT.md\ + \ is identical to the base, the assessment did not land,\n# whoever is at fault.\n# Look for the\ + \ package in the working tree OR in a commit made this\n# tick. Propagation between per-step sandboxes\ + \ is lossy, so a package\n# that exists only as a loose file may not arrive; one committed by\n\ + # the assess step travels in git history instead.\nif git log --oneline main..HEAD -- ops/NEXT.md\ + \ 2>/dev/null | grep -q .; then\n echo \"ASSESS_PACKAGE_COMMITTED: found ops/NEXT.md change in\ + \ this tick's history\"\nelif git diff --quiet main -- ops/NEXT.md 2>/dev/null; then\n # Warn,\ + \ do not fail. This was fatal, and it killed four runs in six\n # while the loop produced nothing\ + \ \u2014 a worse outcome than the risk\n # it guarded against.\n #\n # The risk it guarded was\ + \ \"the builder gets scope nobody wrote this\n # tick\". But scope does not actually come from\ \ ops/NEXT.md: it comes\n # from ops/TARGET.md, which the launcher COMMITS into the uploaded\n\ \ # tree, so it is present in every per-step sandbox and cannot be\n # lost to the propagation\ \ fault. NEXT.md refines the target; it does\n # not define it.\n echo \"ASSESS_WARN_STALE_NEXT:\ @@ -221,66 +256,67 @@ workflows: \ test --workspace 2>&1); rc=$?\n echo \"$out\" | tail -8; ran=1\n if [ $rc -eq 124 ]; then\n\ \ echo \"VERIFY_SUITE_TIMEOUT: the kernel suite exceeded ${VERIFY_SUITE_TIMEOUT:-900}s and was\ \ killed.\"\n echo \" A hanging test is a defect, not a pass \u2014 recording it as a failure.\"\ - \n fi\n [ $rc -eq 0 ] || ok=1\nfi\nif [ -f packages/sdk/package.json ]; then\n # node_modules is not in\ - \ `git ls-files`, so a sandbox has none.\n # Install before testing, and treat a failed install\ - \ as a failed\n # verify rather than letting `npm test` report a confusing error.\n if [ ! -d\ - \ packages/sdk/node_modules ]; then\n echo \"VERIFY_INSTALL: packages/sdk/node_modules absent \u2014 installing\"\ - \n out=$(cd packages/sdk && run_bounded \"npm ci\" npm ci 2>&1); rc=$?\n if [ $rc -ne 0 ]; then\n \ - \ echo \"$out\" | tail -12\n echo \"VERIFY_FAIL: npm ci failed \u2014 cannot test what\ - \ did not install\"\n exit 1\n fi\n fi\n # Restore exec bits. `npm ci` reported success\ - \ (96 packages) and\n # esbuild still failed with EACCES on run 909e18f6, so the mount\n # this\ - \ installs onto does not carry the executable bit. That is\n # the same fault that left ops/cargo.sh\ - \ non-executable. Which\n # layer drops it is NOT established; this repairs the symptom\n # where\ - \ it is observed and is a no-op where the bit survives.\n chmod -R +x packages/sdk/node_modules/.bin 2>/dev/null\ - \ || true\n find packages/sdk/node_modules -type d -name bin -path \"*esbuild*\" \\\n -exec chmod -R\ - \ +x {} + 2>/dev/null || true\n out=$(cd packages/sdk && run_bounded \"sdk suite\" npm test 2>&1); rc=$?\n\ - \ echo \"$out\" | tail -8; ran=1\n if [ $rc -eq 124 ]; then\n echo \"VERIFY_SUITE_TIMEOUT:\ - \ the sdk suite exceeded ${VERIFY_SUITE_TIMEOUT:-900}s and was killed.\"\n fi\n [ $rc -eq 0 ]\ - \ || ok=1\nfi\nif [ \"$ran\" -eq 0 ]; then\n echo \"VERIFY_FAIL: no suite found \u2014 a verify\ - \ that runs nothing cannot pass\"; exit 1\nfi\n# Enforce the NEXT.md contract instead of asking\ - \ for it.\n#\n# The assess prompt has told runs since PR #19 to quote their scope\n# rather than\ - \ cite ops/TARGET.md, because that file lives only in the\n# throwaway launch worktree and is not\ - \ in the delivered diff. Runs\n# kept citing it anyway \u2014 reviewers filed the same finding again\ - \ on\n# #35, #40 and #48. Four recurrences after the warning is enough\n# evidence that prose does\ - \ not hold and a check does.\n#\n# Same for unevidenced test claims: \"all three tests pass\" with\ - \ no\n# transcript. validateNextWorkPackage refuses both shapes.\n#\n# Runs BEFORE the node_modules\ - \ cleanup below, which needs packages/sdk/dist.\nif [ -f ops/NEXT.md ] && [ -f packages/sdk/dist/index.js ]; then\n\ - \ nextout=$(node -e '\n const fs = require(\"node:fs\");\n const { validateNextWorkPackage\ - \ } = require(\"./packages/sdk/dist/index.js\");\n const verdict = validateNextWorkPackage(\n fs.readFileSync(\"\ - ops/NEXT.md\", \"utf8\"),\n (p) => fs.existsSync(p),\n );\n if (!verdict.accepted) {\n\ - \ console.log(\"NEXT_REFUSED \" + verdict.reason);\n process.exit(1);\n }\n console.log(\"\ - NEXT_OK\");\n ' 2>&1) || true\n echo \"$nextout\" | tail -2\n case \"$nextout\" in\n *NEXT_REFUSED*)\n\ - \ echo \"VERIFY_FAIL: ops/NEXT.md was refused \u2014 quote your scope instead of citing\"\n\ - \ echo \" ops/TARGET.md (it is not in the delivered diff), and paste the literal\"\n \ - \ echo \" command output for anything you claim passes.\"\n ok=1\n ;;\n esac\nfi\n\n\ - # Shrink the tree before the flush that matters.\n#\n# The relayfile mount flush fails on a large\ - \ tree \u2014 http 413, or a\n# timeout waiting for the daemon to ack \u2014 and the failure is\n\ - # NON-FATAL, so the run reports success while later steps read stale\n# files and the delivered\ - \ patch silently loses the work. Five runs\n# were lost this way. PR #38 took kernel/target out\ - \ (4914 paths ->\n# 0), which was not enough on its own: run 5ecf7078 still flushed\n# 1932 files\ - \ and still hit 413.\n#\n# Dependency and compiler outputs are already gitignored, and nothing\n\ - # downstream of here reads them: the suites and NEXT validation have\n# finished by this line, while\ - \ commit/handoff touch tracked files.\n# Removing them after the gates run costs nothing and keeps\ - \ them out\n# of every flush from here on.\n#\n# Runs on pass AND fail. The first version only slimmed\ - \ on a green\n# verify, to keep a failing run's tree intact for diagnosis. That\n# defeated the\ - \ fix: run ae982aaa ended VERIFY_FAIL_NONFATAL, so the\n# cleanup was skipped, the tree stayed at\ - \ 1842 files and the flush\n# failed with three 413s \u2014 and the run still committed and delivered\n\ - # a PR. A run that delivers needs its flush to work whether or not\n# its gates passed, and node_modules\ - \ is reinstallable, so there was\n# never anything to preserve for diagnosis here.\n#\n# Guarded\ - \ with || true and placed AFTER the pass/fail decision is\n# computed, so a cleanup problem can\ - \ never change the verdict.\nrm -rf packages/sdk/node_modules packages/sdk/dist packages/surface/dist 2>/dev/null || true\n\ - echo \"VERIFY_TREE_SLIMMED: removed packages/sdk/node_modules and generated dist trees before flush (ok=$ok)\"\ - \n\n# Measure the tree instead of guessing at it.\n#\n# The relayfile flush keeps failing (http\ - \ 413) and three successive\n# hypotheses about WHY were wrong: kernel/target was removed and the\n\ - # count barely moved; node_modules was removed and it still failed;\n# HOME turned out to be /home/daytona,\ - \ outside the tree, so the\n# toolchain was never in it. Meanwhile the bootstrap reports ~1900\n\ - # changed files against a repo of only 257 tracked ones, and the\n# delivered patch is 12 files.\ - \ Nothing in the log says what the rest\n# are, so this prints it rather than inviting a fourth\ - \ guess.\necho \"TREE_CENSUS: top directories by file count\"\nfind . -type f -not -path './.git/*'\ - \ 2>/dev/null \\\n | sed -E 's|^\\./||; s|/[^/]*$||' \\\n | cut -d/ -f1-2 \\\n | sort | uniq\ - \ -c | sort -rn | head -12 || true\necho \"TREE_CENSUS_TOTAL: $(find . -type f -not -path './.git/*'\ - \ 2>/dev/null | wc -l | tr -d ' ')\"\n[ \"$ok\" -eq 0 ] && echo VERIFY_PASS || echo \"VERIFY_FAIL_NONFATAL:\ - \ recorded; the next cycle must address it\"\n" + \n fi\n [ $rc -eq 0 ] || ok=1\nfi\nif [ -f packages/sdk/package.json ]; then\n # node_modules\ + \ is not in `git ls-files`, so a sandbox has none.\n # Install before testing, and treat a failed\ + \ install as a failed\n # verify rather than letting `npm test` report a confusing error.\n if\ + \ [ ! -d packages/sdk/node_modules ]; then\n echo \"VERIFY_INSTALL: packages/sdk/node_modules\ + \ absent \u2014 installing\"\n out=$(cd packages/sdk && run_bounded \"npm ci\" npm ci 2>&1);\ + \ rc=$?\n if [ $rc -ne 0 ]; then\n echo \"$out\" | tail -12\n echo \"VERIFY_FAIL: npm\ + \ ci failed \u2014 cannot test what did not install\"\n exit 1\n fi\n fi\n # Restore exec\ + \ bits. `npm ci` reported success (96 packages) and\n # esbuild still failed with EACCES on run\ + \ 909e18f6, so the mount\n # this installs onto does not carry the executable bit. That is\n #\ + \ the same fault that left ops/cargo.sh non-executable. Which\n # layer drops it is NOT established;\ + \ this repairs the symptom\n # where it is observed and is a no-op where the bit survives.\n chmod\ + \ -R +x packages/sdk/node_modules/.bin 2>/dev/null || true\n find packages/sdk/node_modules -type\ + \ d -name bin -path \"*esbuild*\" \\\n -exec chmod -R +x {} + 2>/dev/null || true\n out=$(cd\ + \ packages/sdk && run_bounded \"sdk suite\" npm test 2>&1); rc=$?\n echo \"$out\" | tail -8; ran=1\n\ + \ if [ $rc -eq 124 ]; then\n echo \"VERIFY_SUITE_TIMEOUT: the sdk suite exceeded ${VERIFY_SUITE_TIMEOUT:-900}s\ + \ and was killed.\"\n fi\n [ $rc -eq 0 ] || ok=1\nfi\nif [ \"$ran\" -eq 0 ]; then\n echo \"VERIFY_FAIL:\ + \ no suite found \u2014 a verify that runs nothing cannot pass\"; exit 1\nfi\n# Enforce the NEXT.md\ + \ contract instead of asking for it.\n#\n# The assess prompt has told runs since PR #19 to quote\ + \ their scope\n# rather than cite ops/TARGET.md, because that file lives only in the\n# throwaway\ + \ launch worktree and is not in the delivered diff. Runs\n# kept citing it anyway \u2014 reviewers\ + \ filed the same finding again on\n# #35, #40 and #48. Four recurrences after the warning is enough\n\ + # evidence that prose does not hold and a check does.\n#\n# Same for unevidenced test claims: \"\ + all three tests pass\" with no\n# transcript. validateNextWorkPackage refuses both shapes.\n#\n\ + # Runs BEFORE the node_modules cleanup below, which needs packages/sdk/dist.\nif [ -f ops/NEXT.md\ + \ ] && [ -f packages/sdk/dist/index.js ]; then\n nextout=$(node -e '\n const fs = require(\"\ + node:fs\");\n const { validateNextWorkPackage } = require(\"./packages/sdk/dist/index.js\");\n\ + \ const verdict = validateNextWorkPackage(\n fs.readFileSync(\"ops/NEXT.md\", \"utf8\"),\n\ + \ (p) => fs.existsSync(p),\n );\n if (!verdict.accepted) {\n console.log(\"NEXT_REFUSED\ + \ \" + verdict.reason);\n process.exit(1);\n }\n console.log(\"NEXT_OK\");\n ' 2>&1)\ + \ || true\n echo \"$nextout\" | tail -2\n case \"$nextout\" in\n *NEXT_REFUSED*)\n echo\ + \ \"VERIFY_FAIL: ops/NEXT.md was refused \u2014 quote your scope instead of citing\"\n echo\ + \ \" ops/TARGET.md (it is not in the delivered diff), and paste the literal\"\n echo \" command\ + \ output for anything you claim passes.\"\n ok=1\n ;;\n esac\nfi\n\n# Shrink the tree\ + \ before the flush that matters.\n#\n# The relayfile mount flush fails on a large tree \u2014 http\ + \ 413, or a\n# timeout waiting for the daemon to ack \u2014 and the failure is\n# NON-FATAL, so\ + \ the run reports success while later steps read stale\n# files and the delivered patch silently\ + \ loses the work. Five runs\n# were lost this way. PR #38 took kernel/target out (4914 paths ->\n\ + # 0), which was not enough on its own: run 5ecf7078 still flushed\n# 1932 files and still hit 413.\n\ + #\n# Dependency and compiler outputs are already gitignored, and nothing\n# downstream of here reads\ + \ them: the suites and NEXT validation have\n# finished by this line, while commit/handoff touch\ + \ tracked files.\n# Removing them after the gates run costs nothing and keeps them out\n# of every\ + \ flush from here on.\n#\n# Runs on pass AND fail. The first version only slimmed on a green\n#\ + \ verify, to keep a failing run's tree intact for diagnosis. That\n# defeated the fix: run ae982aaa\ + \ ended VERIFY_FAIL_NONFATAL, so the\n# cleanup was skipped, the tree stayed at 1842 files and the\ + \ flush\n# failed with three 413s \u2014 and the run still committed and delivered\n# a PR. A run\ + \ that delivers needs its flush to work whether or not\n# its gates passed, and node_modules is\ + \ reinstallable, so there was\n# never anything to preserve for diagnosis here.\n#\n# Guarded with\ + \ || true and placed AFTER the pass/fail decision is\n# computed, so a cleanup problem can never\ + \ change the verdict.\nrm -rf packages/sdk/node_modules packages/sdk/dist packages/surface/dist\ + \ 2>/dev/null || true\necho \"VERIFY_TREE_SLIMMED: removed packages/sdk/node_modules and generated\ + \ dist trees before flush (ok=$ok)\"\n\n# Measure the tree instead of guessing at it.\n#\n# The\ + \ relayfile flush keeps failing (http 413) and three successive\n# hypotheses about WHY were wrong:\ + \ kernel/target was removed and the\n# count barely moved; node_modules was removed and it still\ + \ failed;\n# HOME turned out to be /home/daytona, outside the tree, so the\n# toolchain was never\ + \ in it. Meanwhile the bootstrap reports ~1900\n# changed files against a repo of only 257 tracked\ + \ ones, and the\n# delivered patch is 12 files. Nothing in the log says what the rest\n# are, so\ + \ this prints it rather than inviting a fourth guess.\necho \"TREE_CENSUS: top directories by file\ + \ count\"\nfind . -type f -not -path './.git/*' 2>/dev/null \\\n | sed -E 's|^\\./||; s|/[^/]*$||'\ + \ \\\n | cut -d/ -f1-2 \\\n | sort | uniq -c | sort -rn | head -12 || true\necho \"TREE_CENSUS_TOTAL:\ + \ $(find . -type f -not -path './.git/*' 2>/dev/null | wc -l | tr -d ' ')\"\n[ \"$ok\" -eq 0 ] &&\ + \ echo VERIFY_PASS || echo \"VERIFY_FAIL_NONFATAL: recorded; the next cycle must address it\"\n" timeoutMs: 1200000 - name: commit-1 type: deterministic diff --git a/workflows/drive.yaml b/workflows/drive.yaml index a52be0aa6..ccf2b7996 100644 --- a/workflows/drive.yaml +++ b/workflows/drive.yaml @@ -191,11 +191,74 @@ workflows: # in its final token (its gate only recognises ASSESS_DONE), so the # Lead writes ops/NEEDS_HUMAN.md instead and this step reads it. set -u + # An escalation is trusted only when TWO INDEPENDENT SIGNALS AGREE: + # the file exists AND this tick is what wrote it. + # + # Existence alone is not a signal. ops/NEEDS_HUMAN.md was committed to + # main on 2026-09-06 (082c62aa) and nothing in this repo has ever + # deleted it — no `rm`, no `git rm`, and `git log --diff-filter=D` + # over that path is empty. ops/launch-gate.sh builds each run's + # worktree from origin/main and does not strip it, so every tick from + # 2026-09-12 onward escalated here before doing any work, on a + # question a human had already answered. PRs #417, #420, #422, #424, + # #426, #427 and #428 are seven consecutive cloud runs whose entire + # diff is this file and ops/NEXT.md, re-litigating the same conflict. + # None merged. A full cloud run was burned on each. + # + # A stale file must never be able to masquerade as a live escalation. + # "This tick" is the same test the ops/NEXT.md freshness check below + # uses — a commit in `main..HEAD` — widened by the uncommitted case, + # because propagation between per-step sandboxes is lossy and assess + # may write the file and fail to commit it. Losing a live escalation + # is the worse error of the two, so an unproven-fresh file that is + # dirty in the working tree still parks the run. + # + # The default is to TRUST the escalation. Only a positive, SUCCESSFUL + # answer from git may downgrade it to stale, because "git printed + # nothing" and "git could not answer" look identical otherwise — and + # a sandbox is exactly where git cannot answer. SYNC_MODE=snapshot + # runs `git init` over an extracted tarball, so a step that runs + # before main exists, a missing .git, or any git failure would + # silently classify a LIVE escalation as stale and walk the builder + # straight past a human decision. That inverts the tradeoff above, + # so an unprovable escalation parks the run. if [ -f ops/NEEDS_HUMAN.md ]; then - echo "ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make." - echo "--- ops/NEEDS_HUMAN.md ---" - cat ops/NEEDS_HUMAN.md - exit 75 + escalation=unprovable + if git rev-parse --git-dir >/dev/null 2>&1 \ + && git rev-parse --verify --quiet main >/dev/null 2>&1; then + tick_log=$(git log --oneline main..HEAD -- ops/NEEDS_HUMAN.md 2>/dev/null) + if [ $? -ne 0 ]; then + escalation=unprovable + elif [ -n "$tick_log" ]; then + escalation=this_tick_committed + else + tick_dirty=$(git status --porcelain -- ops/NEEDS_HUMAN.md 2>/dev/null) + if [ $? -ne 0 ]; then + escalation=unprovable + elif [ -n "$tick_dirty" ]; then + escalation=this_tick_uncommitted + else + escalation=stale + fi + fi + fi + if [ "$escalation" = stale ]; then + echo "ASSESS_STALE_NEEDS_HUMAN_IGNORED: ops/NEEDS_HUMAN.md exists but this tick did not write it." + echo " It is the committed record of an escalation that has already been answered," + echo " not a live one, so it does not park this run. Delete it from main once its" + echo " question is resolved — a resolved escalation left in the tree is a lie the" + echo " next assessor has to spend a run disproving." + else + if [ "$escalation" = unprovable ]; then + echo "ASSESS_ESCALATION_FRESHNESS_UNPROVABLE: git could not say whether this tick" + echo " wrote ops/NEEDS_HUMAN.md (no repo, no main, or git failed). Failing safe and" + echo " treating it as live: ignoring a real escalation is the worse of the two errors." + fi + echo "ASSESS_BLOCKED_NEEDS_HUMAN: the Lead escalated a decision it cannot make ($escalation)." + echo "--- ops/NEEDS_HUMAN.md ---" + cat ops/NEEDS_HUMAN.md + exit 75 + fi fi if [ ! -f ops/NEXT.md ]; then echo "ASSESS_FAIL: no ops/NEXT.md — an assessment that named no work package did not assess"