From f8b7ce534f7ec558c0f4e92140b8d74d179da0c3 Mon Sep 17 00:00:00 2001 From: kjgbot Date: Sun, 6 Sep 2026 16:24:35 +0200 Subject: [PATCH] docs(next): point drive runs at #174 instead of human-blocked credential work MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The previous package named the review-swarm credential. That work is real and it is blocked on a repository administrator: minting a Cloud credential and storing an Actions secret are not agent-permitted, and the Lead may not edit the gate that judges its work. Four consecutive drive runs read it, correctly concluded they were blocked, and each produced a NEEDS_HUMAN saying so — #199, #202, #207, #208. That is four cycles spent re-deriving one fact. A package that names human-blocked work turns every run into a report. #174 is the opposite: a real intermittent hang in crash-resume, reopened today with fresh evidence, needing no credential and no gate access. It reproduces at roughly one run in eight on main, which makes it tractable by repetition rather than by insight. The package carries the evidence a run needs and the trap that made this look like a regression: the failure rate did not change when seven commits landed in ten minutes, the sample size did. A shell-only commit failed while the next passed with identical kernel code. Definition of done requires proving a fix by repetition and explicitly permits stopping if it cannot be reproduced, because a hang nobody reproduced is not fixed by a change nobody can test. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01FtQSAcGDta5VH9xiZFT4sR Session-Id: c228933d-4f94-4d83-9a9a-daf3c83b94f1 --- ops/NEXT.md | 127 ++++++++++++++++++++++++++-------------------------- 1 file changed, 63 insertions(+), 64 deletions(-) diff --git a/ops/NEXT.md b/ops/NEXT.md index ca355999d..a75b36db5 100644 --- a/ops/NEXT.md +++ b/ops/NEXT.md @@ -1,85 +1,84 @@ -# NEXT — give the review gate a credential +# NEXT — fix the crash-resume hang (#174) -**Scope:** one Actions secret and two `env:` lines in -`.github/workflows/review-swarm.yml`. Nothing else. +**Scope:** `kernel/relayflowd/`, the crash-resume test suite, and nothing else. -**The Relayflow Lead cannot do this one.** RFC-0001 decision #6 and the -charter's second hard rail: it cannot edit the gates that judge its work. +## Why this and not gate 3 -## The headline +The previous package pointed at the review-swarm credential. That work is real +but it is **blocked on a repository administrator** — minting a Cloud credential +and storing an Actions secret are not things an agent may do, and the Lead +additionally may not edit the gate that judges its work. -**The review swarm has never succeeded.** +Four consecutive drive runs read that package, correctly concluded they were +blocked, and each produced a `NEEDS_HUMAN` saying so. That is four cycles spent +re-deriving the same fact. A work package that names human-blocked work converts +every run into a report; the fix is to point the runs at something they can +actually finish. -``` -TOTAL runs: 76 failure: 75 cancelled: 1 successes: 0 -first 2026-08-30T20:22:22Z -latest 2026-09-06T04:03:35Z -``` - -Treat any claim that gate 3 is "architecturally complete" against that number. -Most of its nine requirements describe behaviour downstream of a launch that has -never happened, so nothing past authentication has ever executed. +The credential decision is tracked and waiting elsewhere. Do not work on it here. -## What already shipped (2026-09-06) +## The problem -Four layers, each revealing the next: +`llm::sigkill_sweep_covers_before_and_between_the_rung_b_steps` hangs +intermittently on GitHub runners. Issue **#174**, reopened 2026-09-06 with fresh +evidence after being closed. -| step | failed because | closed by | -|---|---|---| -| `Validate cloud authentication` | repo had zero Actions secrets | `RELAY_WORKSPACE_KEY` added | -| `Prepare review input` | gate scripts were mode `100644`, exit 126 | #172 | -| `Launch cloud swarm` | CLI never installed, exit 127 | #198 | -| `Launch cloud swarm` | pinned runtime read no API key | #198 (pin → 11.10.3) | +``` +thread 'llm::sigkill_sweep_covers_before_and_between_the_rung_b_steps' +panicked at relayflowd/tests/crash_resume/llm.rs:121:27 +test result: FAILED. 33 passed; 1 failed +``` -Also landed: #203 (whole-line verdict matching, `jq -er` on the poll response), -#202 (a missing reviews directory yields `MISSING` rather than a `find` error). +Line 121 is the `no step.dispatch after resume` path — the worker never receives +a dispatch after the daemon is SIGKILLed and resumed. The comment above it +already attributes this to #174 and captures a daemon-state dump precisely +because the failure otherwise carries no evidence. -## The one thing left +## The evidence, and what makes it tractable now -The job now has a CLI that can read an API key, and no key to read. -`agent-relay cloud run` falls back to the interactive device flow and dies after -ten minutes: +It reproduces at roughly one run in eight on `main`: ``` -Device login expired before it was approved. Run the command again to get a new code. +main, cloud-runtime-artifact.yml, last 8 runs: 7 success, 1 failure ``` -`@agent-relay/cloud@11.10.3` resolves `CLOUD_API_KEY` through -`WorkflowApiKeyClient.fromEnv`, which `workflowApiClient` prefers over the stored -login. With the variable set, the device flow is never reached. +Earlier this looked like a regression from a specific commit, because `main` +normally runs about once a day and seven commits landed within ten minutes. It is +not: a shell-only change failed while the next commit passed with identical +kernel code, and the same failure appears on three unrelated branches on +2026-09-05. **The rate did not change; the sample size did.** + +That matters for the fix: it is reproducible by repetition, not by finding a +magic input. Run the crash-resume suite in a loop and it will show up. ## What to do -1. **Mint the credential.** `AgentWorkforce/cloud` → - `docs/runbooks/relay-ci-workflow-credential.md`, profile - `CI_TOKEN_PROFILE=workflow-invoke`. Non-human, workspace-bound, scoped to - exactly `workflow:invoke:read` and `workflow:invoke:write`. The runbook notes - provisioning and rotation "require no browser login". -2. **Store it.** An operator mints; **a repository administrator stores it**. The - runbook is explicit that an agent is not authorized to create or update - GitHub secrets. -3. **Set both variables** on the `Launch cloud swarm` step: `CLOUD_API_URL` and - `CLOUD_API_KEY`. -4. **Fix the preflight, which currently cannot fail.** `Validate cloud - authentication` tests that `RELAY_WORKSPACE_KEY` is non-empty, never examines - the credential `cloud run` uses, and never attempts an authentication — it - passed green on run 34007204726, whose authentication then failed ten minutes - later. Assert both variables, the way `AgentWorkforce/relay` does: - - ```bash - test -n "$CLOUD_API_URL" - test -n "$CLOUD_API_KEY" - ``` - -**Precedent:** `AgentWorkforce/relay`'s `.github/workflows/relayflow-pr-proof.yml` -runs this exact shape in production — published CLI, `CLOUD_API_URL` and -`CLOUD_API_KEY` in the environment, no interactive login. +1. Reproduce it locally. `cd kernel && sh ../ops/cargo.sh test -p relayflowd --test crash_resume` + in a loop until it fails. Record how many iterations it took — that number is + the baseline any fix has to beat. +2. Find where the dispatch is lost. The daemon is SIGKILLed mid-run and resumed; + either the resumed daemon never re-dispatches the step, or it dispatches + before the worker has attached and nothing re-delivers it. +3. Fix it in `kernel/relayflowd/`. Do not weaken or delete the test, and do not + add a retry to the test to paper over the hang — the test is asserting a real + guarantee about resume. +4. Prove the fix by repetition, not by one green run. State the iteration count + before and after. ## Definition of done -1. A review-swarm run reaches a step after `Launch cloud swarm` — the first - non-zero success in this workflow's history. -2. Paste the literal step list showing `Launch cloud swarm` succeeded. -3. If it fails, paste the literal error and STOP. Do not weaken the gate to make - it green. A gate that passes without running is the failure this whole - sequence has been climbing out of. +1. `cargo test --workspace` green from `kernel/`. +2. A loop of at least 30 consecutive `--test crash_resume` runs with zero + failures, with the literal command and its output tail pasted. +3. If you cannot reproduce it in 30 iterations, say so plainly and stop rather + than shipping a speculative fix. A hang nobody reproduced is not fixed by a + change nobody can test. + +## Constraints + +- `kernel/` only. Do not touch `.github/workflows/`, `packages/`, or the + publish pipeline. +- Do not edit `testdata/tick-heartbeat.*` or `hello-ladder.*` — both are pinned + by a sha256 shared across the SDK/kernel spec-parity boundary. +- `ops/reviews/`, `ops/DRIVE-LOG.md` and `ops/BACKLOG.md` are records of what was + true when written. Do not rewrite them.