From 954e104592a422b108c1679601e9e6b84a7193b2 Mon Sep 17 00:00:00 2001 From: Relayflow Lead Date: Mon, 14 Sep 2026 16:20:51 -0700 Subject: [PATCH 1/3] =?UTF-8?q?ops:=20verify-features=20drive=20=E2=80=94?= =?UTF-8?q?=20Grok=20loop=20variant,=20directive,=20ordered=20plan?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Sets up an unattended drive loop that works toward "every feature verifiable, then autonomous merge" on a Grok credit pool: - ops/VERIFY-FEATURES-PLAN.md: measured starting state, the four verify tiers, and seven ordered work packages (WP-V1 expect(kind) … WP-V7 auto-merge) each with a runnable definition of done and blocked-by. - ops/DIRECTIVES.md: the standing directive the assess step honors first. - ops/gen-drive-cloud.py --cli/--out: emits same-steps variants of the cloud loop; workflows/drive-cloud-grok.yaml is generated, not hand-edited (verified equal to drive-cloud.yaml apart from cli/name/channel). - ops/launch-gate.sh DRIVE_WORKFLOW and ops/autodrive.sh AUTODRIVE_BRIEF_FILE: env overrides so this loop can coexist with the default one. ops/AUTODRIVE_BRIEF-VERIFY.md is its brief. node --test ops/*.test.mjs: 64 pass, 0 fail. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_0166HKhhZVRRp81PTPeLw43E --- ops/AUTODRIVE_BRIEF-VERIFY.md | 5 + ops/DIRECTIVES.md | 9 + ops/VERIFY-FEATURES-PLAN.md | 179 +++++++++++++++++ ops/autodrive.sh | 10 +- ops/gen-drive-cloud.py | 31 ++- ops/launch-gate.sh | 12 +- workflows/drive-cloud-grok.yaml | 332 ++++++++++++++++++++++++++++++++ 7 files changed, 567 insertions(+), 11 deletions(-) create mode 100644 ops/AUTODRIVE_BRIEF-VERIFY.md create mode 100644 ops/VERIFY-FEATURES-PLAN.md create mode 100644 workflows/drive-cloud-grok.yaml diff --git a/ops/AUTODRIVE_BRIEF-VERIFY.md b/ops/AUTODRIVE_BRIEF-VERIFY.md new file mode 100644 index 000000000..93bdcb05c --- /dev/null +++ b/ops/AUTODRIVE_BRIEF-VERIFY.md @@ -0,0 +1,5 @@ +Work the standing directive of 2026-09-14 in ops/DIRECTIVES.md: make every feature in this repo verifiable and get to autonomous merge. The ordered work list with runnable definitions of done is ops/VERIFY-FEATURES-PLAN.md (WP-V1 … WP-V7). + +Selection rule for this tick: read ops/STATE.md, ops/DRIVE-LOG.md and `gh pr list --state open` if available; take the FIRST work package whose `blocked-by` packages are all merged on main. If an earlier package's PR is open and review is waiting on fixes, that fix IS the tick. One package per tick, never two. Quote the package's definition of done into ops/NEXT.md verbatim — cite no path that is not in the tree. + +Hard rails, unchanged: never edit `.github/workflows/review-swarm*.yml`, `workflows/review-swarm.yaml`, `ops/preswarm-check/`, or anything in ops/IMMUTABLE_PATHS; never merge; if a package cannot land as written, write ops/NEEDS_HUMAN.md with the exact question and still end with ASSESS_DONE. Every claim of a passing command carries the literal command and its captured output. diff --git a/ops/DIRECTIVES.md b/ops/DIRECTIVES.md index f71f0d73d..c8ce50218 100644 --- a/ops/DIRECTIVES.md +++ b/ops/DIRECTIVES.md @@ -3,3 +3,12 @@ Directives from Khaliq to the Relayflow Lead. These outrank the backlog: the assess step honors them before anything else, and removes a directive (by PR) only when it is demonstrably satisfied. + +## 2026-09-14 — every feature verifiable, then autonomous merge + +Work `ops/VERIFY-FEATURES-PLAN.md` in order: take the first work package +(WP-V1 … WP-V7) whose `blocked-by` are all merged on `main`, or fix an open +PR from an earlier package if review is waiting on it. One package per tick. +Quote the package's definition of done into `ops/NEXT.md`. Satisfied when +WP-V7 is merged and the plan's human steps are recorded done; remove this +directive in that PR. diff --git a/ops/VERIFY-FEATURES-PLAN.md b/ops/VERIFY-FEATURES-PLAN.md new file mode 100644 index 000000000..2d19bb4b4 --- /dev/null +++ b/ops/VERIFY-FEATURES-PLAN.md @@ -0,0 +1,179 @@ +# VERIFY-FEATURES — every feature in flows verifiable, then autonomous merge + +Directive from Khaliq, 2026-09-14. Goal state: + +> A PR merges without a human when unit tests, integration tests, and **flow +> tests** are all green at its head. A flow test is a `.flow.ts` that uses an +> agent to exercise a feature end to end, with deterministic gates judging the +> result. Every feature and piece of functionality in this repo is mapped to +> the tests that prove it, so a break is known before it lands — the same shape +> as `relay`'s `verify-features` (manifest + audit + tiered procedures + +> `checks.jsonl` / `verdict.json`). + +This file is the ordered work list. Each drive tick takes the **first work +package whose `blocked-by` are all merged**, or fixes an open PR from an +earlier package if review is waiting. Never two packages in one tick. + +## Where we started (2026-09-14, measured) + +- `main` has no branch protection and no rulesets; no check is required. + PR #383 merged having run only `review` + `guard` — every test workflow is + path-filtered, so a diff outside `kernel/`, `packages/sdk/**`, + `packages/surface/**`, `regressions/**`, `testdata/**` runs zero tests. +- Unit: SDK 101 test files / 120 src; kernel 240 `#[test]`s incl. + `crash_resume`; **surface 5 test files / 70 src**; ts-plugin 1; + `create-flow` and `relayflows` wrappers 0. +- Integration: `live-kernel`, `daemon-lifecycle-live`, `yaml-local-agent-live`, + `webhook-live`, `real-cli-adapters` exist, but CI sets + `RELAYFLOWS_ALLOW_ANALYZER_SKIP=1` and never sets + `RELAYFLOWS_REAL_CLI_ADAPTERS`. CI has never proved an agent step against a + real model. +- Flow tests: `regressions/` is dormant (typecheck only); `examples/` gallery + 1/3 passing by hand; nothing in `workflows/` runs in CI. Zero flows run in CI. +- No feature manifest, no surface audit, no change→feature mapping, no + coverage measurement (no vitest coverage provider, no `cargo llvm-cov`). +- Model access in CI: GitHub runners have no `claude`/`codex`/`grok`. The + review swarm gets a model by placing itself in Agent Relay Cloud + (`agent-relay cloud run workflows/review-swarm.yaml --sync-code`). + +## Tiers (pin these words; SKIP is never PASS) + +| Tier | Name | Environment | Proves | +| --- | --- | --- | --- | +| 0 | unit | vitest / cargo, no I/O | one module's contract | +| 1 | integration | real `relayflowd` + deterministic wrapper CLIs from `testdata/preflight/`, no model | components compose | +| 2 | flow test | `flows run --local-agent` against the real kernel + a real agent CLI + pinned model; gates on journal facts | the feature works for a user | +| 3 | cloud flow test | tier 2 placed via `agent-relay cloud run` | cloud-only features (deploy/bucket, `run --cloud`, webhook admission) | + +Every check records `pass` / `fail` / `skip` + reason in `checks.jsonl`; +`verdict.json` is the only authoritative result. An agent's prose never +greens a test — only a deterministic gate does. + +## Work packages, in order + +Each has a runnable definition of done. Quote the DoD into `ops/NEXT.md`; +paste literal command output in the PR (AGENTS.md "Evidence is captured"). + +### WP-V1 — surface `expect(kind)`: declared-failure assertion +- **Why first:** `regressions/README.md` gap #1. Negative flow tests ("this + refuses with `model_unavailable`") cannot be written without it, and every + refusal path in the manifest needs one. +- **Scope:** `packages/surface/src/step.ts` (+ types), lowering in + `packages/sdk/src/named-gate-lowering.ts` / `named-gates.ts`, + `docs/SURFACE.md` §2 entry. A step may declare + `.expect({ kind: '' })`; the run passes iff the step fails + with exactly that declared kind and fails otherwise. +- **DoD:** `cd packages/surface && bun run test` and `cd packages/sdk && npm test` + pass with new tests covering: matching kind passes; wrong kind fails; step + that succeeds fails the expectation. One regression pair in `regressions/` + rewritten to use it, typechecking under `bun run typecheck:regressions`. +- **blocked-by:** none. + +### WP-V2 — feature manifest v1 + surface audit +- **Scope:** `.agentworkforce/features/manifest.yaml`, + `.agentworkforce/features/verify/procedures.md`, + `scripts/audit-feature-manifest.mjs`, `scripts/audit-feature-manifest.test.mjs`. +- **Manifest entry shape:** `id`, `name`, `category`, `criticality` + (`critical|hot|standard`), `location` (globs), `verify_tier` (0–3), and + `verify: { unit: [...], integration: [...], flow: [...] }` listing test + files / flow ids, or `unverified: `. +- **Categories to enumerate (derive, do not guess):** CLI commands from the + `USAGE` block in `packages/sdk/src/cli.ts` (`add build deploy run[yaml|ts|digest|cloud|local-agent|reuse] check[watch|json] serve-webhook tick resume replay observer hn-monitor`); + surface verbs from `packages/surface/src/context.ts` and `step.ts` + (`f.run`, three `f.llm` forms, `f.agent`, named + predicate `.gate`, + `.expect`, `f.done`, `f.cloud.*`, `f.memory.*`, slack); YAML dialect + every + helper in `packages/surface/src/helpers/index.ts` + `triggers/`; kernel + behaviours (journal, resume, leases, durable timers, streams, hello ladder + a/b/c, budget); adapters (`claude`, `codex`, `wrapper`, headless per + `packages/sdk/src/adapters/`); plugins/`flows add`; webhook receiver + (auth, rate limit, provider-shape, admission); bundle/digest; MCP; observer + link; daemon lifecycle. +- **Audit:** derives the surface from `flows --help`, the exported + `FlowContext` type, `helpers/index.ts`, `triggers/`, and diffs against the + manifest. Exit `0` clean, `1` drift, `2` audit could not run (never read as + clean). Fails on any feature with no `verify` entry and no `unverified`. +- **DoD:** `node scripts/audit-feature-manifest.mjs` exits 0 on the branch; + `node --test scripts/audit-feature-manifest.test.mjs` passes and includes a + case where a command is added to USAGE and the audit reports drift. +- **blocked-by:** WP-V1 (so `.expect` is in the surface the audit derives). + +### WP-V3 — flow-test harness + first critical flows +- **Scope:** `flow-tests/.flow.ts` (v2 dialect), `scripts/flow-tests.mjs`, + `scripts/flow-tests.test.mjs`, `flow-tests/README.md`. +- **Runner:** reads the manifest; `--feature ...` or `--all`; runs each + flow via `flows run --local-agent --json --data-dir `; writes + `.workflow-artifacts/flow-tests/checks.jsonl` and `verdict.json` + (`pass|fail|skip` + reason per feature; overall = no `fail`, and no `skip` + on a `critical` feature). Zero retries. Per-flow budget cap and timeout. +- **First flows (all `critical`):** hello ladder a (deterministic), b (llm + + gate), c (agent); kill -9 mid-run then `flows resume` completes only + unfinished work with exact budget; `flows check` refuses a missing / + unauthenticated CLI (uses WP-V1 `.expect`); `serve-webhook` rejects a bad + signature and admits a good one; budget cap parks the run with + `completionReason` set. +- **DoD:** `node scripts/flow-tests.mjs --all` on an authenticated host + writes a `verdict.json` with every listed feature `pass`; the runner's own + tests pass; `regressions/` pairs run under the same runner with red/green + semantics (red passes while the bug is open, green fails). +- **blocked-by:** WP-V1, WP-V2. + +### WP-V4 — CI runs every tier on every PR +- **Scope:** `.github/workflows/tests.yml` (new umbrella: kernel, SDK, + surface, ts-plugin, schema, `create-flow`/`relayflows` smoke — no path + filters), `.github/workflows/flow-tests.yml` (places `workflows/flow-tests.yaml` + in Agent Relay Cloud exactly like `review-swarm.yml`, waits, posts + `verdict.json` to the PR, fails the job on any `fail` or critical `skip`), + and in that placed job `RELAYFLOWS_REAL_CLI_ADAPTERS=1` with + `RELAYFLOWS_ALLOW_ANALYZER_SKIP` unset so the live integration tests gate. +- **Selection:** changed files → manifest `location` globs → feature ids → + their flows; `critical` always; `--all` on `push: main` and in the merge + queue. The PR gets one sticky comment: "touches X, Y; flow tests A, B: PASS". +- **DoD:** on the PR that adds it, both workflows run and are green; the + posted comment names the features the PR touched. +- **blocked-by:** WP-V3. +- **Human step (Khaliq, cannot be done from a sandbox):** ruleset on `main` + requiring `tests`, `flow-tests`, `linux-x64-artifact`, `packed-consumer`, + `review`, `guard`, `npm` at the head sha; enable merge queue. Write + `ops/NEEDS_HUMAN.md` naming the exact check names when WP-V4 merges. + +### WP-V5 — coverage ratchet +- **Scope:** `@vitest/coverage-v8` in `packages/sdk` and `packages/surface`, + `cargo llvm-cov` for `kernel/`, `scripts/coverage-ratchet.mjs` comparing + the PR's numbers to `main`'s committed baseline + (`.agentworkforce/coverage-baseline.json`); a drop fails, a rise updates + the baseline in the same PR. +- **DoD:** the ratchet runs in `tests.yml`; a deliberate test deletion on a + scratch branch is shown failing it (paste output). +- **blocked-by:** WP-V4. + +### WP-V6 — fill the thin layers +- Surface: tests for every helper client in `packages/surface/src/helpers/` + (request shape, error mapping), `cloud.ts`, `memory.ts`, `runtime.ts`. +- `create-flow` and `relayflows` wrappers: scaffold-and-run smoke tests. +- ts-plugin: diagnostics for each rule it ships. +- Manifest `unverified:` count must reach zero for `critical` and `hot`. +- **DoD:** audit exits 0 with no `unverified` on critical/hot; coverage + ratchet rises for surface. +- **blocked-by:** WP-V5. May be split across ticks by package. + +### WP-V7 — autonomous merge +- **Scope:** `.github/workflows/auto-merge.yml`: on `check_suite`/`status` + completion for an open PR, if every required check is green **at the head + sha**, the swarm verdict is PASSED, `verdict.json` has no `fail` and no + critical `skip`, and the diff touches nothing under `ops/IMMUTABLE_PATHS` + or `.github/workflows/**`, then `gh pr merge --squash --auto`. +- **DoD:** the workflow's decision function is a script with unit tests + covering each refusal; a dry-run mode logs the decision without merging. +- **blocked-by:** WP-V4, and a **human step**: Khaliq amends RFC-0001 settled + decision #16 to name this mechanical policy (decision #6 forbids the Lead + from writing its own authority). Ship the workflow behind a + `AUTO_MERGE_ENABLED` repository variable defaulting to off until then. + +## What this loop may not do + +- Merge anything. `ops/autodrive.sh` delivers PRs; a human (or, once WP-V7 is + enabled by the RFC amendment, the workflow) merges. +- Edit `.github/workflows/review-swarm*.yml`, `workflows/review-swarm.yaml`, + or `ops/preswarm-check/` — the gates that judge this work (decision #6). +- Widen a package. If a package cannot land as written, say so in + `ops/NEEDS_HUMAN.md` and stop. diff --git a/ops/autodrive.sh b/ops/autodrive.sh index 776f04b35..287a88ffa 100644 --- a/ops/autodrive.sh +++ b/ops/autodrive.sh @@ -28,11 +28,15 @@ INTERVAL="${AUTODRIVE_INTERVAL:-300}" MAX_LIVE="${AUTODRIVE_MAX_LIVE:-1}" STOP_FILE="${AUTODRIVE_STOP:-/tmp/autodrive.stop}" STATE="${AUTODRIVE_STATE:-/tmp/autodrive-seen.txt}" +# The brief is a file so a running loop can be re-aimed without a restart; +# BRIEF_FILE lets two loops with different briefs (and DRIVE_WORKFLOW +# variants) coexist without fighting over ops/AUTODRIVE_BRIEF.md. +BRIEF_FILE="${AUTODRIVE_BRIEF_FILE:-ops/AUTODRIVE_BRIEF.md}" touch "$STATE" say() { echo "[$(date -u +%H:%M:%S)] $*"; } -say "autodrive starting (interval ${INTERVAL}s, max ${MAX_LIVE} live run, stop: $STOP_FILE)" +say "autodrive starting (interval ${INTERVAL}s, max ${MAX_LIVE} live run, brief: $BRIEF_FILE, workflow: ${DRIVE_WORKFLOW:-workflows/drive-cloud.yaml}, stop: $STOP_FILE)" while [ ! -f "$STOP_FILE" ]; do cd "$REPO" 2>/dev/null || { say "FATAL: $REPO missing"; exit 1; } @@ -106,9 +110,9 @@ while [ ! -f "$STOP_FILE" ]; do # loop. Five generic-brief cycles ("read STATE.md, pick one small thing") # produced nothing deliverable, while every run that produced real code had # a specific, scoped task. Vague instructions cost a full cycle each. - brief=$(cat ops/AUTODRIVE_BRIEF.md 2>/dev/null) + brief=$(cat "$BRIEF_FILE" 2>/dev/null) if [ -z "$brief" ]; then - say "NO BRIEF: ops/AUTODRIVE_BRIEF.md is missing or empty — not launching blind" + say "NO BRIEF: $BRIEF_FILE is missing or empty — not launching blind" sleep "$INTERVAL" continue fi diff --git a/ops/gen-drive-cloud.py b/ops/gen-drive-cloud.py index a0e3b6ed7..105d5a92b 100644 --- a/ops/gen-drive-cloud.py +++ b/ops/gen-drive-cloud.py @@ -7,7 +7,15 @@ laptop closed. The operator recovers the work with `agent-relay cloud sync`. Run from the repo root: python3 ops/gen-drive-cloud.py + +Variant with every agent on one CLI (e.g. to spend a provider credit pool): + + python3 ops/gen-drive-cloud.py --cli grok --out workflows/drive-cloud-grok.yaml + +The step bodies are identical; only the agents' `cli` and the swarm name / +channel differ, so the variant is a generated file too, never hand-edited. """ +import argparse import copy import yaml @@ -39,12 +47,15 @@ BASE_STEPS = ["assess", "assess-gate", "build", "verify"] -def build(): +def build(cli=None, suffix=""): d = yaml.safe_load(open("workflows/drive.yaml")) src = {s["name"]: s for s in d["workflows"][0]["steps"]} out = copy.deepcopy(d) - out["name"] = "flows-drive-cloud" + out["name"] = f"flows-drive-cloud{suffix}" + if cli: + for agent in out["agents"]: + agent["cli"] = cli out["description"] = ( "The Lead's tick, shaped for a cloud sandbox with the laptop closed.\n" "A cloud sandbox has no git remote and no GitHub token, so this flow\n" @@ -53,7 +64,7 @@ def build(): "with `agent-relay cloud sync `. Nothing reaches main without a\n" "human. GENERATED from workflows/drive.yaml by ops/gen-drive-cloud.py.\n" ) - out["swarm"]["channel"] = "flows-drive-cloud" + out["swarm"]["channel"] = f"flows-drive-cloud{suffix}" out["swarm"]["timeoutMs"] = 3600000 # 1h — one cycle takes ~10 min; a run # that has not finished in an hour is hung, not slow, and should stop # burning budget rather than sit for eight hours. @@ -131,9 +142,15 @@ def build(): if __name__ == "__main__": - doc = build() - with open("workflows/drive-cloud.yaml", "w") as f: + ap = argparse.ArgumentParser() + ap.add_argument("--cli", help="put every agent on this CLI (claude, codex, grok, ...)") + ap.add_argument("--out", default="workflows/drive-cloud.yaml") + args = ap.parse_args() + suffix = f"-{args.cli}" if args.cli else "" + doc = build(cli=args.cli, suffix=suffix) + regen = "python3 ops/gen-drive-cloud.py" + (f" --cli {args.cli} --out {args.out}" if args.cli else "") + with open(args.out, "w") as f: f.write("# GENERATED from workflows/drive.yaml by ops/gen-drive-cloud.py.\n" - "# Do not hand-edit: change drive.yaml, then regenerate.\n") + f"# Do not hand-edit: change drive.yaml, then regenerate: {regen}\n") yaml.safe_dump(doc, f, sort_keys=False, width=100, default_flow_style=False) - print(f"wrote workflows/drive-cloud.yaml ({len(doc['workflows'][0]['steps'])} steps)") + print(f"wrote {args.out} ({len(doc['workflows'][0]['steps'])} steps)") diff --git a/ops/launch-gate.sh b/ops/launch-gate.sh index 4dec2e13e..22906429b 100755 --- a/ops/launch-gate.sh +++ b/ops/launch-gate.sh @@ -79,5 +79,15 @@ fi echo "LAUNCH_GATE=$gate" echo "LAUNCH_WORKTREE=$work" echo "LAUNCH_TARGET_TRACKED=ok" -agent-relay cloud run workflows/drive-cloud.yaml 2>&1 | grep -E "Run created|Status:" +# DRIVE_WORKFLOW selects the generated cloud variant to launch. The default +# is the canonical claude/codex loop; ops/gen-drive-cloud.py --cli emits +# same-steps variants (e.g. workflows/drive-cloud-grok.yaml) for spending a +# specific provider's credit pool without forking the loop itself. +workflow="${DRIVE_WORKFLOW:-workflows/drive-cloud.yaml}" +if [ ! -f "$workflow" ]; then + echo "LAUNCH_FAIL_NO_WORKFLOW: $workflow is not in the launch worktree (is it on origin/main?)" >&2 + exit 70 +fi +echo "LAUNCH_WORKFLOW=$workflow" +agent-relay cloud run "$workflow" 2>&1 | grep -E "Run created|Status:" echo "LAUNCH_NOTE: worktree kept at $work — remove with 'git worktree remove --force $work'" diff --git a/workflows/drive-cloud-grok.yaml b/workflows/drive-cloud-grok.yaml new file mode 100644 index 000000000..a91c642ec --- /dev/null +++ b/workflows/drive-cloud-grok.yaml @@ -0,0 +1,332 @@ +# GENERATED from workflows/drive.yaml by ops/gen-drive-cloud.py. +# Do not hand-edit: change drive.yaml, then regenerate: python3 ops/gen-drive-cloud.py --cli grok --out workflows/drive-cloud-grok.yaml +version: '1.0' +name: flows-drive-cloud-grok +description: 'The Lead''s tick, shaped for a cloud sandbox with the laptop closed. + + A cloud sandbox has no git remote and no GitHub token, so this flow + + never delivers: it runs 1 full work-package cycles back to back + + in ONE sandbox, committing each to the sandbox branch. Recover the work + + with `agent-relay cloud sync `. Nothing reaches main without a + + human. GENERATED from workflows/drive.yaml by ops/gen-drive-cloud.py. + + ' +swarm: + pattern: dag + channel: flows-drive-cloud-grok + timeoutMs: 3600000 +agents: +- name: lead + cli: grok + preset: analyst + role: The Relayflow Lead. Assesses state, plans one work package, reports honestly. +- name: builder + cli: grok + preset: worker + role: Implements the work package. Rust for kernel/, TypeScript for packages/sdk/. +- name: adversary + cli: grok + preset: reviewer + role: Adversarial reviewer against RFC-0001 and AGENTS.md. +workflows: +- name: drive-cloud-loop + steps: + - name: sync + type: deterministic + command: "# Materialize the repo; never assume it. This step assumed a clone\n# with an `origin` remote\ + \ and so every cloud tick died here with\n# `fatal: 'origin' does not appear to be a git repository`\ + \ (runs\n# 9fc8d996, ff35187a, 06505b94, 4cf36ea7, b33c2c9a).\n#\n# A cloud workflow sandbox does\ + \ NOT get a clone. The platform's own\n# materialization is the code sync: the CLI tars the `git\ + \ ls-files`\n# set and the bootstrap extracts it into the code mount, then runs\n# `git init` over\ + \ it. Files yes, `.git` history and remotes no.\n# A checkout with a remote only exists on a host\ + \ that already has one\n# (laptop, fleet node). Both shapes are supported below; neither is\n# assumed,\ + \ and an unmaterialized sandbox fails closed and typed\n# rather than failing later as an unexplained\ + \ tool error.\nset -eu\necho \"SYNC_WORKDIR=$(pwd)\"\n\nmissing=\"\"\nfor required in AGENTS.md\ + \ docs/RFC-0001-everything-is-a-relayflow.md ops/DIRECTIVES.md kernel packages/sdk; do\n [ -e \"\ + $required\" ] || missing=\"$missing $required\"\ndone\nif [ -n \"$missing\" ]; then\n echo \"SYNC_FAIL_NOT_MATERIALIZED:\ + \ the repo is not present in this execution environment.\" >&2\n echo \" missing:$missing\" >&2\n\ + \ echo \" cwd: $(pwd)\" >&2\n echo \" A cloud run must upload the working tree: \\`agent-relay\ + \ cloud run\\`\" >&2\n echo \" syncs code by default; \\`--no-sync-code\\` produces exactly this\ + \ state.\" >&2\n exit 78\nfi\necho \"SYNC_MATERIALIZED=ok\"\n\n# Fail fast on a stale tree. A per-step\ + \ sandbox can be seeded from an\n# older orchestrator archive, and five consecutive runs burned\ + \ ~20\n# minutes each producing diffs that reverted merged work \u2014 a stale\n# tree diffed against\ + \ fresh main looks like a wholesale revert. The\n# guards at delivery caught them, but only after\ + \ the cost was paid.\n#\n# ops/FORBIDDEN_PATHS lists paths that must NOT exist. Their presence\n\ + # here means this sandbox is not the tree we uploaded, and nothing\n# built on it can be trusted.\n\ + if [ -f ops/FORBIDDEN_PATHS ]; then\n stale=\"\"\n while IFS= read -r forbidden; do\n case\ + \ \"$forbidden\" in ''|\\#*) continue ;; esac\n [ -e \"$forbidden\" ] && stale=\"$stale $forbidden\"\ + \n done < ops/FORBIDDEN_PATHS\n if [ -n \"$stale\" ]; then\n echo \"SYNC_FAIL_STALE_TREE: this\ + \ sandbox contains paths that do not exist on the base:\" >&2\n for p in $stale; do echo \" \ + \ $p\" >&2; done\n echo \" The workspace was seeded from an older archive, so it is not the\ + \ tree\" >&2\n echo \" that was uploaded. A diff computed from it reverts merged work.\" >&2\n\ + \ echo \" Failing now rather than spending a full cycle to produce an unusable diff.\" >&2\n\ + \ exit 75\n fi\nfi\n\ngit rev-parse --git-dir >/dev/null 2>&1 || git init -q\ngit config user.email\ + \ \"lead@relayflows.local\"\ngit config user.name \"Relayflow Lead\"\n\nif git remote get-url origin\ + \ >/dev/null 2>&1; then\n # Real checkout (laptop / fleet node): take the true origin/main.\n \ + \ echo \"SYNC_MODE=remote\"\n git fetch --quiet origin\n git checkout --quiet -B main origin/main\n\ + \ base=$(git rev-parse --short origin/main)\nelse\n # Sandbox snapshot: there is no remote to\ + \ fetch and nothing to\n # rebase onto. The snapshot IS the base. Commit it so the tick has\n \ + \ # a parent to diff against \u2014 `git diff main` in the review step\n # needs a `main` that\ + \ exists.\n echo \"SYNC_MODE=snapshot\"\n if ! git rev-parse --verify --quiet HEAD >/dev/null\ + \ 2>&1; then\n git add -A\n git commit --quiet -m \"snapshot base for this tick\" || true\n\ + \ fi\n git branch --quiet -f main HEAD 2>/dev/null || git checkout --quiet -b main\n base=$(git\ + \ rev-parse --short HEAD)\nfi\n\ngit checkout --quiet -B \"flow/drive-${base}-$(date +%m%d%H%M)\"\ + \necho \"SYNC_BASE=$base\"\necho \"SYNC_BRANCH=$(git rev-parse --abbrev-ref HEAD)\"\necho SYNCED\n" + - name: assess-1 + type: agent + agent: lead + dependsOn: + - sync + task: "You are the Relayflow Lead (charter/LEAD.md). Assess the repo.\n\nYOUR SCOPE IS THE TASK YOU\ + \ WERE GIVEN. Two launchers exist and they\ndeliver it differently: ops/launch-gate.sh commits an\ + \ ops/TARGET.md\nnaming one gate, while the autodrive loop passes the task directly\nand writes\ + \ NO TARGET.md. If ops/TARGET.md is absent that is normal \u2014\nit is not missing context and\ + \ there is nothing to go looking for.\n\nEither way: QUOTE the scope into ops/NEXT.md, never cite\ + \ the path.\nTARGET.md lives only in the throwaway launch worktree and is NOT in\nthe delivered\ + \ diff, so a reviewer sees a reference to a file that\ndoes not exist. Review flagged that on PR\ + \ #19 and again on #35, #40\nand #48 \u2014 it is now enforced in verify, which REFUSES a NEXT.md\ + \ that\ncites a path not present in the tree. Anything you rely on must\nappear in the package itself.\n\ + \nAnd when you state that something passes, paste the literal command\nand its output. \"Three tests\ + \ pass\" with no captured output is not a\nclaim a reviewer can check, and it was also flagged on\ + \ PR #19. This\nis AGENTS.md's central standard, applied to your own reporting. It is the operator's\ + \ scoping decision and it overrides your\nown judgement about priority \u2014 several runs execute\ + \ in parallel, each\npinned to a different gate, and a run that wanders outside its target\nwill\ + \ collide with a sibling. Stay inside it or, if the target is\ngenuinely unreachable, say so in\ + \ ops/NEEDS_HUMAN.md rather than\nsilently choosing different work.\nThen read ops/STATE.md \u2014\ + \ it is ground truth about gates and open\nPRs for an environment with no git history, and it names\ + \ the known\nsandbox faults that are NOT reasons to block. Then read\nops/DIRECTIVES.md \u2014 standing\ + \ human directives outrank the backlog;\nif one is unsatisfied, it IS the work package.\nThen read\ + \ docs/bootstrap-report.md and ops/DRIVE-LOG.md if they exist,\n`git log --oneline -15`, `gh pr\ + \ list --state open` and open PR review\nstate, kernel/ and packages/sdk/ test status. Then write\ + \ ops/NEXT.md: the\nSINGLE highest-priority work package toward the current gate\n(gate 1 until\ + \ its done-when in RFC-0001 \xA73 holds), with: objective,\nfiles in scope, definition of done (must\ + \ include passing commands),\nand what is explicitly OUT of scope for this tick. If an open PR is\n\ + awaiting fixes from review, the work package is fixing it \u2014 never\nstart new work over unfinished\ + \ work. If work is blocked on a human\ndecision, write ops/NEEDS_HUMAN.md stating the exact question\ + \ and the\noptions \u2014 and then STILL end with ASSESS_DONE.\n\nCOMMIT YOUR WORK PACKAGE BEFORE\ + \ YOU FINISH:\n git add -A && git commit -m \"assess: work package for this tick\"\nEach step runs\ + \ in its OWN sandbox and files reach the next step only\nthrough the executor's propagation, which\ + \ is lossy: on runs a2089144\nand 2560e02d your predecessor wrote ops/NEXT.md, said so truthfully,\n\ + and the file never arrived \u2014 one of those runs finished with a\nzero-file patch. Committing\ + \ puts the package in git history rather\nthan leaving it as a loose working-tree file. If the commit\ + \ fails,\nsay so in your output rather than finishing silently. The assess-gate step\nbelow reads\ + \ that file and parks the run with a typed outcome.\nALWAYS end with ASSESS_DONE, blocked or not:\ + \ this gate cannot tell a\ndifferent final token from a crashed agent, so on run 54ebd998 the\n\ + Lead correctly reported BLOCKED_NEEDS_HUMAN three times and was\nscored as failing three times.\ + \ Saying you are blocked is a result,\nnot a failure \u2014 but it must be said in the file, not\ + \ the token.\n" + verification: + type: output_contains + value: ASSESS_DONE + timeoutMs: 1800000 + - name: assess-gate-1 + type: deterministic + dependsOn: + - assess-1 + command: "# A typed park, not a crash. The assess step cannot express \"blocked\"\n# in its final\ + \ token (its gate only recognises ASSESS_DONE), so the\n# Lead writes ops/NEEDS_HUMAN.md instead\ + \ and this step reads it.\nset -u\nif [ -f ops/NEEDS_HUMAN.md ]; then\n echo \"ASSESS_BLOCKED_NEEDS_HUMAN:\ + \ the Lead escalated a decision it cannot make.\"\n echo \"--- ops/NEEDS_HUMAN.md ---\"\n cat\ + \ ops/NEEDS_HUMAN.md\n exit 75\nfi\nif [ ! -f ops/NEXT.md ]; then\n echo \"ASSESS_FAIL: no ops/NEXT.md\ + \ \u2014 an assessment that named no work package did not assess\"\n exit 1\nfi\n# The assessment\ + \ must have WRITTEN this tick's package, not merely\n# left the previous one in place. On run 457a6102\ + \ assess reported\n# \"The work package is written to ops/NEXT.md\" and the very next step\n# read\ + \ the OLD file \u2014 the logs carry the reason:\n# \"relayfile flush failed after the command\ + \ succeeded (exit 1);\n# a later agent step may see stale files\"\n# The builder then correctly\ + \ refused to invent scope, but only after\n# a whole build step had been spent. Catch it here instead:\ + \ if\n# ops/NEXT.md is identical to the base, the assessment did not land,\n# whoever is at fault.\n\ + # Look for the package in the working tree OR in a commit made this\n# tick. Propagation between\ + \ per-step sandboxes is lossy, so a package\n# that exists only as a loose file may not arrive;\ + \ one committed by\n# the assess step travels in git history instead.\nif git log --oneline main..HEAD\ + \ -- ops/NEXT.md 2>/dev/null | grep -q .; then\n echo \"ASSESS_PACKAGE_COMMITTED: found ops/NEXT.md\ + \ change in this tick's history\"\nelif git diff --quiet main -- ops/NEXT.md 2>/dev/null; then\n\ + \ # Warn, do not fail. This was fatal, and it killed four runs in six\n # while the loop produced\ + \ nothing \u2014 a worse outcome than the risk\n # it guarded against.\n #\n # The risk it guarded\ + \ was \"the builder gets scope nobody wrote this\n # tick\". But scope does not actually come from\ + \ ops/NEXT.md: it comes\n # from ops/TARGET.md, which the launcher COMMITS into the uploaded\n\ + \ # tree, so it is present in every per-step sandbox and cannot be\n # lost to the propagation\ + \ fault. NEXT.md refines the target; it does\n # not define it.\n echo \"ASSESS_WARN_STALE_NEXT:\ + \ ops/NEXT.md did not change from the base commit.\"\n echo \" The assess step's package did not\ + \ survive the step boundary (a known\"\n echo \" platform fault: per-step sandboxes lose both\ + \ loose files and git objects).\"\n echo \" Proceeding, because ops/TARGET.md is committed in\ + \ the tree and carries this\"\n echo \" run's scope. The builder is not working blind \u2014 it\ + \ is working from the\"\n echo \" target rather than from a refinement of it.\"\n if [ -f ops/TARGET.md\ + \ ]; then\n echo \"--- ops/TARGET.md (the scope that did survive) ---\"\n head -8 ops/TARGET.md\n\ + \ else\n echo \"ASSESS_FAIL_NO_SCOPE: neither a fresh ops/NEXT.md nor an ops/TARGET.md.\"\n\ + \ echo \" With no scope from either source the builder WOULD be working blind.\"\n exit 1\n\ + \ fi\nfi\n# A package with no definition of done cannot be verified, and the\n# builder cannot\ + \ honestly report BUILD_DONE against it.\n# A package must be verifiable, but do not dictate its\ + \ wording. This\n# check demanded the literal phrase \"definition of done\" and so\n# rejected a\ + \ CORRECT assessment three times on run 30475b25 \u2014 one\n# that reported gate 2's primitives\ + \ already complete and proposed\n# moving to gate 3, and was right on both counts. A gate that\n\ + # rejects true reports is as bad as one that accepts false ones.\n#\n# Accept either shape: a runnable\ + \ command (that is what \"verifiable\"\n# actually means), or an explicit statement that this tick\ + \ has no\n# buildable package.\nif grep -qiE \"definition of done|definition-of-done|done when|done-when|acceptance\ + \ criteria\" ops/NEXT.md \\\n || grep -qE \"(cargo|npm|node|sh|pytest) [a-z]\" ops/NEXT.md \\\n\ + \ || grep -qiE \"no buildable work|nothing to build|assessment only|gate .* is (green|complete)\"\ + \ ops/NEXT.md; then\n :\nelse\n echo \"ASSESS_FAIL_NO_DOD: ops/NEXT.md names neither a runnable\ + \ command nor a\"\n echo \" statement that this tick has no buildable package. A work package\ + \ that\"\n echo \" cannot be verified cannot be built against.\"\n exit 1\nfi\necho \"ASSESS_GATE_PASS\ + \ ($(grep -m1 -oE 'WP-[0-9]+[^|]*' ops/NEXT.md || echo 'work package'))\"\n" + timeoutMs: 120000 + - name: build-1 + type: agent + agent: builder + dependsOn: + - assess-gate-1 + maxIterations: 3 + task: "Read ops/NEXT.md, AGENTS.md, and the relevant parts of\ndocs/RFC-0001-everything-is-a-relayflow.md.\ + \ Implement exactly that\nwork package \u2014 nothing more. Run the definition-of-done commands\n\ + yourself and iterate until they pass. Keep files small and\nsingle-purpose. End with BUILD_DONE\ + \ only when the definition of done\npasses locally; paste the passing output.\n" + verification: + type: output_contains + value: BUILD_DONE + timeoutMs: 5400000 + - name: verify-1 + type: deterministic + dependsOn: + - build-1 + command: "# CLOUD VARIANT (generated): a FAILED verify is recorded and the\n# run continues. Nothing\ + \ is delivered from a sandbox, so a failure\n# here cannot ship; the next cycle's assess treats\ + \ it as the work\n# package. On a delivering environment verify stays fatal.\n# A gate that cannot\ + \ fail is not a gate. Never pipe a test command\n# into tail inside the status check: the pipeline's\ + \ status is tail's.\nset -u\nran=0; ok=0\n\n# Bound every long-running command, not just the suites.\ + \ Run\n# 6d045b23 sat in verify for 29+ minutes: its suites were bounded but\n# `cargo build` and\ + \ `npm ci` were not, so a cold sandbox installing a\n# toolchain and compiling from scratch had\ + \ no ceiling at all.\n# Bounding half the step is not bounding the step.\n#\n# timeoutMs is NOT\ + \ enforced by the platform \u2014\n# observed three times on 2026-08-28 (verify-1 at 31min against\ + \ a\n# 20min bound, review-1 at 36min against 30min, plus an unbounded\n# toolchain install). And\ + \ the kernel suite now contains a test that\n# intermittently hangs under sandbox timing:\n# an_entry_appended_during_watch_registration_is_delivered_exactly_once\n\ + # ran past 60s in cloud while passing locally in 0.54s. Without a\n# bound here, one hanging test\ + \ consumes the entire run budget.\nrun_bounded() {\n _label=\"$1\"; shift\n if command -v timeout\ + \ >/dev/null 2>&1; then\n timeout \"${VERIFY_SUITE_TIMEOUT:-900}\" \"$@\"\n elif command -v\ + \ gtimeout >/dev/null 2>&1; then\n gtimeout \"${VERIFY_SUITE_TIMEOUT:-900}\" \"$@\"\n else\n\ + \ echo \"VERIFY_WARN: no timeout(1); $_label runs unbounded\" >&2\n \"$@\"\n fi\n}\n\nif\ + \ [ -d kernel ]; then\n # Invoke through `sh`: in a cloud sandbox this script was present\n #\ + \ but not executable (observed on run 4cf36ea7). Git tracks it as\n # mode 100755, so the exec\ + \ bit is lost somewhere in materialization\n # \u2014 which stage is NOT established, so no mechanism\ + \ is claimed here.\n # `sh