Skip to content

feat: harvest 5 — v1.0.0: guards that resolve paths, the block as a gate, review bots on aliases, the Opus 5.5 recalibration, and the first release - #75

Merged
j4th merged 135 commits into
mainfrom
feat/harvest-5-v1.0.0
Oct 2, 2026
Merged

j4th merged 135 commits into
mainfrom
feat/harvest-5-v1.0.0

Conversation

@j4th

@j4th j4th commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Closes #58
Closes #60
Closes #61
Closes #62
Closes #63
Closes #64
Closes #65
Closes #66
Closes #67
Closes #68
Closes #69
Closes #70
Closes #71
Closes #72
Closes #73
Closes #74

Harvest 5 lands every issue the three sister repositories filed since v0.5.0, plus a review of the whole kit, as v1.0.0, the kit's first release. The design is docs/superpowers/specs/2026-09-30-cascade-kit-harvest-5-design.md, and the trace beside it has one row per ask; the plans are under docs/superpowers/plans/. Tags and the GitHub Release follow the merge, and only on the operator's go-ahead (plan F5).

Summary by cluster

Decisions (settled with the operator, 2026-09-29/30)

  • D41 Back-tags v0.1.0–v0.5.0 on the past harvest merges and v1.0.0 on this one, after merge. D42 One PR, atomic commits, block green after each. D43 Ship the shared path resolver (reverses harvest 4's no-helper non-goal). D44 Ship all three behavioural hook fixtures. D45 Ship the 2–4-arm finish-ab and the headless runner. D46 review-sweep dedups on file and line and carries every title. D47 Review bots on family aliases, never best, effort explicit, ANTHROPIC_MODEL never set. D48 A paths: trigger whose re-includes of .claude/** and docs/adr/** come last. D49 The Opus 5.5 / Sonnet 5.5 recalibration. D50 Record the always-loaded bytes; warn above 140,000. D51 The block as a gate. D52 Target hygiene. D53 Kit issues cited as context-builder-kit#N. D54 A PR that closes an issue carries its key in the branch. D55 The frozen-corpus backstop's bypass list. D56 The whole-kit review's 68 findings join v1.0.0, the unverified re-checked at their commit. D57 disable-model-invocation on the five commands. D58 .mcp.json.example rebuilt (${VAR}, type: "http", no GitHub server). D59 Split knowledge-backend.md to pay for the recalibration (target: no net growth). D60 Install the drop-in set from a tagged archive; upgrade with git merge-file. D61 An unreadable payload denies (hard-deny) or asks (ask-gate). D62 Bracket-split the private target's name at HEAD.

Always-loaded rules — before and after

File v0.5.0 (main) v1.0.0 Δ
cbk-conventions.md 30,058 31,064 +1,006
knowledge-backend.md 20,336 9,513 −10,823
orchestration.md 18,562 25,084 +6,522
pr-review.md 28,281 30,448 +2,167
simplification.md 3,213 3,564 +351
tooling.md 12,416 13,288 +872
workflows.md 17,762 19,842 +2,080
Total 130,628 132,803 +2,175

D59's target was no net growth, and it is missed by 2,175 bytes (1.7%). The knowledge-backend.md split freed 10,823 bytes. The recalibration's per-model and per-surface effort facts and the review round's corrections took more than that; no operative text was cut to make the number. The total sits under the 140,000-byte warning line.

Probes (spec § Probes at execution)

  • P1 — a hook payload's cwd follows the Bash tool's cd (Claude Code 2.1.286, 2026-09-30, variant B). require-repo-root-for-agents.sh now denies a dispatch made after a cd, and protect-main-branch.sh judges the checkout the shell is in.
  • P2 — a workflow agent has no Agent tool, plain or worktree-isolated (outcome A, 2.1.286). pr-review.md § The floor says an unattended floor that must fan out is a headless claude -p session.
  • P3 — the record-model step ran green in a target's review run, but the rendered job-summary line was not read back: designed-unexercised, dated 2026-09-30.
  • P4 — unshare -rm on ubuntu-24.04: denied, so the three bind-mount cases SKIP in CI (47 cases ok there, 50 locally where they run), https://github.com/j4th/context-builder-kit/actions/runs/36953707599 — recorded, dated, in the fixture's header (e5f4090).
  • P5 — claude mcp list (2.1.286): Linear and Notion answer 401 (auth required, so the endpoints are live), context7 200; mcp-server-time 2026.8.18 is the newest release and past the settle window.
  • P6 — the review sweep checked 153 double-quoted platform quotations in the rules raw, with grep -F, and found every one verbatim. The four claims it could not verify were re-read at their sources by four verifiers on 2026-10-01 and corrected (6856188).

Review gate

  • /simplify — ran: 4 cleanup agents, 28 findings (13 after dedup), 4 applied (c4d8c36, cd8a8ff, 576d697, 6b954a2) · the agents ran on Sonnet at the session's effort (xhigh); the Agent tool takes no effort
  • pr-review-toolkit:review-pr — ran: 5 agents (code-reviewer, pr-test-analyzer, comment-analyzer, silent-failure-hunter, type-design-analyzer; code-simplifier left to /simplify), 78 findings, triaged 63/9/4/0/2 · the agents set no model or effort, so they ran on the session's Opus 5.5 at its effort, xhigh (orchestration.md § The effort axis)
  • review-sweep — ran: 11 finders + 8 verifiers, 6 confirmed / 2 refuted / 36 unverified, dropped coverage: none, bounds 3/8 · the changed-file list leaves out docs/superpowers/ (the planning record, not shipped); every unverified finding was triaged at the caller (below)
  • Not a second floor: one verification workflow, run 2026-10-01, re-read at their sources the four rule-text claims the sweep could not verify (4 read-only verifiers, workhorse tier, high; all four corrected, 6856188). One refute-by-default verifier checked the trace's two non-landed rows, and both held. On the operator's follow-up (below), one more verification workflow checked that delta: 3 read-only verifiers, workhorse tier, high (the reader's code, an independent recomputation of every re-derived figure, every other claim). It found 5 code defects and 2 wording errors, all applied in a037350 with red-first cases, and confirmed every figure but one range, which was corrected.

Triage

Toolkit — 78 findings: Apply 63 · Apply with care 9 · Surface 4 · Defer 0 · Reject 2. Sweep — 44: Apply 28 · Apply with care 3 · Surface 1 · Reject 12. Every applied finding is its own commit, red first where it is behavioural (the red-first table below), with the block green after each.

Applied, by area

  • Hooks: 12f289d (a non-string payload field, which had run as a command, and an empty payload are refused) · 55a89d2 (fail-open warnings shown as a systemMessage) · 9a366a8 (backslash-newline continuations) · 349619e (an unborn branch's root commit) · d752a79 (the ADR guard's two path readings, each pinned) · 7f1dade (matchers checked against their tools) · 0e2e324 (bun.lockb, npm-shrinkwrap.json; the ask-gate's cases).
  • Harness: ec17e0a and 2a798c7 (agent-cost: one bill per response, 1-hour writes at 2x, [1m] only) · 2bed51b (run-arms-headless: strict config, arm crash recorded, gh deny forms, CLAUDE* provider and auth kept, worktree-add failure handled) · 5684420 (finish-ab: arguments, efforts, anon charset, the executed cell, judge scores) · 533ece4 (review-sweep: arguments, roster, key collisions, path normalization).
  • Review templates: 24dc6bb (the rule quotation, kept live) · 602dfc4 (claude.yml gates in the job if:) · 83408c6 (a per-path dispatch rule) · e07bed1 (a bold verdict marker, **SKIPPED**, the moved-base notice).
  • Block: 0e5c278 (.mcp.json jq status) · 700d321 (a misspelled rule citation) · 9428a63 (:(glob) case; bounds pinned by behaviour) · e6da512 (pointer claims on all four halves) · c6d7367 (Linear taxonomy) · 8063291 (the axis name).
  • Docs: 6b36fdf (blueprint § Workstreams and § Manual setup) · d3749f0 (the floor's docs-only scoping; a "not run" gate form) · 11bba65 (rust-toolchain.toml coverage) · c098399 (CLAUDE.md) · 6856188 (four platform claims) · 9a31317 (the action-release citation, bump timing) · ae489c8 (README: install, upgrading, docs/, the ADR opt-out) · 6ba4971 (README example) · 7147791 (CHANGELOG).
  • Apply with care: the payload-field fix and its three duplicate reports (12f289d), the systemMessage channel across six hooks (55a89d2), the arm-crash wrapper (2bed51b), the two template sections (6b36fdf), and the README install rewrite (ae489c8, which a scratch-repository run exercised).

Surface — not applied; the reviewer's case, verbatim in substance

  1. finish-ab — "panel balance ignores judge effort, so reading order can be fully confounded with effort" (type-design feat: second harvest of real-run cascade learnings into the kit #5). A per-effort Latin square changes the panel's shape.
  2. The four project checks the branch added "have never run, red or green" (pr-test Review gate: the orchestrated sweep is framed as a substitute for direct dispatch, contradicting /finish's own "does not skip" rule #8). Testing them needs a filled-target fixture tree, which is new test infrastructure.
  3. The formatter floor "is pinned only by a source-text grep" (pr-test review-sweep.js fans out one workhorse-tier verify agent per finding, uncapped and undeduped #10). Driving it with relative and .. paths matters once a target fills a formatter arm.
  4. README's Tuitor worked example, which V9.20 removed from .claude/ (code-reviewer Review gate: the orchestrated sweep is framed as a substitute for direct dispatch, contradicting /finish's own "does not skip" rule #8). The README's own illustration is a taste call.
  5. run-arms-headless — worktree_root "must be gitignored" is stated but never checked; the mcp_config allowlist checks server names only; and the runner's facts are appended to the arm's notes rather than kept in a structured result.runner (type-design No proportionality rule scaling the review mechanism to diff size — and session-level effort silently overrides judgement #12, No rules file carries paths: frontmatter and the kit never mentions the mechanism — every rule loads every session #15).
  6. The ADR guard refuses a new file in a not-yet-existing docs/adr subdirectory when an ADR there shares its basename (sweep).

Reject

  • review-sweep merges same-line findings into one verdict (code-reviewer feat: axis parity pass — reconcile the kit with its exercised reality #7, type-design feat: harvest real-run cascade learnings into the kit #4, sweep ×4): by design, since D46's verify prompt asks which report the evidence proves; refuted by the sweep's own verifier.
  • The KB ask-gate misses claude.ai connector tool names (sweep): refuted; the hook's Not seen: line now names it.
  • A missing resolve-path.sh fails a hard-deny guard open: the contract (§ Hook authoring): a missing dependency fails open and names its backstop.
  • Advisory hooks skip an unparseable payload: advisory by contract; the check task is the backstop.
  • Opus 4.x and provider-qualified ids are unpriced: they are named and excluded from the total, never guessed.
  • A ~ in file_path is taken literally: the harness sends absolute paths; ~ is a shell expansion.
  • Version stamps disagree: each stamp dates the release it was verified on.
  • The fast path's --fallback-model sonnet equals its model: CLI 2.1.287 accepts the pair (exit 0), and the action passes the flag as a CLI argument, never as the SDK option that refuses it.
  • The assertion's two-dot diff: correct, because the action compares contents, so a base that moved skips the review too. The notice wording was fixed (e07bed1).

/simplify skipped (behaviour or scope): every verification-block cleanup, since simplification.md keeps the pass out of .claude/rules/ (a follow-up issue candidate); a shared fixture library, which would be a new drop-in file with a sync note; the review template's output consolidation; review-sweep's reports[] interface; parallel worktree_setup; the lock-file guard on the resolver, which is a behaviour change (a follow-up candidate).

Operator-directed follow-up — "fix the agent-cost double counting too" (2026-10-01)

The per-line double count was already fixed (ec17e0a). Two more defects turned up on this machine's transcripts:

  • A second double count: forks (6dfcade, a037350). A fork's transcript opens with copies of its parent's history, same message ids, and the reader deduplicated only within a file. So a parent's spend was billed again for every fork: $2.48, $4.28 and $5.40 on the three fork-bearing session directories, the last a fifth of the real total. Each response is now billed once per directory, at its best copy. Forks are read after their parents at every level. A fork whose parent is missing is cut at the message that hands it its task, and one that cannot be cut is unpriced and named. Fork copies carry the parent's final usage where the parent's own line is a mid-stream snapshot, so parents are now billed for output the old reader never saw.

  • An undercount, named rather than hidden (c5a3f26). From Claude Code 2.1.278 most subagent responses are written mid-stream ("stop_reason": null, output usually under ten tokens), and nothing else in the transcript records the final count. Such rows are billed as recorded and named a floor.

  • The kit's dated cost figures, re-derived (eaa5cb8). The A/B runs behind orchestration-reference.md § Applied instances were recomputed from their original transcripts, and the figures replaced:

    Run Recorded Re-derived
    /finish executors about $50 each $32.54 and $33.01
    /finish judges $18–23 $9.12–$11.81
    Framing, top:workhorse 3.2× 1.48×
    Rough-in, xhigh judge 15% more 8% more

    On those runs the per-line count was the whole error: the recorded figures ran 1.5× to about 2.2× high.

Trace

371 rows: 369 landed, 1 not holding (#69/c5859811353/H1 — --disallowedTools is variadic), 1 non-goal (#58/c5901493591/residue-2 — D31's fail-fast stands), 0 open. Every landed row was audited against the tree at HEAD. The five the audit found partial carry the review round's completing commit, and release/7's tags are pushed after merge. A refute-by-default verifier could not show either non-landed row required.

Red-first table (130 rows)
Task Test The failure it showed first
V1.1 run-verification-block-fixture.sh FAIL: a filled target whose project sub-block never ran fails (want exit 1, got 0)
V1.2 block, Stop-hook negative case the Stop-hook check passed a hook that could not look (exit 0 with a WARNING) — in a gate that is red
V1.3 block budget line (padded clone) always-loaded total: 140630 bytes, warn-lines=0
V1.4 block pin (disposition row) .claude/rules/cbk-conventions-reference.md ships a paths: placeholder, but scaffold's rule-file disposition table has no row for it
V1.4 fresh scaffold (helper) grep: CLAUDE.md: No such file or directory / verification: block exited 2
V1.5 block pin (templates name the runner) .claude/skills/blueprint/references/templates/tooling.md does not name .claude/workflows/tests/run-verification-block.sh (the template that wires the block into check)
V1.6 block pin (spellchecker note) § Verification lacks the bracket idiom's file-scoped spellchecker exemption
V1.7 block pin (kit verify.yml) .github/workflows/verify.yml lacks defaults.run.shell: bash or runs-on: ubuntu-24.04
V2.1 adr-ci-body-fixture.sh FAIL: modify a non-ASCII-named ADR (want exit 1, got 0)
V2.1 unfixed ADR body (direct) stale base: exit=1 (the unfixed body fails a PR that touched no ADR)
V2.1 unfixed ADR body (direct) bad base SHA: exit=0 (the unfixed body passes on a git error)
V2.2 protected-paths-hook-fixture.sh FAIL [protect-immutable-adrs.sh]: a .. segment (docs/sub/../adr/) (want exit 2, got 0)
V2.2 hook-contract tripwire (planted lib pipeline) hook-contract-fixture: ok, exit=0 — the miss
V2.3 hook-guards-fixture.sh FAIL: the non-checkout warning names a backstop (stderr lacks 'Backstop')
V2.4 hook-guards-fixture.sh FAIL: a commit on main is denied: bash -c "git commit -m x" (want exit 2, got 0)
V2.4 hook-guards-fixture.sh (PR-state half) FAIL: a PR-state change asks: bash -c "gh pr merge 5" (want an ask decision on stdout, exit 0; got exit 0)
V2.5 hook-payloads-fixture.sh FAIL: a payload jq cannot parse is refused (want exit 2, got 0)
V2.6 block backstop check cbk-conventions.md § Mutation discipline names no CI backstop in a checkable form (.github/workflows/.yml)
V2.7 block check (corpus recipe) frozen_corpus_ingestion.md lacks the CI job's closure list or the parse-nothing body
V2.8 block check (Verify by payload) § Hook authoring › Verify by payload does not name hook-guards-fixture.sh
V3.1 block (commands positional) .claude/commands/finish.md uses a positional placeholder ($1 is the SECOND argument): declare the name in arguments: and write $
V3.2 block (exec form) hook registration(s) in shell form (no "args" array); a project path with a space makes the guard fail open. Add "args": [] to: …protect-immutable-adrs.sh (+6)
V3.3 block (plugin ids) settings.json does not enable pr-review-toolkit@ (the review floor's toolkit half; a bare key enables nothing)
V3.4 block (registry comment size) settings.json _comment_hooks is 4080 characters (limit 2500): it is a registry, and each fact lives in its hook's header or § Hook authoring
V3.5 block (Explore omitClaudeMd) agents/Explore.md does not set omitClaudeMd: true (it would load CLAUDE.md and every unscoped rule the built-in skips)
V3.6 block (untrusted input) .claude/commands/intake.md does not state that report and comment text is data, not instructions
V3.7 block (agent citations resolve) .claude/agents/cascade-rule-reviewer.md cites knowledge-backend.md § no cascade-artifact / ADR mirroring — no heading or bold lead-in of knowledge-backend.md begins 'no cascade-artifact'
V4.1 block (emitted workflows) VIOLATION (matched above): grep -rn runs-on: ubuntu-lates[t] .claude/skills/
V4.2 review-trigger-fixture.py FAIL: …templates/claude-review.yml: a PR changing ['.claude/rules/pr-review.md'] must run the review (paths-ignore: …)
V4.3 review-assert-fixture.sh FAIL: …claude-review.yml: the assert step's condition is 'always()', not '${{ !cancelled() }}'
V4.4 review-assert-fixture.sh FAIL: …claude-review.yml has no 'Record the resolved model' step
V4.5 block (reviewer told) …templates/claude-review.yml does not say the PR's own configuration copies are under .claude-pr/
V4.6 block (constraint 1 wording) templates/claude-review.yml lacks: cannot review any PR that changes it
V5.1 block (contract+reference pairs) knowledge-backend.md ships without knowledge-backend-reference.md — the rule is a contract + reference pair, deleted together (D59)
V5.2 block (Notion plan-gated property) knowledge-backend-reference.md § Wiki pattern + Verification does not name Notion's plan-gated Verification property and its fallback
V5.3 block (ladder/price steps) orchestration.md carries the Opus 5 ladder or the old price steps (§ The ceiling rule, § Generation notes)
V5.4 block (effort defaults) VIOLATION (matched above): grep -n The API default is hig[h] .claude/rules/orchestration.md
V5.5 block (surfaces heading) VIOLATION (matched above): grep -rn three surfaces + resolution orde[r]|§ The three surfaces an[d] .claude/
V5.6 block (caps/teammate/read-only) orchestration.md lacks the workflow concurrency variable, the teammate trap, or the read-only-agents bullet
V5.7 block (text-only end of turn) orchestration.md § Anti-patterns lacks the text-only end-of-turn row (context-builder-kit#69)
V5.8 block (triad checklist) VIOLATION (matched above): grep -n the harness's task lis[t] .claude/rules/workflows.md
V5.9 block (delegating implementation) workflows.md § Subagent dispatch lacks the delegating-implementation clause (context-builder-kit#69)
V5.10 block (workflow agent has no Agent tool) pr-review.md § The floor does not name the workflow agent among the contexts with no Agent tool (probe P2)
V5.11 block (/simplify changelog history) VIOLATION (matched above): grep -n as of 2.1.26[3] .claude/rules/simplification.md
V5.12 block (grounding rule 4) rough-in research-phase.md § Grounding existence claims lacks rule 4 (which binary a tool resolves to)
V5.13 block (no slot in a reference half) VIOLATION (matched above): grep -n [[Rr]ecor[d] .claude/rules/orchestration-reference.md
V5.14 block (Primary sources rows) orchestration-reference.md quotes code.claude.com/docs/en/agent-teams but § Primary sources has no row for it
V6.1 agent-cost-fixture.sh FAIL: fff: cost 14.25, want 11.2 (Opus 5.5: 1M in @ $4 + 1M cache write @ 1.25x + 1M cache read @ 0.05x + 100k out @ $20 — never Opus 5)
V6.1 block list-price diff under busybox > LOpus 5 25 (review/portability/67)
V6.1 block list-price diff the list-price diff read nothing (PRICE: 0 keys; quoted bullet: 4 models) — a format changed; update this extraction
V6.1 block cache-read diff the cache-read diff read too little (CACHE_READ: default 0.1; | quoted sentence: ;) — a format changed; update this extraction
V6.2 finish-ab-shape.mjs Error: finish-ab: args.arms must name exactly two arms
V6.3 run-arms-headless-fixture.sh FAIL: .claude/workflows/finish-ab/run-arms-headless.py is missing (python3 would exit 2 on it, which case 1 would read as a refusal)
V6.3 finish-ab-shape.mjs scenario 19 FAIL: run-arms-headless.py's ARM_SCHEMA could not be read (Command failed: python3 -B -c import importlib.util, json, sys)
V6.4 block arm-count check VIOLATION (matched above): grep -nE [Tt]wo executor[s]|two-ar[m]|rank the tw[o]|either worktre[e]|two arms shar[e] …
V7.1 review-sweep-accounting.mjs AssertionError [ERR_ASSERTION]: the degrade path logs WHY the roster read threw, not only that it degraded (context-builder-kit#72 item 3)
V7.2 review-sweep-accounting.mjs AssertionError [ERR_ASSERTION]: the caller's focus reaches the finder's prompt
V7.2 block (caller-finder shape) .claude/rules/pr-review.md does not state the caller-finder shape {key, prompt?, agentType?}
V7.3 review-sweep-accounting.mjs AssertionError [ERR_ASSERTION]: paraphrases on one line take one verify slot (got 3)
V7.3 block (dedup key) pr-review.md invariant (3) does not key dedup on file and line, or does not triage a merged finding by the report its verifier named
V7.4 review-sweep-accounting.mjs AssertionError [ERR_ASSERTION]: find:code-review carries the read-only clause
V7.5 review-sweep-accounting.mjs AssertionError [ERR_ASSERTION]: find:code-review says a check-in gets no answer
V7.6 block (interrupted run re-run fresh) pr-review.md does not say an interrupted floor or sweep is re-run fresh with its reviewer memory discarded
V7.7 block (one verification workflow per delta) pr-review.md § The floor › Once lacks the one-verification-workflow rule
V7.8 block (review layers independent) pr-review-reference.md § Anti-patterns lacks 'Folding one review layer into another'
V7.9 block (family of directories enumerated) the roster guidance for a family of directories is missing from pr-review.md's craft rule or pr-review-reference.md § Authoring
V7.10 block (gate consolidation is a new draft) rough-in references/research-phase.md does not treat a change made at the gate as a new draft
V7.11 block (pr-respond NOT-list) pr-respond.md's NOT-list contradicts Step 7's round-block append
V7.12 block (rubric names its classes) pr-review.md § Triage rubric does not name its four classes and the Apply-with-care variant
V7.13 block (docs-only skips sweep, not floor) VIOLATION (matched above): grep -nE the simplify pass is (enoug[h]|sufficien[t]) …
V7.15 block (STANDARDS citations) VIOLATION (matched above): grep -nE STANDARDS.md`? § (Step [0-9]|PR review proces[s]) …
V8.1 block (container images) cbk-conventions-reference.md § Dependency settle-window does not name the COPY --from gap
V8.2 block (mise task_config) blueprint templates/tooling.md lacks the mise [task_config] shell line or its version floor
V8.3 block (.mcp.json shape) .mcp.json.example: a url entry with no type is skipped by Claude Code: notion context7
V8.4 block (.gitignore harness block) github-starter-templates.md lacks the harness-block section or its fence
V8.5 block (formatter floor) format-on-edit.sh's floor does not skip .claude/workflows/, or its arms do not name the forcing flag
V9.1 block (executor quartet) .claude/commands/finish.md: Step 1 does not admit a meta's child as the fifth title form
V9.2 block (capstone close marker) .claude/commands/finish.md: item 8 does not name the milestone issue in a capstone PR's close markers
V9.3 block (D54 branch rule) cbk-conventions.md § Branch naming or its Quick reference row lacks the D54 branch rule
V9.4 block (STANDARDS headings) VIOLATION (matched above): grep -rnE STANDARDS.md`? § (Testing philosoph[y]|PR feedback loo[p]|PR review proces[s]|Commit and branch convention[s]) .claude/
V9.5 block (retired frame section) VIOLATION (matched above): grep -rnE Deferred meta-issue[s] .claude/skills/
V9.6 block (adr-new relation grains) adr-new/SKILL.md lacks ## Relation grains (the pointer to § ADR relation grains)
V9.7 block (dated rail, schema field) VIOLATION (matched above): grep -rn as of current cascade versio[n]|subIssueProgres[s] .claude/skills/
V9.7 block (four backends.md copies, mutation) rough-in/references/backends.md drifted from scaffold's copy (the four copies are byte-identical)
V9.8 block (eight sections) .claude/skills/scaffold/SKILL.md does not carry the executor's section list (Context / Assumptions / … / PR contract)
V9.9 block (AC form) …/rough-in-spec-template.md: the default acceptance criteria are not numbered [R<#>.AC] (the rough-in contract's form)
V9.10 block (one name per axis value) VIOLATION (matched above): grep -rnE github-only profil[e]|[Mm]arkdown-only profil[e]|profile: markdown-onl[y]|profile fiel[d] is .claude/
V9.11 block (methodology register) VIOLATION (matched above): grep -rn methodology_register.m[d]|sdd.m[d]|… .claude/
V9.12 block (CI-skip spellings) cbk-conventions.md § [skip ci] rule does not name: [ci skip]
V9.13 block (Linear free plan) manual_steps.md: the Linear free-plan line cites no dated pricing page
V9.14 block github_only_profile.md's label taxonomy lacks triage (a flow applies it)
V9.15 block VIOLATION (matched above): grep -rnE four detection state[s]|State [4] (no MCP)|because state [2])|Full automation (GitHub MCP|Skip the \.github/ issue template[s]|with five HITL gate[s] .claude/skills/scaffold/
V9.16 block VIOLATION (matched above): grep -n produces 3-7 R-issue[s] .claude/skills/rough-in/references/test_cases.md
V9.17 block framing/SKILL.md's phase-exit checklist lacks the restamp standing item
V9.18 block blueprint/SKILL.md counts six foundation docs where its table lists seven on two axes
V9.19 block .claude/rules/cbk-conventions-reference.md § .gitignore anchoring has no pointer heading in cbk-conventions.md and is not named in its preamble
V9.20 block VIOLATION (matched above): grep -rn -i tuito[r]|anubi[s]|per-Pilo[t]|Servo.set_angl[e] .claude/
V9.21 block (bare kit-issue citation) VIOLATION (matched above): grep -rnE (^|[:space:#[0-9]{1,3}\b .claude/rules .claude/hooks … .github
V9.21 block (sibling name, mutation) VIOLATION (matched above): git grep -n -i -E you-are-hea[r]|echospher[e] -- .claude .github README.md CLAUDE.md
V9.22 F1 denylist (counted) docs/superpowers/plans/2026-09-21-harvest-4-applied-defects.md:1 (hits=0)
V10.1 block (Kit commit form) scaffold_output_template.md's Kit commit row does not take the vX.Y.Z (sha) form
V10.1 block (Syncing names Sync notes) cbk-conventions-reference.md § Syncing the kit does not name CHANGELOG.md's Sync notes
V10.2 block (CHANGELOG sections) CHANGELOG.md has no [0.1.0] section with a ### Sync notes heading
V10.2 sim-sync.sh (Review Focus 1) FAIL: CHANGELOG.md has no [1.0.0] section with Sync notes
V10.3 block (README inventory) README.md's inventory tree does not name .claude/hooks/lib/resolve-path.sh
V10.3 install-test.sh (D60) FAIL: README.md's install block has no single-line curl download to swap for the local archive
V10.4 block (CLAUDE.md gate words) CLAUDE.md does not name run-verification-block.sh (its gate, its release record and its harvest audit)
V10.5 block (LICENSE holder) VIOLATION (matched above): grep -nF [name of copyright owner] LICENSE
F2 12f289d hook-guards / hook-payloads fixtures an array-valued field and an empty payload: the guard exited 1 and let the call through (the c4d8c36 regression)
F2 ec17e0a agent-cost-fixture.sh (rrr) one response written as two lines was billed twice
F2 55a89d2 hook-payloads-fixture.sh (shows) a fail-open warning reached stderr only, never stdout's systemMessage
F2 9a366a8 hook-guards-fixture.sh a git \-continued commit and a continued gh pr ready passed both Bash guards
F2 349619e hook-guards-fixture.sh an unborn branch's root commit was called "not a git repo"
F2 d752a79 protected-paths-hook-fixture.sh (mutants) dropping either path reading survived all 48 cases
F2 7f1dade block (matcher coverage) a guard's matcher narrowed off its tool left the block green
F2 0e2e324 hook-guards-fixture.sh / block bun.lockb and npm-shrinkwrap.json passed; the KB ask-gate's cases could be switched off silently
F2 2bed51b run-arms-headless-fixture.sh (11j–u, 18–20) an unknown key, a bad budget, dry-run with resume, an arm crash: accepted or lost before the fix
F2 5684420 finish-ab-shape.mjs a misspelled argument, an unknown effort, a judge with empty scores: accepted
F2 533ece4 review-sweep-accounting.mjs malformed arguments accepted; one defect at two path spellings counted twice
F2 24dc6bb review-trigger-fixture.py claude-review.yml quotes pr-review.md § What NOT to flag as 'Not automatically light …', which the rule no longer says
F2 602dfc4 review-assert-fixture.sh claude.yml's job if: lacks github.event.sender.type == 'User'
F2 83408c6 review-assert-fixture.sh (mutants) each path's dispatch rule against its tool list: 2 mutants killed (the rules were written beside the check)
F2 e07bed1 review-assert-fixture.sh a tracker comment that names a verdict word in passing (want exit 1, got 0); the docs-only skip line opens with no bold marker; the notice lacks 'differs from its base's'
F2 0e5c278 block (scratch clone) "headers": "Authorization: Bearer ghp_…" → exit 0
F2 700d321 block (scratch clone) cbk-convention.md § ADR index sync → exit 0
F2 9428a63 adr-ci-body-fixture.sh (no-:(glob) mutant) FAIL: edit a file in a directory named like an ADR (want exit 0, got 1)
F2 e6da512 block V9.19 (all four halves) orchestration-reference.md § Generation notes — the sources has no pointer heading in orchestration.md and is not named in its preamble
F2 2a798c7 agent-cost-fixture.sh (sss, ttt) sss: cost 10.0, want 13.0; ttt: must be unpriced, got 4.0
F2 c6d7367 block V9.14 (Linear) linear_planning.md's label taxonomy lacks workstream:<slug>
F2 8063291 block (axis name, against HEAD~) the new pattern matches 8 lines of the pre-fix text
Sync notes preview (CHANGELOG [1.0.0])

Sync notes

  • Hand-merge these files. This release rewrites text a filled target changes in each of them. Merge each three ways from the release your Kit commit row records, and keep what the project filled:
    • .github/workflows/claude-review.yml and .github/workflows/claude.yml, your copies of blueprint's review templates. Keep your label names, turn cap, sizing, allowed tools and the filled REVIEW_LOGIN. Take the model and effort lines (claude.yml now runs --effort high), the Record the resolved model step, the rebuilt "Assert the review posted" step, the fork guard, concurrency moved into the job, defaults: run: shell: bash, runs-on: ubuntu-24.04, the step ids review and claude, the prompt's restore section with its two new bracketed fills, the notes above claude_args, the label step's dispatch_rule output with the two prompt lines that read it, the assertion's bold verdict marker and the prompt's **SKIPPED** docs-only line, and in claude.yml the bot and skip-claude gates moved into the job if: (the gate step is gone). The paths: filter replaces paths-ignore: put your own prose-only negations above the re-includes of .claude/** and docs/adr/**. The kit sub-block now runs review-trigger-fixture.py and review-assert-fixture.sh against your filled .github/workflows/claude-review.yml, so a target that syncs the rules before merging this workflow is red until it does.
    • The CI workflow blueprint generated from its skeleton (usually .github/workflows/ci.yml). It was never a kit copy, so there is nothing to merge: edit it to match the skeleton. That means defaults: run: shell: bash, runs-on: ubuntu-24.04, and every uses: pinned to a full commit SHA with a trailing version comment, in place of a tag such as actions/checkout@v4.
    • .mcp.json, your copy of .mcp.json.example. Keep the servers your axes use. The kit ships no GitHub entry, hosted entries carry "type": "http", and time runs pinned through uvx. Turn every literal credential into a ${VAR} reference, export the variable before launching claude, and commit the file. The kit sub-block now checks the committed file's shape: a url entry without a type, an unpinned npx, uvx or bunx server, or a literal credential turns it red.
    • .claude/rules/orchestration.md and .claude/rules/orchestration-reference.md. Keep your posture row and your dated applied instances; take the kit's text for its own kit-shipped observations, whose cost figures are re-derived. Re-point any citation of § The three surfaces + resolution order to § The dispatch surfaces + resolution order.
    • .claude/rules/tooling.md. Keep your filled stack sections; take § MCP configuration's merge rule and § Automated review on the git host.
    • .claude/settings.json. Keep your own hooks' registrations, rewritten in exec form, and re-add their names and tiers to _comment_hooks. Then diff the hook event keys against your pre-merge copy (§ Syncing the kit).
    • .claude/rules/cbk-conventions.md and .claude/rules/cbk-conventions-reference.md. Keep your filled values and your project sub-block; take everything in the block above it. In cbk-conventions.md, take § Branch naming's rule that any issue a PR closes, cascade or not, puts its key in the branch (the bare issue number on the github-issues axis; <type>/<short-slug> only for work no issue tracks) with its Quick reference row, and the reworded substring trap in the CI-skip rule's section, which names all five skip tokens and GitHub's skip trailer and cites GitHub's page. The kit sub-block pins both, so a copy that keeps the old wording is red.
    • .claude/rules/pr-review.md and .claude/rules/pr-review-reference.md. Keep your roster entries; take the floor, the sweep and the invariants.
    • .claude/rules/simplification.md and .claude/rules/workflows.md. Take the kit's text; keep a filled § Cost+scope-explicit.
    • .claude/rules/knowledge-backend.md. It is now two halves. On the notion axis, take both, knowledge-backend.md and .claude/rules/knowledge-backend-reference.md, and move each filled value under its heading. On the none axis, where the rule was deleted, add neither half.
    • .claude/commands/finish-procedure.md and .claude/skills/rough-in/references/finish-procedure.md. Their citations now point at the conventions: a copy whose CONTRIBUTING.md citation was filled by hand takes the kit's text.
    • .claude/skills/adr-new/SKILL.md. Take the kit's text: it reads the index's rows in both forms, and its ## Refines vs Supersedes vs Extends section is now ## Relation grains, a pointer to .claude/rules/cbk-conventions-reference.md § ADR relation grains. Repoint any citation of "adr-new § Refines vs Supersedes" in your filled rules, reviewers or docs at adr-new § Relation grains or § ADR relation grains; an ADR that cites it stays as written, and the correction goes in docs/adr/corrections.md.
    • .claude/hooks/format-on-edit.sh and .claude/hooks/analyze-on-edit.sh. Keep your filled case arms; take the exec-form Register: stanza and the exit 2 on a formatter failure or an analyzer error (exit 0 hid it in the debug log), and in format-on-edit.sh the skip floor, which now names .claude/workflows/ (an analyzer only reads, so analyze-on-edit.sh has no floor).
    • .github/workflows/adr-immutability-check.yml. Re-copy it from the kit, then restore only your comments and the pinned job name: the job body is replaced. It reads a :(glob) pathspec, diffs three-dot from the merge base with --no-renames, checks out blobless, runs under shell: bash on a named runner image, and fails closed on any git error.
  • Files a target may already run ahead of the kit. These carry this release's harness and guard ports: .claude/workflows/agent-cost.py, .claude/workflows/tests/agent-cost-fixture.sh, .claude/workflows/finish-ab/finish-ab.js, .claude/workflows/tests/finish-ab-shape.mjs, .claude/workflows/finish-ab/run-arms-headless.py, .claude/workflows/tests/run-arms-headless-fixture.sh, .claude/workflows/review-sweep.js, .claude/workflows/tests/review-sweep-accounting.mjs, .claude/hooks/protect-immutable-adrs.sh and .claude/hooks/lib/resolve-path.sh. For each one your target already carries, diff your copy against v1.0.0's. Where the kit's application differs from yours, take the kit's, and keep only your project's own fills. From this sync on, v1.0.0 is these files' merge base, not your install (§ Syncing the kit).
  • New files. Add .claude/hooks/lib/resolve-path.sh, which is sourced, never run. A .gitignore with an unanchored lib/ line hides it from git add without a word: add !/.claude/hooks/lib/ after that line, and after the sync commits check that git ls-files .claude/hooks/lib/ names the helper (if it does not, git check-ignore -v .claude/hooks/lib/resolve-path.sh names the line that hides it). Add the new fixtures under .claude/workflows/tests/, each an add row: the three hook fixtures protected-paths-hook-fixture.sh, hook-guards-fixture.sh and hook-payloads-fixture.sh; adr-ci-body-fixture.sh; review-assert-fixture.sh and review-trigger-fixture.py; run-verification-block-fixture.sh; extract-run-block.sh, which the ADR and review-assert fixtures source; and, unless your target already runs them ahead (above), .claude/workflows/finish-ab/run-arms-headless.py and run-arms-headless-fixture.sh, which are required, because the block runs the fixture and the fixture fails without the runner. The block runs every one of them but the sourced helper. Append the harness block from scaffold's references/github-starter-templates.md § .gitignore to your .gitignore, below every stack section, and state its pin assertions in the commit body; the kit's own .gitignore is not in the drop-in set.
  • Kit-owned code stays out of your formatters. Exclude .claude/workflows/** from every repo-wide formatter and linter, forced for explicit paths (§ Syncing the kit). A wired format-on-edit.sh merges the new skip-floor line.
  • mise runs tasks under bash with pipefail. A target on mise adds the [task_config] block from blueprint's templates/tooling.md step 1 to its mise.toml and pins mise ≥ 2026.7.15 wherever tasks run, CI's setup action included. Nothing in the kit checks a target's mise.toml.
  • Dependabot covers more than the settle-window said. rust-toolchain.toml is covered by Dependabot's rust-toolchain ecosystem, and so is a container base-image tag; a target whose filled § Dependency settle-window lists either as uncovered corrects that section, which now names the three image gaps.
  • The cost reader keys by model version. A target synced at v0.5.0 prices legacy Fable 5 cache reads at Fable 5.1's 0.025x until it takes v1.0.0's agent-cost.py, which reads them at the standard 0.1x. A target that runs run-arms-headless.py sets its config's worktree_setup to its own per-worktree step, and adds any write-capable MCP server it wires beyond the kit's github, linear and notion to WRITE_SERVERS, which makes the file a merge row in its sync table.
  • A filled target owes both sentinels. .claude/workflows/tests/run-verification-block.sh is a copy row with a third rail: it fails a target whose output lacks verification: project sub-block complete. .claude/workflows/tests/run-verification-block-fixture.sh, which the block runs against the runner, is an add row. Wire the runner as a verification task your check depends on, per blueprint's templates/tooling.md. If you never stamped the bracketed manifest-and-lockfile globs in cbk-conventions-reference.md's paths:, stamp them now (the bootstrap checklist's disposition pass has the row), or the project sub-block is red. The block's CLAUDE.md check now fires in a target only once docs/cbk/blueprint.md exists: from then on your CLAUDE.md must mention cbk-conventions as a backticked path, never an @ import.
  • Fill the three hook backstop slots. A guard's fail-open warning names the backstop that still stands, and three of them are the project's to name: [the project's CI lockfile check — …] in protect-lock-files.sh, and [the base branch's ruleset — pull requests only — where one exists] in both warnings of protect-main-branch.sh. The bootstrap checklist's disposition pass gains a Hook backstop slots row: replace each bracket with the real check or ruleset, or with none and the reason. The project sub-block refuses a slot left bracketed. A filled slot makes those two hooks merge rows in your sync table.
  • The commands no longer start from a description. /finish, /intake, /enrich, /pr-respond and finish-procedure carry disable-model-invocation: true: type them.
  • Kit issue citations are qualified. Kit text cites its own issues as context-builder-kit#N, because a bare number links to the target's own issue. The rewrite touches about thirteen files a target copies byte for byte, among them the hooks, the workflow scripts and their tests, the executor pair and adr-new: take the kit's side of every such hunk. A copy you patched by hand takes the kit's form. The kit sub-block's bare-citation check runs on the kit's own tree only; in yours a bare number is your own issue.
  • Re-copy the issue templates. .github/ISSUE_TEMPLATE/cascade-meta.md and .github/ISSUE_TEMPLATE/cascade-rough-in.md are byte copies of scaffold's references/issue-templates/, and both changed: re-copy them. The project sub-block is red while your cascade-meta.md still cites the retired § Deferred meta-issues.
  • An unreadable payload now denies. A hard-deny guard refuses a payload jq cannot parse, and an ask-gate asks. A missing jq still fails open, naming its backstop.
  • The ADR guard judges what a path is, not how it is spelled. protect-immutable-adrs.sh sources lib/resolve-path.sh, resolves the path lexically and physically, and denies an Edit, Write or MultiEdit to an existing numbered ADR in any checkout, whatever spelling, relative path or symlink names it; its Not seen: line lists what it cannot see (a Bash write, a case-insensitive filesystem, a hand edit, a symlink swapped between check and write), which the ADR CI job refuses at the pull request. A target that customized its ADR hook takes the kit's hook (a copy row) plus the helper (an add row), carries anything it still needs as a named exception (§ Syncing the kit), and checks the helper is tracked (New files, above).
  • A dispatch after a cd is judged where the shell is. A hook payload's cwd follows the Bash tool's cd (the probe is recorded in require-repo-root-for-agents.sh's Timing: paragraph). So the launch-root guard now denies an Agent, Task or Workflow dispatch made after a cd into a subdirectory, and protect-main-branch.sh judges a commit against the checkout the shell is in. Return to the repository root, as its own command, before dispatching.
  • The review model is the action's pin. The templates name the family aliases, so a review runs whatever model the pinned anthropics/claude-code-action release's Claude Code resolves them to. opus resolves to Opus 5.5 on the Anthropic API from Claude Code v2.1.280, and sonnet to Sonnet 5.5 from v2.1.284 (https://code.claude.com/docs/en/model-config § Version history, read 2026-09-30). The action installs Claude Code 2.1.280 from its v1.0.232 and 2.1.284 from its v1.0.236 (src/entrypoints/run.ts, claudeCodeVersion, read 2026-09-30). A workflow pinned below v1.0.232 still reviews on the previous model under opus (Opus 5 from Claude Code v2.1.219), and below v1.0.236 on Sonnet 5 under sonnet: bump the pinned SHA.
  • A guard's fail-open warning is now shown. Every guard prints its fail-open warning as a systemMessage on stdout as well as on stderr, which on exit 0 reaches only the debug log; the Stop hook gathers its notes the same way. The guards are copy rows: take the kit's. A hook of your own follows § Hook authoring's fail-open bullet. Beside it, the main-branch guard allows the root commit of an unborn branch and reads a command continued with a backslash-newline, and the lock-file guard names npm-shrinkwrap.json and bun.lockb.
  • A blueprint written before this release has no § Workstreams table. /finish, /intake and framing look a workstream slug up there. blueprint.md is immutable, so append the table — workstream, slug, layer — under its append-only ## Amendments, from the slugs its slug gate confirmed. On the in-repo-markdown axis, setup progress belongs on the roadmap's step-0 row, never in commits to blueprint.md.
  • A Linear target relabels its workstreams. Scaffold used to create one area:<slug> label per workstream; every Linear flow applies workstream:<slug>. Rename them, and create the enhancement, source:<name>, triage and appetite:small / medium / big labels the flows apply.
  • The experiment runners refuse before they spend. run-arms-headless.py refuses a config with an unknown key, a budget or continuation count out of range, or a dry run combined with resume, and records each arm in executed.json as {anon, cell, result}. finish-ab.js refuses an unknown argument or effort and an executed record whose cell does not match, and drops a judge whose scores do not cover every arm. review-sweep.js refuses malformed arguments: a list that is not a list, a bound that is not a whole number, a verifyModel other than opus, sonnet or haiku. A saved config or call that relied on the old leniency fails loudly; fix it rather than the check.
  • Cost figures read before this release overstate. agent-cost.py summed every transcript line, and Claude Code writes one response as several; it billed a fork again for the parent history its transcript opens with; and it priced every cache write at the 5-minute rate, where a 1-hour write bills at 2x. The kit's own dated figures were re-derived from their original transcripts: they ran 1.5x to about 2.2x high, all from the per-line count (orchestration-reference.md § Applied instances). Re-run a figure of your own on its original transcripts before comparing it across releases. From Claude Code 2.1.278 a subagent's responses are mostly recorded mid-stream, so a figure read from a current transcript is a floor on output and cost; the reader names the rows it applies to.
  • The ## Review gate block has a "not run" form. An issue-less branch writes its /simplify and toolkit lines not run — <reason>; a cascade PR's docs-only diff still runs the floor (pr-review.md § What NOT to flag).

🤖 Generated with Claude Code

j4th and others added 30 commits September 30, 2026 08:57
The design for landing the sixteen open issues (#58's residue, #60–#74)
and the whole-kit review as the kit's first tagged release: decisions
D41–D62, ten clusters in commit order, six execution-time probes, and a
371-row trace of every ask by issue and comment id.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A master plan (constraints, review focus, ownership map, baseline, and
the closing tasks F1-F5: budget and sanitization gate, the review floor
once, the trace audit, the draft PR, and the post-merge tags and
release) and ten cluster plans, V1-V10, 94 tasks in commit order. Every
ported hunk is re-authored inline; every trace row has a home; each
cluster file was dry-run on a scratch copy of the kit and reconciled
against its neighbours.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the block runs

A filled target (docs/cbk/scaffold.md exists) now also owes `verification:
project sub-block complete`: the done sentinel prints whether or not the
project sub-block ran, so a guard that stopped firing turned every project
check off with the gate green. run-verification-block-fixture.sh drives the
three rails in six cases on synthetic blocks, and the kit sub-block runs it.
§ Verification › Run it says three rails, names the residual (a deleted or
renamed scaffold.md), and cites context-builder-kit#58 instead of a bare #58.

Trace: #73/body/1, critic/22 (D51).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… gate

detect-forked-agent-memory.sh fails open by design (exit 0 with a WARNING)
when it could not look: no checkout, an unenterable root, a partial scan, an
unmonitored walk. The block's check read only the exit code, so it passed on
nothing while the hook named the block as its backstop. stop_hook_clean reads
the hook's stderr; an in-block negative case, outside any checkout under a git
ceiling, must itself come back red, so a regression is caught where the check
runs. The preamble states the rule.

Trace: #73/body/2, the block's half (D51). The hook's WARNING contract and
the partial-scan fixture case are V2's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…oaded bytes

The always-loaded loop printed the standing cost but nothing flagged growth.
Above 140,000 bytes it now prints a WARN line and the run goes on: the number
is a signal to path-scope, split or delete a rule, not a gate. Shown by
padding an always-loaded rule to 140,630 bytes in a throwaway clone: no WARN
before, one WARN after, exit 0 both times.

Decision: D50 (no trace row of its own).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cation block

Two causes kept a fresh scaffold red. The disposition pass never asked about
cbk-conventions-reference.md's bracketed manifest-and-lockfile glob, which the
project sub-block's stamped-globs check rejects: the table gains the row, § 4's
first sentence counts three path-scoped placeholders, and the phase-exit item
and Test 1 name it. And the kit sub-block required a CLAUDE.md mention that
blueprint, not scaffold, writes: a scaffolded target now owes it once
docs/cbk/blueprint.md exists, and the kit tree always does. A kit-sub-block pin
fails when a rule ships a paths: placeholder the disposition table has no row
for. Shown on a simulated scaffold (drop-in set, ADR starters, scaffold.md,
globs stamped): all three sentinels; unstamped, red naming the file.

Trace: Review Focus 4; the Gate settled call (disposition row); the defect
behind review/portability/68 (V9's row, landed here).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…check

Nothing in the templates ran the block, so a target ran it only if someone
remembered. Blueprint's tooling template: the minimum task set, light mode
included, gains a verification task that check depends on, whose body is the
runner, with the tools it needs (bash, git, awk, mktemp; jq, node, python3
for the block's fixtures); § Rules makes it a leg of check, run in CI through
check and once before the HITL presentation. The phase-exit checklist and Test
1 assert it. Scaffold's verification matrix runs the runner after the
disposition pass and expects both sentinels. Run it names the two homes, and a
kit-sub-block pin keeps both templates naming the runner.

Trace: #58/c5901493591/R4 (D51).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cker, scoped to one file

An absent check brackets one letter of the phrase it hunts so the grep never
matches its own line; a target that spellchecks .claude/ reads each fragment
as a typo. § Verification now says to exempt the idiom in this one file, never
repo-wide, and shows the typos form: a [type.<name>] table whose extend-glob
names cbk-conventions-reference.md and whose extend-ignore-re matches the
idiom, quoted from typos' reference (read 2026-09-30) and exercised against
typos 1.50.3. A kit-sub-block pin keeps the note.

Trace: #58/c5901493591/R5 (D51).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The kit's own verify.yml ran with GitHub's unset shell (bash -e, no pipefail)
on ubuntu-latest, a label that moves under the gate; its workflow templates
take the same two pins later in this PR. verify.yml gains
defaults.run.shell: bash and runs-on: ubuntu-24.04, each with its source read
2026-09-30 (the workflow-syntax shell table; the runner-images migration note
and the 24.04 image readme). A kit-tree-only
check in the block pins both lines. The job name, the required check, is
unchanged.

Decision: D52, the kit's-own-workflows settled call (no trace row of its own).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ails closed

The job's body is a :(glob) pathspec with --diff-filter=a and --no-renames, and
any output fails it: the awk match it replaces read a quoted non-ASCII name and a
spaced name as "no ADR changed". The diff is three-dot, so an ADR merged to the
base after the branch point no longer reads as a deletion. A git error now fails
the step, where mapfile of a process substitution read it as a pass. The checkout
is blobless; the job runs under shell: bash on a named runner image.
adr-ci-body-fixture.sh runs the job's own body, extracted by the shared
extract-run-block.sh; the block runs the fixture and pins what it cannot see.

Red first: FAIL: modify a non-ASCII-named ADR (want exit 1, got 0)

Trace: #60/c5876324022/adr-job, #60/c5881158391/port-glob-body,
#60/c5892401033/adr-job-glob-fixture, #60/c5901494943/3, #61/body/fix,
#61/body/grep, review/security/19; the extractor for #67/c5892401572/2.
Decisions: D55, D44.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…helper

protect-immutable-adrs.sh matched the payload's path text against a prefix, so
every spelling context-builder-kit#60 measured passed: a .. segment, . and //, a
trailing-slash or foreign project dir, a relative path, a symlinked directory, a
symlink to the file, a hardlink, /proc/self/root and an unparseable payload. The
guard now sources .claude/hooks/lib/resolve-path.sh: the payload read exactly in
one jq, the path resolved lexically and physically, /proc and unreadable symlinks
refused, the same file under another name caught by bash's -ef over the hook's own
checkout, the project dir and a linked worktree. The deny is path-independent: an
existing docs/adr/NNNN-*.md in any checkout; the existence test runs on the
resolved path, so a new ADR stays creatable. A missing helper or jq fails open
naming the ADR job. hook-contract-fixture.sh's early-reader check reads lib/*.sh;
section Hook authoring states the sourced-helper contract and the lib/ gitignore
trap, and the block asserts the helper is not ignored.

Red first: FAIL [protect-immutable-adrs.sh]: a .. segment (docs/sub/../adr/) (want exit 2, got 0)

Trace: #60/body/fix, #60/body/alt, #60/body/fixture, #60/c5876324022/table,
#60/c5876324022/closed-1, closed-2, closed-3, closed-4,
#60/c5876324022/closed-5, closed-6, #60/c5876324022/closed-note,
#60/c5881158391/port-helper, #60/c5892401033/helper,
#60/c5892401033/adr-path-independent, #60/c5892401033/linked-worktrees,
#60/c5892401033/ef-inode, #60/c5892401033/gitignore-trap,
#60/c5892401033/fixture, #60/c5901494943/2a, 2b, 2c, 2d, 2e,
#62/body/extra-adr-hook,
#58/c5901493591/R1 (the path guard), release/5 (section Hook authoring),
#69/c5859756889/apply-h4/1 (a target's own hooks, held to the same contract),
review/consistency/48 (the tally, rewritten count-free).
Decisions: D43, D44, D61.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… refused or asked

The lock-file, main-branch and PR-state guards failed open on a payload jq cannot
parse (one lone UTF-16 surrogate escape); the two denies now refuse it and the
ask-gate asks, and section Hook authoring says why an unreadable payload is not an
environment defect. Each fail-open warning names what still stands: a bracketed
slot for the project's CI lockfile check and for the base branch's ruleset, filled
at scaffold's rule-file disposition pass (a new bootstrap row) and refused unfilled
by the project sub-block; the PR-state gate says no mechanical backstop exists.
The lock-file fallback's remediation names the real route, a case arm above the
*.lock) arm. The knowledge-backend ask-gate's header carries its ten verbs, and
protect-main-branch.sh gains the Depends: line every header carries.
hook-guards-fixture.sh drives every branch of the four guards, finds the ask-gate
through the registry, and runs one deny per guard from a root containing a space.

Red first: FAIL: the non-checkout warning names a backstop (stderr lacks 'Backstop')

Trace: #62/body/lock-files, #62/body/pr-state, #62/body/main-branch,
#58/c5901493591/R8, #58/c5901493591/R1 (the four guards), critic/8,
#66/body/2 (the ask-gate header), release/5 (protect-lock-files.sh).
Decisions: D61, D44.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…racters

The left anchor admitted only a start, a space or ;&|, so a commit on main passed
the hard-deny inside bash -c "…", sh -c '…', a subshell, $(…), backticks or through
/usr/bin/git, and the right anchor let git commit; and git commit&&… through; the
PR-state ask-gate shared the left anchor. Both anchors are now negated classes:
anything that cannot continue a word before, anything but a word character, . or
- after, so git commit-tree and git config commit.gpgsign stay allowed. Each
header's Not seen: line names what a pattern guard cannot see (eval, a variable, an
alias). The remediation's comments say "an issue this PR closes" and "work no issue
tracks". The fixture gains every spelling, both ways.

Red first: FAIL: a commit on main is denied: bash -c "git commit -m x" (want exit 2, got 0)

Trace: review/security/20, #68/body/1a (the guard's remediation).
Decision: D44.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t hooks, by payload

require-repo-root-for-agents.sh failed open on a payload jq cannot parse; it now
refuses it. hook-payloads-fixture.sh drives every branch the launch-root guard and
the forked-memory detector document, a root containing a space included, and pins
the literal WARNING on each could-not-look branch, the partial scan among them
(context-builder-kit#58 items 1, 2 and 4). Probe P1 settled whether a hook
payload's cwd follows the Bash tool's cd on the release CLI; the guard's Timing
paragraph, the main-branch guard's cwd comment and section Hook authoring state
the answer, dated. The detector's rule 1 now states what it guarantees and points
at its Residual; its bare citations are qualified.

Red first: FAIL: a payload jq cannot parse is refused (want exit 2, got 0)

Trace: #58/c5901493591/R1 (the two hooks), #58/c5901493591/R6,
#58/c5901493591/R13, #73/body/2 (the WARNING pins), release/5 (the two hooks).
Decisions: D61, D44. Probe: P1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A CI job named as a hook's backstop but never landed is how a target's frozen
corpus went two weeks unguarded (context-builder-kit#60, suggestion 1). The kit
sub-block now reads the mutation table, the hook registry and every hook header,
and fails on a .github/workflows or scripts path that does not exist; the checker
is asked about bogus names first, and a mutation section that names no checkable
backstop is red. The table's ADR row names its workflow by path. The project
sub-block resolves `mise run <task>` names against mise.toml where one exists.

Red first: cbk-conventions.md § Mutation discipline names no CI backstop in a checkable form (.github/workflows/<file>.yml)

Trace: #60/c5876324022/corpus-s1, #60/c5881158391/port-s1,
#60/c5892401033/suggestion-1-built.
Decision: D51.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… checklist row per item

Consultation's enforcement set told a target to extend the ADR job, which then
inherited its awk path match. The recipe now builds the corpus guard on
lib/resolve-path.sh's root-scoped mode and the corpus CI job on the ADR job's
parse-nothing body, and lists the nine closures two targets' red-team rounds
measured, ending with the additions call each target makes. The bootstrap
checklist gains one verification row per enforcement item, so no item is ticked
off with the others; brownfield audit and consultation's Test 4 say the same.
The block asserts the closures and the six rows.

Red first: frozen_corpus_ingestion.md lacks the CI job's closure list or the parse-nothing body

Trace: #60/c5901494943/1, #60/c5876324022/corpus-s2,
#60/c5876324022/echosphere-z, #60/c5892401033/corpus-leg,
#60/c5892401033/leg-defects, critic/6, critic/7.
Decision: D55.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d the three-edit wiring

Verify by payload said the payload-reachable branches were asserted in
hook-contract-fixture.sh, which drove none of the ADR guard, the knowledge-backend
ask-gate or the main-branch allow cases. It now names the fixture of each hook
family, all run by the block, and what a target does when it adds a hook. The
advisory-hook bullet says that wiring one is three edits: the stanza, the
ADVISORY_WIRED name and the two-views paragraph, each checked by the block. The
block asserts both.

Red first: § Hook authoring › Verify by payload does not name hook-guards-fixture.sh

Trace: #58/c5901493591/R1 (section Verify by payload), #58/c5901493591/R11
(section Hook authoring; the exemplars' Register: headers are V8's).
Decision: D44.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ke-only

`$1` is the second argument (code.claude.com/docs/en/skills § Available
string substitutions, read 2026-09-30), so `/finish 42` left `$1` literal and
`/finish 42 --skip-review` rendered `#--skip-review`. Every command and both
bundled executor copies now declare `arguments:` and write `$issue`, `$pr` or
`$source`. The five commands carry `disable-model-invocation: true` (D57): they
open branches, issues and PRs. `/finish` item 4 reads `--skip-review` from
`$ARGUMENTS`, so pr-review.md's break-glass flag is honoured (Review Focus 5,
probed on 2.1.285). The block pins all three.

Trace: review/claude-code/9, review/claude-code/14.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every registration was shell form with an unquoted ${CLAUDE_PROJECT_DIR}.
Under a project path containing a space the shell split it, the script was
not found, and each PreToolUse guard's non-2 exit let the tool call through:
every hard-deny guard failed open with no signal (probed on 2.1.285). Each
registration now carries "args": [], the exec form the hooks page prescribes
for a hook that references a path placeholder (code.claude.com/docs/en/hooks
§ Exec form and shell form, read 2026-09-30). The block refuses a placeholder
registration without an args array (Review Focus 2). § Hook authoring now
states the form: the stdin bullet's registered call is the exec-form entry,
and the Project-relative paths bullet says to register in exec form and why
(V2 declined these two sentences and granted the region; they land here).

Trace: review/claude-code/10.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…erge rule is the documented one

enabledPlugins is keyed <plugin>@<marketplace> (code.claude.com/docs/en/
plugin-marketplaces, read 2026-09-30). The bare "pr-review-toolkit" key
enabled nothing, so the review floor's toolkit half was never declared; the
keys are now pr-review-toolkit@claude-plugins-official and
commit-commands@claude-plugins-official, and the comment gives the install
commands and points at simplification.md § Plugin instead of restating a CLI
version. tooling.md's always-loaded rule said a repeated list in local
settings "silently replaces" the committed one; the settings page says lists
merge, except four model-list keys, and a local false is the documented
per-machine plugin opt-out. Always-loaded: +209 bytes.

Trace: review/claude-code/13, review/release/29, review/claude-code/16,
#69/body/F9 (settings.json part).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
_comment_hooks restated the tier definitions and per-hook facts each hook
header and § Hook authoring already carry (4,080 characters). It is now the
registry: each hook under its tier with its event and matcher, the exec form,
pointers, and the facts other clusters handed to this file — the ADR guard
resolves the path both ways through .claude/hooks/lib/resolve-path.sh, its CI
backstop is .github/workflows/adr-immutability-check.yml, and an unparseable
payload makes a hard-deny guard refuse (D61). The block caps it at 2,500
characters and requires it to name every sourced helper.

enabledMcpjsonServers drops github (D58: gh is the kit's GitHub interface),
its comment no longer tells a reader to comment out JSON, and the block
checks that every listed server is declared in .mcp.json.

Trace: #66/body/2 (settings.json part), #60/body/fix (registry sentence),
#60/c5892401033/helper (registry names the helper),
#60/c5876324022/corpus-s1 (registry names the ADR workflow),
#70/body/6a (enabledMcpjsonServers).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… an override

A project agent named Explore replaces the built-in (code.claude.com/docs/en/
sub-agents, read 2026-09-30). The kit's override kept searches on the
smallest tier but, unlike the built-in, loaded the whole CLAUDE.md hierarchy
and every unscoped rule on each dispatch. It now sets omitClaudeMd: true
(v2.1.271+), its description is the routing sentence (the provenance moved to
a maintainer paragraph in the body), and the bootstrap checklist asks whether
to keep it, so an adopting project knows it ships a live override.

Trace: review/claude-code/15, review/release/63 (agent and checklist parts).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s the rubric's reference half

/intake reads an outside reporter's text and /pr-respond any commenter's, yet
neither stated the rule the research phases carry: fetched text is data, and
an imperative inside it is surfaced, never acted on. Both now do. /intake
writes its own reproduction, names it at the Step 3 gate, and never runs a
command quoted in the report. /pr-respond applies a finding only from the PR's
author, a collaborator with write access or the project's review bot; any
other author's finding is Surface at most. /pr-respond also reads
pr-review-reference.md's calibration, path-conditional aggressiveness and
anti-patterns explicitly: a triage is not a file read, so the path-scoped half
never loaded (code.claude.com/docs/en/memory, read 2026-09-30).

Trace: review/security/21, review/claude-code/52.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ted, /simplify's dimensions are sourced

The cascade-rule reviewer cited a knowledge-backend.md section that does not
exist; it now cites § The code-adjacent split — canonical, and the block
checks that every § citation in an agent resolves. The logging reviewer
flagged per-tick debug calls that logging.md § Level taxonomy permits and
contradicted its own level-taxonomy item; it now flags above debug and names
telemetry as the better form. simplification.md asserted five behaviours of
the bundled skill with no source; it now quotes the commands page's four
dimensions and states the rest as the project's own rules. The rule index
row no longer claims a triage the file does not contain. Three dangling
docs/STANDARDS.md section citations are gone. Always-loaded: −149 bytes.

Trace: review/claude-code/58, review/claude-code/59, review/claude-code/60,
review/consistency/40 (simplification.md and logging reviewer parts).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…trings in the emitted workflows

Every workflow the kit emits (both review templates, blueprint's CI skeleton, scaffold's CI stub) sets
`defaults: run: shell: bash`, so each run: step has pipefail, and runs on ubuntu-24.04 rather than the
moving -latest label. The review templates' label checks read a here-string instead of piping into an
early-exiting grep -q, and the label-removal step captures its label read, so a failed read is an
error and a label deleted from the repository is the sourced benign branch. The CI skeleton's two
`uses:` lines take the fill-slot form @[full-commit-sha] # v[version], as the review templates pin
theirs, instead of a mutable tag. The kit sub-block carries the three assertions, each shown red by a
mutant.

Trace: #70/body/1a, #70/body/1b, #70/body/1c, #70/body/2, #70/body/3, #67/c5901495773/8, critic/1, critic/2,
review/claude-code/56

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…urrency and a fork guard

The review trigger is a `paths` filter whose re-includes (.claude/**, docs/adr/**) come after the
markdown negation, so rules, skills, commands, agents and ADRs are reviewed while docs-only PRs are
not; review-trigger-fixture.py evaluates the filter as GitHub does over twelve PR shapes and runs in
the kit sub-block, red first on the old paths-ignore. `concurrency` moves to the job in both
templates, so a run the job's if: skips cannot cancel a live review or response, and a fork PR, which
gets no secrets, is skipped rather than failed. The prompt's skip clause and blueprint's review-bot
paragraph name the same shape.

Trace: #67/c5901495773/1a, #67/c5901495773/1b, #67/c5901495773/1c, review/security/18, review/security/22

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the verdict and says why

"Assert the review posted" fetches the PR's comment pages once and counts them with jq -s (gh api
--paginate --jq ran per page, so a review past comment 30 read as missing); matches the review app's
comment by its verdict marker on updated_at, so the tracker alone never passes and a tracker ending
in the verdict does; runs on !cancelled(); and, when no summary landed, says why: the session ran
without a verdict, or no session started, told apart by a diff of the workflow against the PR base
(a validation skip) or not (auth or setup). The Auto-review step gains `id: review`, Constraint 4 is
restamped with the exercised paths, and review-assert-fixture.sh runs the step's own body on
synthetic comment pages through V2's extract-run-block.sh, red first on the old step.

Trace: #67/c5877251222/1, #67/c5877251222/2, #67/c5877251222/3, #67/c5877251222/4,
#67/c5881158070/1 (corrected by c5892401572/1), #67/c5892401572/2, #67/c5892401572/3,
#67/c5901495773/confirm-2, critic/12, critic/13

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rded model

Both review templates name the model by family alias (opus by default and on the deep path, sonnet on
the fast path and as the fallback; never best), and every branch passes --effort, claude.yml at high,
because the default effort is per model and Opus 5.5 and Sonnet 5.5 default to medium. The header says
the action's pinned SHA is the model pin, never to set ANTHROPIC_MODEL (it wins over --model), and what
the action-bump PR re-checks; the deep branch says what choosing fable costs and why xhigh is named now
that ultracode sets no effort. A "Record the resolved model" step on both templates (continue-on-error)
writes the model the session started on, its Claude Code version and every model that answered to the
job summary, and the prompt's model line says it shows the alias. The fast branch's effort=high is
labelled a pin pending a Sonnet 5.5 effort sweep, since Sonnet 5.5's levels are recalibrated. The
fixture runs the record step on synthetic execution files, red first. Probe P3 recorded.

Trace: #67/body/1a, #67/body/1b, #67/body/1c, #67/body/1d, #67/body/2a, #67/body/2b, #67/body/2c,
#67/c5804984992/1, #67/c5881158070/2a, #67/c5881158070/2b + #67/c5892401572/Effort,
#67/c5901495773/confirm-1, #67/c5901495773/5 (template half), #69/c5859756889/3-C4,
#69/c5881157875/3a (template half), critic/18 (the fast branch)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e base branch's

The review prompt says the action restored .claude/, .mcp.json and the root CLAUDE.md from the base
branch and kept the PR's copies under .claude-pr/, that a nested CLAUDE.md (and, where a project has
one, a symlinked AGENTS.md) reads at the PR head, that a rule the PR changes is reviewed as the PR
states it with the version named, and that CI's check job, read with gh pr checks, is what runs the
PR's own hooks and fixtures. Both headers state the mechanism with its source. The summary reads CI
and reports only commands actually run; a note above claude_args pairs every prompt gate with an
--allowedTools entry and sizes the degrade clause's N (40) against --max-turns 60, re-measured on
Claude Code 2.1.286. The two kit-issue citations are qualified.

Trace: #68/body/2a (+c5861166655/2, c5881158244/1, c5892402074/2), #68/body/2b, #68/body/2c,
#68/body/2d, #68/c5881158244/2 (+c5892402074/2), #68/c5901496131, #67/c5901495773/2,
#67/c5901495773/3, #58/c5901493591/R12 = #67/c5901495773/4, #67/c5901495773/7, critic/14, critic/15,
release/5 (the two template citations)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…g.md § Automated review

The review template's Constraint 1 says the workflow cannot review any PR that changes it, quoting
the action's own log line for the skip, and tooling.md § Automated review says the same, names the
assert step that goes red on such a PR, and adds that the pinned action fixes the model and the
Claude Code the bot runs, so a harness fact the bot asserts is checked against the operator's version.
Constraints 3 and 6 are sourced to the action at the pinned SHA, and the track_progress rail is dated:
labeled is accepted from v1.0.188. The block's review-automation check and a tooling.md check follow
the corrected wording. Always-loaded +452 bytes (tooling.md).

Trace: #67/c5901495773/6, #67/c5901495773/9, review/portability/35

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
j4th and others added 25 commits October 1, 2026 19:21
…e job's if:, before the concurrency group

The bot and label checks sat in a step, which runs only after the job has joined the job-level
concurrency group. A bot comment that mentions @claude passed the job `if:`, joined the group,
and with cancel-in-progress cancelled the reply in flight before the step could skip it. The
gates now sit in the job `if:`, as claude-review.yml's already do. `sender.type == 'User'`
mirrors the action's own gate, which refuses any non-User actor unless `allowed_bots` names it
(action.yml default "", checkHumanActor in src/github/validation/actor.ts, read 2026-10-01), so a
bot is skipped green instead of failing red. review-assert-fixture.sh pins the gates in the
job `if:` and refuses a step-level gate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s tool list allows

The prompt all three label paths share forbade subagents ("Do the review in this one session",
"inline yourself (one session, no subagents)"), while the deep path drops `--disallowedTools
Agent` and turns on ultracode precisely so it can fan out. The label step now writes a
`dispatch_rule` beside the tool list: the default and fast paths say do not dispatch, and the
deep path says it may, with every subagent returning before the summary posts. The prompt
interpolates that rule in both places. review-assert-fixture.sh runs the label step on five
label sets and checks the precedence (deep over fast over default, `claude-deep-review-wip` no
deep label), that each path's rule agrees with its tool list, and that the prompt states no
dispatch rule of its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…old skip line, and names a moved base

"Assert the review posted" counted any comment whose body contained APPROVE, REQUEST CHANGES or
NEEDS DISCUSSION, so a tracker note naming a verdict word in passing read as a landed review.
The marker is now a bold verdict, as the prompt prescribes; a verdict after a label on its line
still lands. The docs-only skip line the prompt asks for carried no verdict and so turned a
correct skip red, with a "no summary posted" notice: it now opens with a bold `**SKIPPED**`,
which the assertion accepts. The fixture checks every bold word on the prompt's verdict line and
the skip line against the step, so the two vocabularies cannot drift apart.

The validation-skip notice blamed "this PR" for changing the workflow. The diff is tree against
tree, which is right: the action compares contents, so a base that changed the file after the
PR branched skips the review too. The notice now says the copies differ, and a fixture case
pins the moved base, so a merge-base diff would go red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rmed entry cannot pass a token by

The block runs without errexit, and the credential check piped jq into `sort -u`, so a jq error
— `"headers": "Authorization: Bearer ghp_…"`, a string where an object belongs, gives "string
has no keys" — left `bad` empty and the literal token passed (reproduced on a scratch clone:
exit 0). The two checks beside it had the same shape without the pipe. Each jq now runs alone,
a failure stops the block with the reason, and the sort runs afterwards. On the clone the string
header now fails, and so does the same token written as an object.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the kit tree

An agent's `§` citation to a rule file that does not exist was skipped everywhere, which is
right only on a filled target, where the disposition pass may have deleted the rule. On the
kit tree every rule exists, so `cbk-convention.md § ADR index sync` passed silently
(reproduced on a scratch clone). It now fails there with the file it names, and is still
skipped on a target with docs/cbk/scaffold.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd the sweep's bounds are pinned by behaviour only

The nested-file case passed with or without `:(glob)`, since the file was an addition and
additions are allowed. The fixture's base now holds docs/adr/0009-notes/x.md, and a case edits
it: only the `:(glob)` pathspec leaves it out, because a plain pathspec's `*` matches across
`/`. With the magic removed the case goes red.

The block grepped review-sweep.js for `maxPerDimension ?? 3`, `maxVerify ?? 8`, FIND_EFFORT
and RETRY_EFFORT. The accounting harness already asserts the 3/8 default bounds and the
finder efforts as behaviour (scenarios 14, 23 and the default-bounds scenario). A
behaviour-preserving refactor broke the greps, and `??` turned to `||` passed them, so they
are dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…two that overclaimed say what is true

orchestration-reference.md's preamble said orchestration.md keeps a pointer heading for every
section there, and workflows.md's rule index repeated it for all four halves. Two sections
added since the split, § Generation notes — the sources and § Cost terms and run hygiene, have
none. The V9.19 check looped over only the conventions and review halves, so nothing caught
it. The loop now covers all four halves. It goes red on the orchestration preamble, which now
names the two later sections, and the index row says the pointer is kept for every section
moved at the split. On a target a deleted pair is skipped; the kit tree must hold all four.

knowledge-backend-reference.md said its sections "were moved verbatim on 2026-09-30", but the
same change re-sourced four of them: it now says "moved … and re-sourced in the same change".

Always-loaded total 132,473 → 132,494 bytes (workflows.md, +21).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the [1m] variant at list

Every cache write was priced at the 5-minute 1.25x. Claude Code records the per-TTL split in
`usage.cache_creation`, and a session on the 1-hour TTL writes all of its cache there, which the
pricing page bills at "2x base input price" (read 2026-10-01). Writes are now split, the
5-minute part at 1.25x and the 1-hour part at 2x; a write with no breakdown stays at 1.25x.

tier() accepted any bracketed suffix, so an id whose variant the table does not know priced at
list. It now accepts only `[1m]`, which the same page says bills at standard rates ("Claude 4.6
and later models ... include the full 1M token context window at standard pricing"); any other
variant is unpriced and named. The docstring's "each assistant message carries usage" now says
what the dedup reads: lines that repeat a response's message id, the last line standing for it.

agent-cost-fixture.sh adds sss (a TTL-split write, $13.00) and ttt (`[2m]`, unpriced), both red
first: $88.15 over 13 of 20 priced agents. orchestration-reference.md's note on past figures adds
the 1.25x write rate, which pulled a 1-hour session's figure the other way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ual setup every citation names

More than twenty surfaces cite `blueprint.md` § Workstreams: /finish Step 1 validates a title's
slug against it, framing takes one row, and the conventions say the slug list lives there. The
blueprint SKILL describes "the workstreams table … slugs", and the template's own
planning-axis text names "the Workstreams table inside this file". The template never emitted
one. It now does, one row per core project with its confirmed slug and layer, plus authoring
guidance.

On the in-repo-markdown axis, seven surfaces put the handoff content in `blueprint.md` § Manual
setup, which the template also lacked. Two of them said setup progress is committed to
blueprint.md, which § Mutation discipline makes immutable after its commit. The template gains
the section for that axis only, written once. Progress moves to the roadmap's step-0 row, which
is freely mutable: planning-backend-commit.md, scaffold SKILL.md's disclosure and the roadmap
template now say so. Blueprint's test case 1 gains the missing-Workstreams failure.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the gate block has a "not run" form

pr-review.md § What NOT to flag and its reference half said a docs-only diff still runs the floor,
both skills invoked. cbk-conventions.md § Branch naming, whose first issue-less example is a docs
sweep, writes the floor "not run — docs-only sweep, no code changed", and the starter PR template
agrees. Both docs-only statements now say "on a cascade PR" and point an issue-less branch at
§ Branch naming. The `## Review gate` shapes in § The floor had no "not run" form, so the
/simplify and toolkit lines gain `or: not run — <reason> (an issue-less branch only)`.

Always-loaded total 132,494 → 132,731 bytes (pr-review.md, +237).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…flows apply

linear_planning.md created one `area:<slug>` label per workstream area, "mirroring the GitHub
label set". The GitHub set uses `workstream:<slug>`, and every Linear flow applies
`workstream:<slug>`. The list also lacked `enhancement` and `source:*` (intake) and `triage`
(the lifecycle axis), and framing's Linear call sets `appetite:small|medium|big`, which
neither taxonomy created. The taxonomy now lists the GitHub axes as the Linear flows use them:
the work type rides Linear's own Feature/Bug/Improvement label, and the review-control labels
stay on the repository. enrich.md's two `area:<slug>` mentions, and the label examples in
scaffold's SKILL.md and output template (which also lacked `chore`), use `workstream:`.

The V9.14 check grepped only the GitHub profile. It now checks the Linear taxonomy for every
label a Linear flow applies and refuses `area:` labels anywhere in scaffold or the commands.
It went red on the old text first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e fact, and the fact's record names it

blueprint's handoff-issue.md listed `rust-toolchain.toml` among pins no Dependabot ecosystem
maintains. The conventions quote Dependabot's support for it, and tooling.md 6b and the starter
templates agree. The line now names the real gaps (the toolchain manager's pins, a `COPY
--from` image, a dev container's image) and calls `rust-toolchain.toml` the covered exception.
The multi-surface record in § Dependency settle-window listed three surfaces and missed this
fourth, which is how the sweep missed it; it now names all four.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ish's --skip-review, and scopes CI to main

- The harvest-trace rule now audits a `landed` row against the branch head, not its commit
  message: the ask is located at a file:line and given a verdict, and a row that does not hold
  is reopened. This is the audit method a harvest's critic asked to be the gate (trace
  critic/9).
- `/finish` takes the issue number and an optional `--skip-review` (finish.md's argument-hint),
  not "one argument".
- verify.yml runs on pull requests to main and pushes to it, not "every pull request".
- scaffold SKILL.md's label-axis list names the lifecycle axis its GitHub profile creates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The review sweep left four rule-text claims unverified. One read-only verifier per claim
(workhorse tier, high effort) re-read the primary source on 2026-10-01:

- orchestration-reference.md, subagent model resolution: "all checked against the org's model
  allowlist (an excluded value is skipped and the agent runs on the inherited model)". The
  sub-agents page § Choose a model checks only the first three values. Since v2.1.222 a
  blocked family alias runs on "the newest version of that family the allowlist permits", and
  only other values fall back "on the inherited model instead".
- orchestration.md, the over-verification anti-pattern, quoted "verify your work" and "use a
  subagent to double-check". Neither string is in the Opus 5 or Opus 5.5 guide. The row now
  quotes the guide's own examples and names § Task scope and over-verification.
- cbk-conventions-reference.md § Hook authoring credited stdin delivery to § Exec form and
  shell form, which says only that the command is spawned with no shell. "Input arrives on
  stdin" is § Hook lifecycle; each fact now cites its own section.
- § Hook authoring, the fatal settings diagnostic: the fatal path still holds in the 2.1.287
  bundle, but the trigger is narrower than "any non-hooks key carrying matcher". It fires on a
  non-empty PreToolUse/PermissionRequest key, or a matcher object not under a hook-event key,
  up to three levels deep. Nine keys are exempt, and the file loses only its own settings. The
  dating extends to 2.1.287; "silently" is dropped, since the settings page documents a
  Settings Error dialog.

Always-loaded total 132,731 → 132,803 bytes (orchestration.md, +72).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and the block keeps it so

The axis was renamed `in-repo-markdown` when the three named profiles gave way to two
independent axes, but "markdown-only" survived as its name in four places in the conventions'
reference half, in intake.md and enrich.md, and in rough-in's test cases (trace
review/consistency/3). blueprint SKILL.md also said "every profile" and summarised "profile" at
inheritance. Each now names the axis; the adjective (a table that lives only in markdown) stays.
A block check refuses the old name for a backend, mode, project or planning axis in the skills,
commands and rules; it matches eight lines of the pre-fix text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… Code, and when a target's bump lands

The CHANGELOG and the review template cite `src/entrypoints/run.ts` (`claudeCodeVersion`) for
which claude-code-action release installs which Claude Code. orchestration-reference.md cited
`base-action/action.yml`. Both pin the same version, but the templates `uses:` the root action,
whose installer is run.ts, so the reference half now cites it too and names the other as a
second pin. Re-read at v1.0.231/232/235/236 on 2026-10-01: 2.1.278 → 2.1.280, and 2.1.283 →
2.1.284.

The trace row asked for the cooldown timing as well (#67/c5881158070/2c). With the kit's
dependabot.yml (github-actions monthly, `cooldown` 7 days), a target's review model moves one
to five weeks after the action release that moves its alias, plus the bump PR's merge. This is
derived from the schedule, not measured.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… four claims say what the kit does

- Install (#4): any existing `.claude/` was refused and sent to Upgrading, which starts from a
  Kit commit row such a repository does not have. Claude Code writes
  `.claude/settings.json` on its own, so this is the common case. The snippet now refuses only
  a repository that already carries the kit (`.claude/rules/cbk-conventions.md`). Otherwise it
  copies file by file, never overwriting, and lists each file it skipped for a hand merge.
  Run against a scratch repository holding its own settings.json: kept and listed, hook exec
  bits preserved, and a second run refused.
- Upgrading (#17): read the CHANGELOG entry of every release after the recorded one, as the
  conventions and the CHANGELOG say, not "from" it.
- `docs/` (#18) is the ADR starters plus each harvest's design, trace and plans, not "the
  kit's own decision records".
- Customization (#13): the lead-in said four files need editing while item 2 says `finish.md`
  needs none.
- The ADR opt-out (#19) would turn the block red. With no numbered ADR the guard and its CI job
  never fire, so the README says keep them, and lists what removal really takes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…what the kit does

- New files (#2): `run-arms-headless.py` and its fixture are required `add` rows. The block
  runs the fixture, and the fixture fails without the runner. They were listed only as files a
  target "may already run ahead".
- The advisory hooks (#10): only `format-on-edit.sh` has the `.claude/workflows/` skip floor;
  both now exit 2 on a finding.
- Prices (#14) live in `orchestration-reference.md`, not `orchestration.md`.
- CI scope (#15): v0.5.0's `verify.yml` runs on pull requests to main.
- First tagged (#16): v0.1.0 is the earliest release with a tag, set after the fact.
- The ADR guard (#20) covers Edit, Write and MultiEdit, whatever names the path; its `Not
  seen:` line lists the rest, which the CI job refuses.
- What landed gains the whole-kit review. The sync notes gain the review round's hand-merges
  (the review templates' dispatch rule, verdict marker, skip line and claude.yml's job gate)
  and seven notes: fail-open warnings shown as a systemMessage, older blueprints' missing
  § Workstreams, Linear's workstream labels, the runners' strict configs, cost figures read
  before this release, the gate block's "not run" form, and the guards' new cases.
- Always-loaded total 132,803 bytes; checked on Claude Code 2.1.286 and 2.1.287 (README too).

The release date stays to be set at merge (#21).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every row now carries a closing status. 369 are `landed` with the commits that name them; the
five rows the landed-row audit found partial carry the review round's completing commit beside
their first (#67/c5881158070/2c 9a31317, critic/9 c098399, review/consistency/3 8063291,
review/consistency/49 e6da512), and release/7 names the tag commands and CHANGELOG headings
that land here, with the tags pushed after merge on the operator's go-ahead. One row is `not
holding` (#69/c5859811353/H1) and one is `non-goal` (#58/c5901493591/residue-2).

A fresh-context, refute-by-default verifier (workhorse tier, high effort, read-only) tried to show
that each non-landed row was required and still live, and failed both times:
`--disallowedTools` is a commander variadic whose parser splits every argv element (2.1.287
bundle and help), and Shape 2 was declined by the reporter and by the spec. On its suggestion
the non-goal row now cites the clause that covers Shape 2, not collect-then-exit, which is
shape 3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…thub-issues branch form

The README's abbreviated scaffold example recorded `**Kit commit**: v0.5.0 (74edf84)`, though
the Quick start now has a reader record the release they install, `v1.0.0 (<sha>)`. Its branch
form `<type>/tui-<N>-<short-slug>` put a team prefix on the GitHub Issues axis, where § Branch
naming makes the key the bare issue number: `<type>/<N>-<short-slug>`. (A review-sweep finding
the verify bound had left unverified; checked and true.)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e P4)

The three bind-mount cases of protected-paths-hook-fixture.sh need an
unprivileged user and mount namespace, and print a SKIP line where the host has
none. The fixture's header now records, dated, what the kit's CI runner did on
this PR's run: on ubuntu-24.04 `unshare -rm` is denied, so they SKIP there (47
cases ok in CI; 50 locally, where they run).

Probe: P4. Decision: D44.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y, so a fork no longer pays for its parent

A fork's transcript (an Agent-tool `fork`, `isFork` in its agent-<id>.meta.json) opens with its
parent's history copied line for line: the parent's responses with the same message ids and
usage. The reader dedups within a file, so it billed that prefix to the parent and again to
every fork. On the three fork-bearing session directories on this machine it overstated the
total by $2.65, $4.49 and $5.54, the last a fifth of the real $25.94: one reviewer's first 33
responses were billed four times.

Each response is now billed once per directory. Forks are read after their parents, at every
level, and keep only what the parent did not record. A fork of the main loop drops the session
transcript's responses beside the subagents directory, without a row. A fork whose parent is
missing is cut at the `<fork-boilerplate>` tool result that hands it its task; one with
neither is unpriced and named. A response with no message id is keyed by its file as well as
its line, so it never matches another transcript. Rows are labelled by the meta.json
description (a workflow stage's label, where the first prompt is harness boilerplate), and a
fork is timed from its own start. A summary line names each row's inherited count.

agent-cost-fixture.sh adds a fork directory in Claude Code's layout: a parent, a fork, a fork
of that fork, a main-loop fork, a fork with a missing parent and one with no boundary. It went
red first: $60.00 billed where $24.00 was spent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…apshot as a floor

A response's last line that says "stop_reason": null was written mid-stream: its input and
cache counts are final, its output (thinking included) a snapshot of 2 to 8 tokens. Through
Claude Code 2.1.276, 89–99% of subagent responses on this machine were written again once they
stopped. From 2.1.278 only 6–31% are, and the transcript holds no other record of the final
count. The harness's own per-agent `subagent_tokens` runs 4k–10k tokens past a transcript's
final context: that is the last response's output, never written. Such a response is billed as
recorded. Each row now counts its snapshot responses, and a summary line names those rows'
output, and their cost, a floor. An absent key is no evidence and counts as none.

agent-cost-fixture.sh adds vvv, one stopped response and one snapshot, red first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es the fork holes a delta check found

A refute-by-default check of 6dfcade and c5a3f26 (three read-only verifiers) found five defects.
Each is now a fixture case that went red first:

- A fork's copy of its parent's history carries the parent's final usage; the parent's own
  last line is often a mid-stream snapshot (31 of 33 and 9 of 9 on two real parents).
  6dfcade's premise, "same message ids and usage", was wrong. Each response is now billed at
  its best copy across the directory (a stopped one, then the larger output), and the floor
  counts only responses with no stopped copy anywhere. A parent with a snapshot of 8 output
  tokens is now billed for the 100,000 its fork copied.
- History a fork excluded (a cut orphan's prefix, a main-loop fork's session responses)
  never entered the billed set, so a fork of that fork billed it again; a fork of an
  unseparable fork billed its parent's work. Excluded history is now accounted for, and
  unseparability passes to children.
- The boundary was the first user line containing the marker anywhere, so an inherited
  grep result quoting it cut too early. It is now structural: a text block opening with
  `<fork-boilerplate>`, beside the tool result for the meta's `toolUseId` where there is one,
  which is the real two-block shape. bp() now emits that shape.
- Every row, forks or not, was timed and labelled from after its last foreign response; a
  non-fork sharing one response with another lost both. Only a fork now starts after its
  copied history, and a cut fork starts at its cut.
- The fork-of-fork case sorted after its parent by name, so depth order went untested; it now
  sorts before.

The docstring now says copies are same-uuid lines carrying final usage, not byte copies. It
also notes that the main-loop fork path is modelled only by the fixture, and that an id-less
response is the one exception to billing once. On the real fork-bearing directories the
per-file overcount is now $2.48, $4.28 and $5.40 (a fifth of $26.08). Corrections to the earlier
commit messages: a snapshot's output is usually under ten tokens, not "2 to 8", and
`subagent_tokens` runs 2k–11k past the final context, not 4k–10k.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…riginal transcripts

orchestration-reference.md § Applied instances carried cost figures the old reader had
inflated, noted "re-run before reusing". They are now re-derived on 2026-10-01 from the three
runs' transcripts (Claude Code 2.1.258 and 2.1.259, which carry final usage on all but a few
responses), and an independent reader recomputed every number:

- framing A/B: output at most 1.6% of a drafter's tokens (was "under 0.5%"); the top tier 73%
  cache writes (78%); 24 against 65 turns pooled (34 vs 52); top:workhorse 1.48x pooled, 1.72x
  and 1.28x by arm (3.2x). The transcripts settle the model: claude-fable-5-1, whose cache
  reads bill at 0.025x. The top-tier procedure drafter read 3.6x the contract one's cached
  tokens (2.7x), 1.3x at the workhorse tier.
- rough-in A/B: the xhigh judge cost 8% more than the high judges' mean (15%).
- /finish A/B: executors $32.54 and $33.01 (about $50), judges $9.12–$11.81 ($18–23).

On these runs the per-line count was the whole error: the recorded figures ran 1.5x to about
2.2x high. The 1-hour-write and fork defects fixed the same day did not touch them, and the
note says so. The CHANGELOG sync note gives the same range and the floor on current
transcripts, and its orchestration hand-merge row says to take the kit's re-derived
observations. The dated planning copies under docs/superpowers/ stay as written.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@j4th
j4th marked this pull request as ready for review October 2, 2026 03:24
@j4th
j4th merged commit 7dbd6b2 into main Oct 2, 2026
2 checks passed
@j4th
j4th deleted the feat/harvest-5-v1.0.0 branch October 2, 2026 03:24
j4th added a commit that referenced this pull request Oct 2, 2026
The heading carried the date the branch was planned on. CLAUDE.md dates a release by its merge
in UTC: PR #75 merged at 2026-10-02T03:24:51Z (gh pr view 75, mergedAt), so the heading reads
2026-10-02. The PR number it predicted, #75, was right. The five back-dated headings already
match their merges' UTC dates.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant