feat: harvest 5 — v1.0.0: guards that resolve paths, the block as a gate, review bots on aliases, the Opus 5.5 recalibration, and the first release - #75
Merged
Conversation
The design for landing the sixteen open issues (#58's residue, #60–#74) and the whole-kit review as the kit's first tagged release: decisions D41–D62, ten clusters in commit order, six execution-time probes, and a 371-row trace of every ask by issue and comment id. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A master plan (constraints, review focus, ownership map, baseline, and the closing tasks F1-F5: budget and sanitization gate, the review floor once, the trace audit, the draft PR, and the post-merge tags and release) and ten cluster plans, V1-V10, 94 tasks in commit order. Every ported hunk is re-authored inline; every trace row has a home; each cluster file was dry-run on a scratch copy of the kit and reconciled against its neighbours. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the block runs A filled target (docs/cbk/scaffold.md exists) now also owes `verification: project sub-block complete`: the done sentinel prints whether or not the project sub-block ran, so a guard that stopped firing turned every project check off with the gate green. run-verification-block-fixture.sh drives the three rails in six cases on synthetic blocks, and the kit sub-block runs it. § Verification › Run it says three rails, names the residual (a deleted or renamed scaffold.md), and cites context-builder-kit#58 instead of a bare #58. Trace: #73/body/1, critic/22 (D51). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… gate detect-forked-agent-memory.sh fails open by design (exit 0 with a WARNING) when it could not look: no checkout, an unenterable root, a partial scan, an unmonitored walk. The block's check read only the exit code, so it passed on nothing while the hook named the block as its backstop. stop_hook_clean reads the hook's stderr; an in-block negative case, outside any checkout under a git ceiling, must itself come back red, so a regression is caught where the check runs. The preamble states the rule. Trace: #73/body/2, the block's half (D51). The hook's WARNING contract and the partial-scan fixture case are V2's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…oaded bytes The always-loaded loop printed the standing cost but nothing flagged growth. Above 140,000 bytes it now prints a WARN line and the run goes on: the number is a signal to path-scope, split or delete a rule, not a gate. Shown by padding an always-loaded rule to 140,630 bytes in a throwaway clone: no WARN before, one WARN after, exit 0 both times. Decision: D50 (no trace row of its own). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cation block Two causes kept a fresh scaffold red. The disposition pass never asked about cbk-conventions-reference.md's bracketed manifest-and-lockfile glob, which the project sub-block's stamped-globs check rejects: the table gains the row, § 4's first sentence counts three path-scoped placeholders, and the phase-exit item and Test 1 name it. And the kit sub-block required a CLAUDE.md mention that blueprint, not scaffold, writes: a scaffolded target now owes it once docs/cbk/blueprint.md exists, and the kit tree always does. A kit-sub-block pin fails when a rule ships a paths: placeholder the disposition table has no row for. Shown on a simulated scaffold (drop-in set, ADR starters, scaffold.md, globs stamped): all three sentinels; unstamped, red naming the file. Trace: Review Focus 4; the Gate settled call (disposition row); the defect behind review/portability/68 (V9's row, landed here). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…check Nothing in the templates ran the block, so a target ran it only if someone remembered. Blueprint's tooling template: the minimum task set, light mode included, gains a verification task that check depends on, whose body is the runner, with the tools it needs (bash, git, awk, mktemp; jq, node, python3 for the block's fixtures); § Rules makes it a leg of check, run in CI through check and once before the HITL presentation. The phase-exit checklist and Test 1 assert it. Scaffold's verification matrix runs the runner after the disposition pass and expects both sentinels. Run it names the two homes, and a kit-sub-block pin keeps both templates naming the runner. Trace: #58/c5901493591/R4 (D51). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cker, scoped to one file An absent check brackets one letter of the phrase it hunts so the grep never matches its own line; a target that spellchecks .claude/ reads each fragment as a typo. § Verification now says to exempt the idiom in this one file, never repo-wide, and shows the typos form: a [type.<name>] table whose extend-glob names cbk-conventions-reference.md and whose extend-ignore-re matches the idiom, quoted from typos' reference (read 2026-09-30) and exercised against typos 1.50.3. A kit-sub-block pin keeps the note. Trace: #58/c5901493591/R5 (D51). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The kit's own verify.yml ran with GitHub's unset shell (bash -e, no pipefail) on ubuntu-latest, a label that moves under the gate; its workflow templates take the same two pins later in this PR. verify.yml gains defaults.run.shell: bash and runs-on: ubuntu-24.04, each with its source read 2026-09-30 (the workflow-syntax shell table; the runner-images migration note and the 24.04 image readme). A kit-tree-only check in the block pins both lines. The job name, the required check, is unchanged. Decision: D52, the kit's-own-workflows settled call (no trace row of its own). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ails closed The job's body is a :(glob) pathspec with --diff-filter=a and --no-renames, and any output fails it: the awk match it replaces read a quoted non-ASCII name and a spaced name as "no ADR changed". The diff is three-dot, so an ADR merged to the base after the branch point no longer reads as a deletion. A git error now fails the step, where mapfile of a process substitution read it as a pass. The checkout is blobless; the job runs under shell: bash on a named runner image. adr-ci-body-fixture.sh runs the job's own body, extracted by the shared extract-run-block.sh; the block runs the fixture and pins what it cannot see. Red first: FAIL: modify a non-ASCII-named ADR (want exit 1, got 0) Trace: #60/c5876324022/adr-job, #60/c5881158391/port-glob-body, #60/c5892401033/adr-job-glob-fixture, #60/c5901494943/3, #61/body/fix, #61/body/grep, review/security/19; the extractor for #67/c5892401572/2. Decisions: D55, D44. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…helper protect-immutable-adrs.sh matched the payload's path text against a prefix, so every spelling context-builder-kit#60 measured passed: a .. segment, . and //, a trailing-slash or foreign project dir, a relative path, a symlinked directory, a symlink to the file, a hardlink, /proc/self/root and an unparseable payload. The guard now sources .claude/hooks/lib/resolve-path.sh: the payload read exactly in one jq, the path resolved lexically and physically, /proc and unreadable symlinks refused, the same file under another name caught by bash's -ef over the hook's own checkout, the project dir and a linked worktree. The deny is path-independent: an existing docs/adr/NNNN-*.md in any checkout; the existence test runs on the resolved path, so a new ADR stays creatable. A missing helper or jq fails open naming the ADR job. hook-contract-fixture.sh's early-reader check reads lib/*.sh; section Hook authoring states the sourced-helper contract and the lib/ gitignore trap, and the block asserts the helper is not ignored. Red first: FAIL [protect-immutable-adrs.sh]: a .. segment (docs/sub/../adr/) (want exit 2, got 0) Trace: #60/body/fix, #60/body/alt, #60/body/fixture, #60/c5876324022/table, #60/c5876324022/closed-1, closed-2, closed-3, closed-4, #60/c5876324022/closed-5, closed-6, #60/c5876324022/closed-note, #60/c5881158391/port-helper, #60/c5892401033/helper, #60/c5892401033/adr-path-independent, #60/c5892401033/linked-worktrees, #60/c5892401033/ef-inode, #60/c5892401033/gitignore-trap, #60/c5892401033/fixture, #60/c5901494943/2a, 2b, 2c, 2d, 2e, #62/body/extra-adr-hook, #58/c5901493591/R1 (the path guard), release/5 (section Hook authoring), #69/c5859756889/apply-h4/1 (a target's own hooks, held to the same contract), review/consistency/48 (the tally, rewritten count-free). Decisions: D43, D44, D61. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… refused or asked The lock-file, main-branch and PR-state guards failed open on a payload jq cannot parse (one lone UTF-16 surrogate escape); the two denies now refuse it and the ask-gate asks, and section Hook authoring says why an unreadable payload is not an environment defect. Each fail-open warning names what still stands: a bracketed slot for the project's CI lockfile check and for the base branch's ruleset, filled at scaffold's rule-file disposition pass (a new bootstrap row) and refused unfilled by the project sub-block; the PR-state gate says no mechanical backstop exists. The lock-file fallback's remediation names the real route, a case arm above the *.lock) arm. The knowledge-backend ask-gate's header carries its ten verbs, and protect-main-branch.sh gains the Depends: line every header carries. hook-guards-fixture.sh drives every branch of the four guards, finds the ask-gate through the registry, and runs one deny per guard from a root containing a space. Red first: FAIL: the non-checkout warning names a backstop (stderr lacks 'Backstop') Trace: #62/body/lock-files, #62/body/pr-state, #62/body/main-branch, #58/c5901493591/R8, #58/c5901493591/R1 (the four guards), critic/8, #66/body/2 (the ask-gate header), release/5 (protect-lock-files.sh). Decisions: D61, D44. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…racters The left anchor admitted only a start, a space or ;&|, so a commit on main passed the hard-deny inside bash -c "…", sh -c '…', a subshell, $(…), backticks or through /usr/bin/git, and the right anchor let git commit; and git commit&&… through; the PR-state ask-gate shared the left anchor. Both anchors are now negated classes: anything that cannot continue a word before, anything but a word character, . or - after, so git commit-tree and git config commit.gpgsign stay allowed. Each header's Not seen: line names what a pattern guard cannot see (eval, a variable, an alias). The remediation's comments say "an issue this PR closes" and "work no issue tracks". The fixture gains every spelling, both ways. Red first: FAIL: a commit on main is denied: bash -c "git commit -m x" (want exit 2, got 0) Trace: review/security/20, #68/body/1a (the guard's remediation). Decision: D44. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t hooks, by payload require-repo-root-for-agents.sh failed open on a payload jq cannot parse; it now refuses it. hook-payloads-fixture.sh drives every branch the launch-root guard and the forked-memory detector document, a root containing a space included, and pins the literal WARNING on each could-not-look branch, the partial scan among them (context-builder-kit#58 items 1, 2 and 4). Probe P1 settled whether a hook payload's cwd follows the Bash tool's cd on the release CLI; the guard's Timing paragraph, the main-branch guard's cwd comment and section Hook authoring state the answer, dated. The detector's rule 1 now states what it guarantees and points at its Residual; its bare citations are qualified. Red first: FAIL: a payload jq cannot parse is refused (want exit 2, got 0) Trace: #58/c5901493591/R1 (the two hooks), #58/c5901493591/R6, #58/c5901493591/R13, #73/body/2 (the WARNING pins), release/5 (the two hooks). Decisions: D61, D44. Probe: P1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A CI job named as a hook's backstop but never landed is how a target's frozen corpus went two weeks unguarded (context-builder-kit#60, suggestion 1). The kit sub-block now reads the mutation table, the hook registry and every hook header, and fails on a .github/workflows or scripts path that does not exist; the checker is asked about bogus names first, and a mutation section that names no checkable backstop is red. The table's ADR row names its workflow by path. The project sub-block resolves `mise run <task>` names against mise.toml where one exists. Red first: cbk-conventions.md § Mutation discipline names no CI backstop in a checkable form (.github/workflows/<file>.yml) Trace: #60/c5876324022/corpus-s1, #60/c5881158391/port-s1, #60/c5892401033/suggestion-1-built. Decision: D51. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… checklist row per item Consultation's enforcement set told a target to extend the ADR job, which then inherited its awk path match. The recipe now builds the corpus guard on lib/resolve-path.sh's root-scoped mode and the corpus CI job on the ADR job's parse-nothing body, and lists the nine closures two targets' red-team rounds measured, ending with the additions call each target makes. The bootstrap checklist gains one verification row per enforcement item, so no item is ticked off with the others; brownfield audit and consultation's Test 4 say the same. The block asserts the closures and the six rows. Red first: frozen_corpus_ingestion.md lacks the CI job's closure list or the parse-nothing body Trace: #60/c5901494943/1, #60/c5876324022/corpus-s2, #60/c5876324022/echosphere-z, #60/c5892401033/corpus-leg, #60/c5892401033/leg-defects, critic/6, critic/7. Decision: D55. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d the three-edit wiring Verify by payload said the payload-reachable branches were asserted in hook-contract-fixture.sh, which drove none of the ADR guard, the knowledge-backend ask-gate or the main-branch allow cases. It now names the fixture of each hook family, all run by the block, and what a target does when it adds a hook. The advisory-hook bullet says that wiring one is three edits: the stanza, the ADVISORY_WIRED name and the two-views paragraph, each checked by the block. The block asserts both. Red first: § Hook authoring › Verify by payload does not name hook-guards-fixture.sh Trace: #58/c5901493591/R1 (section Verify by payload), #58/c5901493591/R11 (section Hook authoring; the exemplars' Register: headers are V8's). Decision: D44. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ke-only `$1` is the second argument (code.claude.com/docs/en/skills § Available string substitutions, read 2026-09-30), so `/finish 42` left `$1` literal and `/finish 42 --skip-review` rendered `#--skip-review`. Every command and both bundled executor copies now declare `arguments:` and write `$issue`, `$pr` or `$source`. The five commands carry `disable-model-invocation: true` (D57): they open branches, issues and PRs. `/finish` item 4 reads `--skip-review` from `$ARGUMENTS`, so pr-review.md's break-glass flag is honoured (Review Focus 5, probed on 2.1.285). The block pins all three. Trace: review/claude-code/9, review/claude-code/14. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every registration was shell form with an unquoted ${CLAUDE_PROJECT_DIR}.
Under a project path containing a space the shell split it, the script was
not found, and each PreToolUse guard's non-2 exit let the tool call through:
every hard-deny guard failed open with no signal (probed on 2.1.285). Each
registration now carries "args": [], the exec form the hooks page prescribes
for a hook that references a path placeholder (code.claude.com/docs/en/hooks
§ Exec form and shell form, read 2026-09-30). The block refuses a placeholder
registration without an args array (Review Focus 2). § Hook authoring now
states the form: the stdin bullet's registered call is the exec-form entry,
and the Project-relative paths bullet says to register in exec form and why
(V2 declined these two sentences and granted the region; they land here).
Trace: review/claude-code/10.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…erge rule is the documented one enabledPlugins is keyed <plugin>@<marketplace> (code.claude.com/docs/en/ plugin-marketplaces, read 2026-09-30). The bare "pr-review-toolkit" key enabled nothing, so the review floor's toolkit half was never declared; the keys are now pr-review-toolkit@claude-plugins-official and commit-commands@claude-plugins-official, and the comment gives the install commands and points at simplification.md § Plugin instead of restating a CLI version. tooling.md's always-loaded rule said a repeated list in local settings "silently replaces" the committed one; the settings page says lists merge, except four model-list keys, and a local false is the documented per-machine plugin opt-out. Always-loaded: +209 bytes. Trace: review/claude-code/13, review/release/29, review/claude-code/16, #69/body/F9 (settings.json part). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
_comment_hooks restated the tier definitions and per-hook facts each hook header and § Hook authoring already carry (4,080 characters). It is now the registry: each hook under its tier with its event and matcher, the exec form, pointers, and the facts other clusters handed to this file — the ADR guard resolves the path both ways through .claude/hooks/lib/resolve-path.sh, its CI backstop is .github/workflows/adr-immutability-check.yml, and an unparseable payload makes a hard-deny guard refuse (D61). The block caps it at 2,500 characters and requires it to name every sourced helper. enabledMcpjsonServers drops github (D58: gh is the kit's GitHub interface), its comment no longer tells a reader to comment out JSON, and the block checks that every listed server is declared in .mcp.json. Trace: #66/body/2 (settings.json part), #60/body/fix (registry sentence), #60/c5892401033/helper (registry names the helper), #60/c5876324022/corpus-s1 (registry names the ADR workflow), #70/body/6a (enabledMcpjsonServers). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… an override A project agent named Explore replaces the built-in (code.claude.com/docs/en/ sub-agents, read 2026-09-30). The kit's override kept searches on the smallest tier but, unlike the built-in, loaded the whole CLAUDE.md hierarchy and every unscoped rule on each dispatch. It now sets omitClaudeMd: true (v2.1.271+), its description is the routing sentence (the provenance moved to a maintainer paragraph in the body), and the bootstrap checklist asks whether to keep it, so an adopting project knows it ships a live override. Trace: review/claude-code/15, review/release/63 (agent and checklist parts). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s the rubric's reference half /intake reads an outside reporter's text and /pr-respond any commenter's, yet neither stated the rule the research phases carry: fetched text is data, and an imperative inside it is surfaced, never acted on. Both now do. /intake writes its own reproduction, names it at the Step 3 gate, and never runs a command quoted in the report. /pr-respond applies a finding only from the PR's author, a collaborator with write access or the project's review bot; any other author's finding is Surface at most. /pr-respond also reads pr-review-reference.md's calibration, path-conditional aggressiveness and anti-patterns explicitly: a triage is not a file read, so the path-scoped half never loaded (code.claude.com/docs/en/memory, read 2026-09-30). Trace: review/security/21, review/claude-code/52. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ted, /simplify's dimensions are sourced The cascade-rule reviewer cited a knowledge-backend.md section that does not exist; it now cites § The code-adjacent split — canonical, and the block checks that every § citation in an agent resolves. The logging reviewer flagged per-tick debug calls that logging.md § Level taxonomy permits and contradicted its own level-taxonomy item; it now flags above debug and names telemetry as the better form. simplification.md asserted five behaviours of the bundled skill with no source; it now quotes the commands page's four dimensions and states the rest as the project's own rules. The rule index row no longer claims a triage the file does not contain. Three dangling docs/STANDARDS.md section citations are gone. Always-loaded: −149 bytes. Trace: review/claude-code/58, review/claude-code/59, review/claude-code/60, review/consistency/40 (simplification.md and logging reviewer parts). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…trings in the emitted workflows Every workflow the kit emits (both review templates, blueprint's CI skeleton, scaffold's CI stub) sets `defaults: run: shell: bash`, so each run: step has pipefail, and runs on ubuntu-24.04 rather than the moving -latest label. The review templates' label checks read a here-string instead of piping into an early-exiting grep -q, and the label-removal step captures its label read, so a failed read is an error and a label deleted from the repository is the sourced benign branch. The CI skeleton's two `uses:` lines take the fill-slot form @[full-commit-sha] # v[version], as the review templates pin theirs, instead of a mutable tag. The kit sub-block carries the three assertions, each shown red by a mutant. Trace: #70/body/1a, #70/body/1b, #70/body/1c, #70/body/2, #70/body/3, #67/c5901495773/8, critic/1, critic/2, review/claude-code/56 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…urrency and a fork guard The review trigger is a `paths` filter whose re-includes (.claude/**, docs/adr/**) come after the markdown negation, so rules, skills, commands, agents and ADRs are reviewed while docs-only PRs are not; review-trigger-fixture.py evaluates the filter as GitHub does over twelve PR shapes and runs in the kit sub-block, red first on the old paths-ignore. `concurrency` moves to the job in both templates, so a run the job's if: skips cannot cancel a live review or response, and a fork PR, which gets no secrets, is skipped rather than failed. The prompt's skip clause and blueprint's review-bot paragraph name the same shape. Trace: #67/c5901495773/1a, #67/c5901495773/1b, #67/c5901495773/1c, review/security/18, review/security/22 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the verdict and says why "Assert the review posted" fetches the PR's comment pages once and counts them with jq -s (gh api --paginate --jq ran per page, so a review past comment 30 read as missing); matches the review app's comment by its verdict marker on updated_at, so the tracker alone never passes and a tracker ending in the verdict does; runs on !cancelled(); and, when no summary landed, says why: the session ran without a verdict, or no session started, told apart by a diff of the workflow against the PR base (a validation skip) or not (auth or setup). The Auto-review step gains `id: review`, Constraint 4 is restamped with the exercised paths, and review-assert-fixture.sh runs the step's own body on synthetic comment pages through V2's extract-run-block.sh, red first on the old step. Trace: #67/c5877251222/1, #67/c5877251222/2, #67/c5877251222/3, #67/c5877251222/4, #67/c5881158070/1 (corrected by c5892401572/1), #67/c5892401572/2, #67/c5892401572/3, #67/c5901495773/confirm-2, critic/12, critic/13 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rded model Both review templates name the model by family alias (opus by default and on the deep path, sonnet on the fast path and as the fallback; never best), and every branch passes --effort, claude.yml at high, because the default effort is per model and Opus 5.5 and Sonnet 5.5 default to medium. The header says the action's pinned SHA is the model pin, never to set ANTHROPIC_MODEL (it wins over --model), and what the action-bump PR re-checks; the deep branch says what choosing fable costs and why xhigh is named now that ultracode sets no effort. A "Record the resolved model" step on both templates (continue-on-error) writes the model the session started on, its Claude Code version and every model that answered to the job summary, and the prompt's model line says it shows the alias. The fast branch's effort=high is labelled a pin pending a Sonnet 5.5 effort sweep, since Sonnet 5.5's levels are recalibrated. The fixture runs the record step on synthetic execution files, red first. Probe P3 recorded. Trace: #67/body/1a, #67/body/1b, #67/body/1c, #67/body/1d, #67/body/2a, #67/body/2b, #67/body/2c, #67/c5804984992/1, #67/c5881158070/2a, #67/c5881158070/2b + #67/c5892401572/Effort, #67/c5901495773/confirm-1, #67/c5901495773/5 (template half), #69/c5859756889/3-C4, #69/c5881157875/3a (template half), critic/18 (the fast branch) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e base branch's The review prompt says the action restored .claude/, .mcp.json and the root CLAUDE.md from the base branch and kept the PR's copies under .claude-pr/, that a nested CLAUDE.md (and, where a project has one, a symlinked AGENTS.md) reads at the PR head, that a rule the PR changes is reviewed as the PR states it with the version named, and that CI's check job, read with gh pr checks, is what runs the PR's own hooks and fixtures. Both headers state the mechanism with its source. The summary reads CI and reports only commands actually run; a note above claude_args pairs every prompt gate with an --allowedTools entry and sizes the degrade clause's N (40) against --max-turns 60, re-measured on Claude Code 2.1.286. The two kit-issue citations are qualified. Trace: #68/body/2a (+c5861166655/2, c5881158244/1, c5892402074/2), #68/body/2b, #68/body/2c, #68/body/2d, #68/c5881158244/2 (+c5892402074/2), #68/c5901496131, #67/c5901495773/2, #67/c5901495773/3, #58/c5901493591/R12 = #67/c5901495773/4, #67/c5901495773/7, critic/14, critic/15, release/5 (the two template citations) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…g.md § Automated review The review template's Constraint 1 says the workflow cannot review any PR that changes it, quoting the action's own log line for the skip, and tooling.md § Automated review says the same, names the assert step that goes red on such a PR, and adds that the pinned action fixes the model and the Claude Code the bot runs, so a harness fact the bot asserts is checked against the operator's version. Constraints 3 and 6 are sourced to the action at the pinned SHA, and the track_progress rail is dated: labeled is accepted from v1.0.188. The block's review-automation check and a tooling.md check follow the corrected wording. Always-loaded +452 bytes (tooling.md). Trace: #67/c5901495773/6, #67/c5901495773/9, review/portability/35 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e job's if:, before the concurrency group The bot and label checks sat in a step, which runs only after the job has joined the job-level concurrency group. A bot comment that mentions @claude passed the job `if:`, joined the group, and with cancel-in-progress cancelled the reply in flight before the step could skip it. The gates now sit in the job `if:`, as claude-review.yml's already do. `sender.type == 'User'` mirrors the action's own gate, which refuses any non-User actor unless `allowed_bots` names it (action.yml default "", checkHumanActor in src/github/validation/actor.ts, read 2026-10-01), so a bot is skipped green instead of failing red. review-assert-fixture.sh pins the gates in the job `if:` and refuses a step-level gate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s tool list allows
The prompt all three label paths share forbade subagents ("Do the review in this one session",
"inline yourself (one session, no subagents)"), while the deep path drops `--disallowedTools
Agent` and turns on ultracode precisely so it can fan out. The label step now writes a
`dispatch_rule` beside the tool list: the default and fast paths say do not dispatch, and the
deep path says it may, with every subagent returning before the summary posts. The prompt
interpolates that rule in both places. review-assert-fixture.sh runs the label step on five
label sets and checks the precedence (deep over fast over default, `claude-deep-review-wip` no
deep label), that each path's rule agrees with its tool list, and that the prompt states no
dispatch rule of its own.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…old skip line, and names a moved base "Assert the review posted" counted any comment whose body contained APPROVE, REQUEST CHANGES or NEEDS DISCUSSION, so a tracker note naming a verdict word in passing read as a landed review. The marker is now a bold verdict, as the prompt prescribes; a verdict after a label on its line still lands. The docs-only skip line the prompt asks for carried no verdict and so turned a correct skip red, with a "no summary posted" notice: it now opens with a bold `**SKIPPED**`, which the assertion accepts. The fixture checks every bold word on the prompt's verdict line and the skip line against the step, so the two vocabularies cannot drift apart. The validation-skip notice blamed "this PR" for changing the workflow. The diff is tree against tree, which is right: the action compares contents, so a base that changed the file after the PR branched skips the review too. The notice now says the copies differ, and a fixture case pins the moved base, so a merge-base diff would go red. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rmed entry cannot pass a token by The block runs without errexit, and the credential check piped jq into `sort -u`, so a jq error — `"headers": "Authorization: Bearer ghp_…"`, a string where an object belongs, gives "string has no keys" — left `bad` empty and the literal token passed (reproduced on a scratch clone: exit 0). The two checks beside it had the same shape without the pipe. Each jq now runs alone, a failure stops the block with the reason, and the sort runs afterwards. On the clone the string header now fails, and so does the same token written as an object. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the kit tree An agent's `§` citation to a rule file that does not exist was skipped everywhere, which is right only on a filled target, where the disposition pass may have deleted the rule. On the kit tree every rule exists, so `cbk-convention.md § ADR index sync` passed silently (reproduced on a scratch clone). It now fails there with the file it names, and is still skipped on a target with docs/cbk/scaffold.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd the sweep's bounds are pinned by behaviour only The nested-file case passed with or without `:(glob)`, since the file was an addition and additions are allowed. The fixture's base now holds docs/adr/0009-notes/x.md, and a case edits it: only the `:(glob)` pathspec leaves it out, because a plain pathspec's `*` matches across `/`. With the magic removed the case goes red. The block grepped review-sweep.js for `maxPerDimension ?? 3`, `maxVerify ?? 8`, FIND_EFFORT and RETRY_EFFORT. The accounting harness already asserts the 3/8 default bounds and the finder efforts as behaviour (scenarios 14, 23 and the default-bounds scenario). A behaviour-preserving refactor broke the greps, and `??` turned to `||` passed them, so they are dropped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…two that overclaimed say what is true orchestration-reference.md's preamble said orchestration.md keeps a pointer heading for every section there, and workflows.md's rule index repeated it for all four halves. Two sections added since the split, § Generation notes — the sources and § Cost terms and run hygiene, have none. The V9.19 check looped over only the conventions and review halves, so nothing caught it. The loop now covers all four halves. It goes red on the orchestration preamble, which now names the two later sections, and the index row says the pointer is kept for every section moved at the split. On a target a deleted pair is skipped; the kit tree must hold all four. knowledge-backend-reference.md said its sections "were moved verbatim on 2026-09-30", but the same change re-sourced four of them: it now says "moved … and re-sourced in the same change". Always-loaded total 132,473 → 132,494 bytes (workflows.md, +21). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the [1m] variant at list
Every cache write was priced at the 5-minute 1.25x. Claude Code records the per-TTL split in
`usage.cache_creation`, and a session on the 1-hour TTL writes all of its cache there, which the
pricing page bills at "2x base input price" (read 2026-10-01). Writes are now split, the
5-minute part at 1.25x and the 1-hour part at 2x; a write with no breakdown stays at 1.25x.
tier() accepted any bracketed suffix, so an id whose variant the table does not know priced at
list. It now accepts only `[1m]`, which the same page says bills at standard rates ("Claude 4.6
and later models ... include the full 1M token context window at standard pricing"); any other
variant is unpriced and named. The docstring's "each assistant message carries usage" now says
what the dedup reads: lines that repeat a response's message id, the last line standing for it.
agent-cost-fixture.sh adds sss (a TTL-split write, $13.00) and ttt (`[2m]`, unpriced), both red
first: $88.15 over 13 of 20 priced agents. orchestration-reference.md's note on past figures adds
the 1.25x write rate, which pulled a 1-hour session's figure the other way.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ual setup every citation names More than twenty surfaces cite `blueprint.md` § Workstreams: /finish Step 1 validates a title's slug against it, framing takes one row, and the conventions say the slug list lives there. The blueprint SKILL describes "the workstreams table … slugs", and the template's own planning-axis text names "the Workstreams table inside this file". The template never emitted one. It now does, one row per core project with its confirmed slug and layer, plus authoring guidance. On the in-repo-markdown axis, seven surfaces put the handoff content in `blueprint.md` § Manual setup, which the template also lacked. Two of them said setup progress is committed to blueprint.md, which § Mutation discipline makes immutable after its commit. The template gains the section for that axis only, written once. Progress moves to the roadmap's step-0 row, which is freely mutable: planning-backend-commit.md, scaffold SKILL.md's disclosure and the roadmap template now say so. Blueprint's test case 1 gains the missing-Workstreams failure. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the gate block has a "not run" form pr-review.md § What NOT to flag and its reference half said a docs-only diff still runs the floor, both skills invoked. cbk-conventions.md § Branch naming, whose first issue-less example is a docs sweep, writes the floor "not run — docs-only sweep, no code changed", and the starter PR template agrees. Both docs-only statements now say "on a cascade PR" and point an issue-less branch at § Branch naming. The `## Review gate` shapes in § The floor had no "not run" form, so the /simplify and toolkit lines gain `or: not run — <reason> (an issue-less branch only)`. Always-loaded total 132,494 → 132,731 bytes (pr-review.md, +237). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…flows apply linear_planning.md created one `area:<slug>` label per workstream area, "mirroring the GitHub label set". The GitHub set uses `workstream:<slug>`, and every Linear flow applies `workstream:<slug>`. The list also lacked `enhancement` and `source:*` (intake) and `triage` (the lifecycle axis), and framing's Linear call sets `appetite:small|medium|big`, which neither taxonomy created. The taxonomy now lists the GitHub axes as the Linear flows use them: the work type rides Linear's own Feature/Bug/Improvement label, and the review-control labels stay on the repository. enrich.md's two `area:<slug>` mentions, and the label examples in scaffold's SKILL.md and output template (which also lacked `chore`), use `workstream:`. The V9.14 check grepped only the GitHub profile. It now checks the Linear taxonomy for every label a Linear flow applies and refuses `area:` labels anywhere in scaffold or the commands. It went red on the old text first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e fact, and the fact's record names it blueprint's handoff-issue.md listed `rust-toolchain.toml` among pins no Dependabot ecosystem maintains. The conventions quote Dependabot's support for it, and tooling.md 6b and the starter templates agree. The line now names the real gaps (the toolchain manager's pins, a `COPY --from` image, a dev container's image) and calls `rust-toolchain.toml` the covered exception. The multi-surface record in § Dependency settle-window listed three surfaces and missed this fourth, which is how the sweep missed it; it now names all four. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ish's --skip-review, and scopes CI to main - The harvest-trace rule now audits a `landed` row against the branch head, not its commit message: the ask is located at a file:line and given a verdict, and a row that does not hold is reopened. This is the audit method a harvest's critic asked to be the gate (trace critic/9). - `/finish` takes the issue number and an optional `--skip-review` (finish.md's argument-hint), not "one argument". - verify.yml runs on pull requests to main and pushes to it, not "every pull request". - scaffold SKILL.md's label-axis list names the lifecycle axis its GitHub profile creates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The review sweep left four rule-text claims unverified. One read-only verifier per claim (workhorse tier, high effort) re-read the primary source on 2026-10-01: - orchestration-reference.md, subagent model resolution: "all checked against the org's model allowlist (an excluded value is skipped and the agent runs on the inherited model)". The sub-agents page § Choose a model checks only the first three values. Since v2.1.222 a blocked family alias runs on "the newest version of that family the allowlist permits", and only other values fall back "on the inherited model instead". - orchestration.md, the over-verification anti-pattern, quoted "verify your work" and "use a subagent to double-check". Neither string is in the Opus 5 or Opus 5.5 guide. The row now quotes the guide's own examples and names § Task scope and over-verification. - cbk-conventions-reference.md § Hook authoring credited stdin delivery to § Exec form and shell form, which says only that the command is spawned with no shell. "Input arrives on stdin" is § Hook lifecycle; each fact now cites its own section. - § Hook authoring, the fatal settings diagnostic: the fatal path still holds in the 2.1.287 bundle, but the trigger is narrower than "any non-hooks key carrying matcher". It fires on a non-empty PreToolUse/PermissionRequest key, or a matcher object not under a hook-event key, up to three levels deep. Nine keys are exempt, and the file loses only its own settings. The dating extends to 2.1.287; "silently" is dropped, since the settings page documents a Settings Error dialog. Always-loaded total 132,731 → 132,803 bytes (orchestration.md, +72). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and the block keeps it so The axis was renamed `in-repo-markdown` when the three named profiles gave way to two independent axes, but "markdown-only" survived as its name in four places in the conventions' reference half, in intake.md and enrich.md, and in rough-in's test cases (trace review/consistency/3). blueprint SKILL.md also said "every profile" and summarised "profile" at inheritance. Each now names the axis; the adjective (a table that lives only in markdown) stays. A block check refuses the old name for a backend, mode, project or planning axis in the skills, commands and rules; it matches eight lines of the pre-fix text. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… Code, and when a target's bump lands The CHANGELOG and the review template cite `src/entrypoints/run.ts` (`claudeCodeVersion`) for which claude-code-action release installs which Claude Code. orchestration-reference.md cited `base-action/action.yml`. Both pin the same version, but the templates `uses:` the root action, whose installer is run.ts, so the reference half now cites it too and names the other as a second pin. Re-read at v1.0.231/232/235/236 on 2026-10-01: 2.1.278 → 2.1.280, and 2.1.283 → 2.1.284. The trace row asked for the cooldown timing as well (#67/c5881158070/2c). With the kit's dependabot.yml (github-actions monthly, `cooldown` 7 days), a target's review model moves one to five weeks after the action release that moves its alias, plus the bump PR's merge. This is derived from the schedule, not measured. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… four claims say what the kit does - Install (#4): any existing `.claude/` was refused and sent to Upgrading, which starts from a Kit commit row such a repository does not have. Claude Code writes `.claude/settings.json` on its own, so this is the common case. The snippet now refuses only a repository that already carries the kit (`.claude/rules/cbk-conventions.md`). Otherwise it copies file by file, never overwriting, and lists each file it skipped for a hand merge. Run against a scratch repository holding its own settings.json: kept and listed, hook exec bits preserved, and a second run refused. - Upgrading (#17): read the CHANGELOG entry of every release after the recorded one, as the conventions and the CHANGELOG say, not "from" it. - `docs/` (#18) is the ADR starters plus each harvest's design, trace and plans, not "the kit's own decision records". - Customization (#13): the lead-in said four files need editing while item 2 says `finish.md` needs none. - The ADR opt-out (#19) would turn the block red. With no numbered ADR the guard and its CI job never fire, so the README says keep them, and lists what removal really takes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…what the kit does - New files (#2): `run-arms-headless.py` and its fixture are required `add` rows. The block runs the fixture, and the fixture fails without the runner. They were listed only as files a target "may already run ahead". - The advisory hooks (#10): only `format-on-edit.sh` has the `.claude/workflows/` skip floor; both now exit 2 on a finding. - Prices (#14) live in `orchestration-reference.md`, not `orchestration.md`. - CI scope (#15): v0.5.0's `verify.yml` runs on pull requests to main. - First tagged (#16): v0.1.0 is the earliest release with a tag, set after the fact. - The ADR guard (#20) covers Edit, Write and MultiEdit, whatever names the path; its `Not seen:` line lists the rest, which the CI job refuses. - What landed gains the whole-kit review. The sync notes gain the review round's hand-merges (the review templates' dispatch rule, verdict marker, skip line and claude.yml's job gate) and seven notes: fail-open warnings shown as a systemMessage, older blueprints' missing § Workstreams, Linear's workstream labels, the runners' strict configs, cost figures read before this release, the gate block's "not run" form, and the guards' new cases. - Always-loaded total 132,803 bytes; checked on Claude Code 2.1.286 and 2.1.287 (README too). The release date stays to be set at merge (#21). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every row now carries a closing status. 369 are `landed` with the commits that name them; the five rows the landed-row audit found partial carry the review round's completing commit beside their first (#67/c5881158070/2c 9a31317, critic/9 c098399, review/consistency/3 8063291, review/consistency/49 e6da512), and release/7 names the tag commands and CHANGELOG headings that land here, with the tags pushed after merge on the operator's go-ahead. One row is `not holding` (#69/c5859811353/H1) and one is `non-goal` (#58/c5901493591/residue-2). A fresh-context, refute-by-default verifier (workhorse tier, high effort, read-only) tried to show that each non-landed row was required and still live, and failed both times: `--disallowedTools` is a commander variadic whose parser splits every argv element (2.1.287 bundle and help), and Shape 2 was declined by the reporter and by the spec. On its suggestion the non-goal row now cites the clause that covers Shape 2, not collect-then-exit, which is shape 3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…thub-issues branch form The README's abbreviated scaffold example recorded `**Kit commit**: v0.5.0 (74edf84)`, though the Quick start now has a reader record the release they install, `v1.0.0 (<sha>)`. Its branch form `<type>/tui-<N>-<short-slug>` put a team prefix on the GitHub Issues axis, where § Branch naming makes the key the bare issue number: `<type>/<N>-<short-slug>`. (A review-sweep finding the verify bound had left unverified; checked and true.) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e P4) The three bind-mount cases of protected-paths-hook-fixture.sh need an unprivileged user and mount namespace, and print a SKIP line where the host has none. The fixture's header now records, dated, what the kit's CI runner did on this PR's run: on ubuntu-24.04 `unshare -rm` is denied, so they SKIP there (47 cases ok in CI; 50 locally, where they run). Probe: P4. Decision: D44. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…y, so a fork no longer pays for its parent A fork's transcript (an Agent-tool `fork`, `isFork` in its agent-<id>.meta.json) opens with its parent's history copied line for line: the parent's responses with the same message ids and usage. The reader dedups within a file, so it billed that prefix to the parent and again to every fork. On the three fork-bearing session directories on this machine it overstated the total by $2.65, $4.49 and $5.54, the last a fifth of the real $25.94: one reviewer's first 33 responses were billed four times. Each response is now billed once per directory. Forks are read after their parents, at every level, and keep only what the parent did not record. A fork of the main loop drops the session transcript's responses beside the subagents directory, without a row. A fork whose parent is missing is cut at the `<fork-boilerplate>` tool result that hands it its task; one with neither is unpriced and named. A response with no message id is keyed by its file as well as its line, so it never matches another transcript. Rows are labelled by the meta.json description (a workflow stage's label, where the first prompt is harness boilerplate), and a fork is timed from its own start. A summary line names each row's inherited count. agent-cost-fixture.sh adds a fork directory in Claude Code's layout: a parent, a fork, a fork of that fork, a main-loop fork, a fork with a missing parent and one with no boundary. It went red first: $60.00 billed where $24.00 was spent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…apshot as a floor A response's last line that says "stop_reason": null was written mid-stream: its input and cache counts are final, its output (thinking included) a snapshot of 2 to 8 tokens. Through Claude Code 2.1.276, 89–99% of subagent responses on this machine were written again once they stopped. From 2.1.278 only 6–31% are, and the transcript holds no other record of the final count. The harness's own per-agent `subagent_tokens` runs 4k–10k tokens past a transcript's final context: that is the last response's output, never written. Such a response is billed as recorded. Each row now counts its snapshot responses, and a summary line names those rows' output, and their cost, a floor. An absent key is no evidence and counts as none. agent-cost-fixture.sh adds vvv, one stopped response and one snapshot, red first. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es the fork holes a delta check found A refute-by-default check of 6dfcade and c5a3f26 (three read-only verifiers) found five defects. Each is now a fixture case that went red first: - A fork's copy of its parent's history carries the parent's final usage; the parent's own last line is often a mid-stream snapshot (31 of 33 and 9 of 9 on two real parents). 6dfcade's premise, "same message ids and usage", was wrong. Each response is now billed at its best copy across the directory (a stopped one, then the larger output), and the floor counts only responses with no stopped copy anywhere. A parent with a snapshot of 8 output tokens is now billed for the 100,000 its fork copied. - History a fork excluded (a cut orphan's prefix, a main-loop fork's session responses) never entered the billed set, so a fork of that fork billed it again; a fork of an unseparable fork billed its parent's work. Excluded history is now accounted for, and unseparability passes to children. - The boundary was the first user line containing the marker anywhere, so an inherited grep result quoting it cut too early. It is now structural: a text block opening with `<fork-boilerplate>`, beside the tool result for the meta's `toolUseId` where there is one, which is the real two-block shape. bp() now emits that shape. - Every row, forks or not, was timed and labelled from after its last foreign response; a non-fork sharing one response with another lost both. Only a fork now starts after its copied history, and a cut fork starts at its cut. - The fork-of-fork case sorted after its parent by name, so depth order went untested; it now sorts before. The docstring now says copies are same-uuid lines carrying final usage, not byte copies. It also notes that the main-loop fork path is modelled only by the fixture, and that an id-less response is the one exception to billing once. On the real fork-bearing directories the per-file overcount is now $2.48, $4.28 and $5.40 (a fifth of $26.08). Corrections to the earlier commit messages: a snapshot's output is usually under ten tokens, not "2 to 8", and `subagent_tokens` runs 2k–11k past the final context, not 4k–10k. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…riginal transcripts orchestration-reference.md § Applied instances carried cost figures the old reader had inflated, noted "re-run before reusing". They are now re-derived on 2026-10-01 from the three runs' transcripts (Claude Code 2.1.258 and 2.1.259, which carry final usage on all but a few responses), and an independent reader recomputed every number: - framing A/B: output at most 1.6% of a drafter's tokens (was "under 0.5%"); the top tier 73% cache writes (78%); 24 against 65 turns pooled (34 vs 52); top:workhorse 1.48x pooled, 1.72x and 1.28x by arm (3.2x). The transcripts settle the model: claude-fable-5-1, whose cache reads bill at 0.025x. The top-tier procedure drafter read 3.6x the contract one's cached tokens (2.7x), 1.3x at the workhorse tier. - rough-in A/B: the xhigh judge cost 8% more than the high judges' mean (15%). - /finish A/B: executors $32.54 and $33.01 (about $50), judges $9.12–$11.81 ($18–23). On these runs the per-line count was the whole error: the recorded figures ran 1.5x to about 2.2x high. The 1-hour-write and fork defects fixed the same day did not touch them, and the note says so. The CHANGELOG sync note gives the same range and the floor on current transcripts, and its orchestration hand-merge row says to take the kit's re-derived observations. The dated planning copies under docs/superpowers/ stay as written. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
j4th
marked this pull request as ready for review
October 2, 2026 03:24
j4th
added a commit
that referenced
this pull request
Oct 2, 2026
The heading carried the date the branch was planned on. CLAUDE.md dates a release by its merge in UTC: PR #75 merged at 2026-10-02T03:24:51Z (gh pr view 75, mergedAt), so the heading reads 2026-10-02. The PR number it predicted, #75, was right. The five back-dated headings already match their merges' UTC dates. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #58
Closes #60
Closes #61
Closes #62
Closes #63
Closes #64
Closes #65
Closes #66
Closes #67
Closes #68
Closes #69
Closes #70
Closes #71
Closes #72
Closes #73
Closes #74
Harvest 5 lands every issue the three sister repositories filed since v0.5.0, plus a review of the whole kit, as v1.0.0, the kit's first release. The design is
docs/superpowers/specs/2026-09-30-cascade-kit-harvest-5-design.md, and the trace beside it has one row per ask; the plans are underdocs/superpowers/plans/. Tags and the GitHub Release follow the merge, and only on the operator's go-ahead (plan F5).Summary by cluster
verification: project sub-block complete. The Stop-hook check reds a hook that could not look. Blueprint's tooling template wires the block intocheck. The budget line warns above 140,000 bytes.lib/resolve-path.shresolves a path lexically and physically, and the ADR guard is rebuilt on it. Guards refuse a payloadjqcannot read. The Bash guards see throughbash -c, quoting, full paths and chained verbs. The ADR CI job uses:(glob)and three-dot, and fails closed. Three behavioural hook fixtures.arguments:anddisable-model-invocation: true; registrations are exec form; plugin keys are qualified; report and comment text is data.--effortper branch; a step records the model that ran; the rebuilt assertion with a fixture; apaths:trigger that reviews rule and ADR markdown, with a fixture; a fork guard; job-level concurrency.knowledge-backend.mdsplit into a contract and a path-scoped reference.finish-ab.jsruns two to four arms under a balanced judge panel,run-arms-headless.pyruns each arm headless in its own worktree, andagent-cost.pyis keyed by model version.## Triage — round Nappend #64; D46).review-sweep.jsdeduplicates on file and line, takes finders that carry their own prompt, and keeps every agent read-only..gitignoreharness block, kit code kept out of a target's formatters, Dependabot's real coverage, mise tasks run underpipefail, and a rebuilt.mcp.json.example.CHANGELOG.mdwith back-dated v0.1.0–v0.5.0, the tagged-archive install, § Syncing the kit in versions, and the trace rule inCLAUDE.md.## Review gateand## Triagebelow.Decisions (settled with the operator, 2026-09-29/30)
best, effort explicit,ANTHROPIC_MODELnever set. D48 Apaths:trigger whose re-includes of.claude/**anddocs/adr/**come last. D49 The Opus 5.5 / Sonnet 5.5 recalibration. D50 Record the always-loaded bytes; warn above 140,000. D51 The block as a gate. D52 Target hygiene. D53 Kit issues cited ascontext-builder-kit#N. D54 A PR that closes an issue carries its key in the branch. D55 The frozen-corpus backstop's bypass list. D56 The whole-kit review's 68 findings join v1.0.0, the unverified re-checked at their commit. D57disable-model-invocationon the five commands. D58.mcp.json.examplerebuilt (${VAR},type: "http", no GitHub server). D59 Splitknowledge-backend.mdto pay for the recalibration (target: no net growth). D60 Install the drop-in set from a tagged archive; upgrade withgit merge-file. D61 An unreadable payload denies (hard-deny) or asks (ask-gate). D62 Bracket-split the private target's name at HEAD.Always-loaded rules — before and after
cbk-conventions.mdknowledge-backend.mdorchestration.mdpr-review.mdsimplification.mdtooling.mdworkflows.mdD59's target was no net growth, and it is missed by 2,175 bytes (1.7%). The
knowledge-backend.mdsplit freed 10,823 bytes. The recalibration's per-model and per-surface effort facts and the review round's corrections took more than that; no operative text was cut to make the number. The total sits under the 140,000-byte warning line.Probes (spec § Probes at execution)
cwdfollows the Bash tool'scd(Claude Code 2.1.286, 2026-09-30, variant B).require-repo-root-for-agents.shnow denies a dispatch made after acd, andprotect-main-branch.shjudges the checkout the shell is in.pr-review.md§ The floor says an unattended floor that must fan out is a headlessclaude -psession.unshare -rmon ubuntu-24.04: denied, so the three bind-mount cases SKIP in CI (47 cases ok there, 50 locally where they run), https://github.com/j4th/context-builder-kit/actions/runs/36953707599 — recorded, dated, in the fixture's header (e5f4090).claude mcp list(2.1.286): Linear and Notion answer 401 (auth required, so the endpoints are live), context7 200;mcp-server-time2026.8.18 is the newest release and past the settle window.grep -F, and found every one verbatim. The four claims it could not verify were re-read at their sources by four verifiers on 2026-10-01 and corrected (6856188).Review gate
/simplify— ran: 4 cleanup agents, 28 findings (13 after dedup), 4 applied (c4d8c36, cd8a8ff, 576d697, 6b954a2) · the agents ran on Sonnet at the session's effort (xhigh); the Agent tool takes no effortpr-review-toolkit:review-pr— ran: 5 agents (code-reviewer, pr-test-analyzer, comment-analyzer, silent-failure-hunter, type-design-analyzer; code-simplifier left to/simplify), 78 findings, triaged 63/9/4/0/2 · the agents set no model or effort, so they ran on the session's Opus 5.5 at its effort,xhigh(orchestration.md§ The effort axis)review-sweep— ran: 11 finders + 8 verifiers, 6 confirmed / 2 refuted / 36 unverified, dropped coverage: none, bounds 3/8 · the changed-file list leaves outdocs/superpowers/(the planning record, not shipped); every unverified finding was triaged at the caller (below)high; all four corrected, 6856188). One refute-by-default verifier checked the trace's two non-landed rows, and both held. On the operator's follow-up (below), one more verification workflow checked that delta: 3 read-only verifiers, workhorse tier,high(the reader's code, an independent recomputation of every re-derived figure, every other claim). It found 5 code defects and 2 wording errors, all applied in a037350 with red-first cases, and confirmed every figure but one range, which was corrected.Triage
Toolkit — 78 findings: Apply 63 · Apply with care 9 · Surface 4 · Defer 0 · Reject 2. Sweep — 44: Apply 28 · Apply with care 3 · Surface 1 · Reject 12. Every applied finding is its own commit, red first where it is behavioural (the red-first table below), with the block green after each.
Applied, by area
systemMessage) · 9a366a8 (backslash-newline continuations) · 349619e (an unborn branch's root commit) · d752a79 (the ADR guard's two path readings, each pinned) · 7f1dade (matchers checked against their tools) · 0e2e324 (bun.lockb,npm-shrinkwrap.json; the ask-gate's cases).[1m]only) · 2bed51b (run-arms-headless: strict config, arm crash recorded,ghdeny forms,CLAUDE*provider and auth kept, worktree-add failure handled) · 5684420 (finish-ab: arguments, efforts, anon charset, the executed cell, judge scores) · 533ece4 (review-sweep: arguments, roster, key collisions, path normalization).claude.ymlgates in the jobif:) · 83408c6 (a per-path dispatch rule) · e07bed1 (a bold verdict marker,**SKIPPED**, the moved-base notice)..mcp.jsonjq status) · 700d321 (a misspelled rule citation) · 9428a63 (:(glob)case; bounds pinned by behaviour) · e6da512 (pointer claims on all four halves) · c6d7367 (Linear taxonomy) · 8063291 (the axis name).rust-toolchain.tomlcoverage) · c098399 (CLAUDE.md) · 6856188 (four platform claims) · 9a31317 (the action-release citation, bump timing) · ae489c8 (README: install, upgrading,docs/, the ADR opt-out) · 6ba4971 (README example) · 7147791 (CHANGELOG).systemMessagechannel across six hooks (55a89d2), the arm-crash wrapper (2bed51b), the two template sections (6b36fdf), and the README install rewrite (ae489c8, which a scratch-repository run exercised).Surface — not applied; the reviewer's case, verbatim in substance
..paths matters once a target fills a formatter arm..claude/(code-reviewer Review gate: the orchestrated sweep is framed as a substitute for direct dispatch, contradicting /finish's own "does not skip" rule #8). The README's own illustration is a taste call.worktree_root"must be gitignored" is stated but never checked; themcp_configallowlist checks server names only; and the runner's facts are appended to the arm's notes rather than kept in a structuredresult.runner(type-design No proportionality rule scaling the review mechanism to diff size — and session-level effort silently overrides judgement #12, No rules file carriespaths:frontmatter and the kit never mentions the mechanism — every rule loads every session #15).docs/adrsubdirectory when an ADR there shares its basename (sweep).Reject
Not seen:line now names it.resolve-path.shfails a hard-deny guard open: the contract (§ Hook authoring): a missing dependency fails open and names its backstop.~infile_pathis taken literally: the harness sends absolute paths;~is a shell expansion.--fallback-model sonnetequals its model: CLI 2.1.287 accepts the pair (exit 0), and the action passes the flag as a CLI argument, never as the SDK option that refuses it./simplify skipped (behaviour or scope): every verification-block cleanup, since
simplification.mdkeeps the pass out of.claude/rules/(a follow-up issue candidate); a shared fixture library, which would be a new drop-in file with a sync note; the review template's output consolidation; review-sweep'sreports[]interface; parallelworktree_setup; the lock-file guard on the resolver, which is a behaviour change (a follow-up candidate).Operator-directed follow-up — "fix the agent-cost double counting too" (2026-10-01)
The per-line double count was already fixed (ec17e0a). Two more defects turned up on this machine's transcripts:
A second double count: forks (6dfcade, a037350). A fork's transcript opens with copies of its parent's history, same message ids, and the reader deduplicated only within a file. So a parent's spend was billed again for every fork: $2.48, $4.28 and $5.40 on the three fork-bearing session directories, the last a fifth of the real total. Each response is now billed once per directory, at its best copy. Forks are read after their parents at every level. A fork whose parent is missing is cut at the message that hands it its task, and one that cannot be cut is unpriced and named. Fork copies carry the parent's final usage where the parent's own line is a mid-stream snapshot, so parents are now billed for output the old reader never saw.
An undercount, named rather than hidden (c5a3f26). From Claude Code 2.1.278 most subagent responses are written mid-stream (
"stop_reason": null, output usually under ten tokens), and nothing else in the transcript records the final count. Such rows are billed as recorded and named a floor.The kit's dated cost figures, re-derived (eaa5cb8). The A/B runs behind
orchestration-reference.md§ Applied instances were recomputed from their original transcripts, and the figures replaced:/finishexecutors/finishjudgesOn those runs the per-line count was the whole error: the recorded figures ran 1.5× to about 2.2× high.
Trace
371 rows: 369 landed, 1 not holding (
#69/c5859811353/H1—--disallowedToolsis variadic), 1 non-goal (#58/c5901493591/residue-2— D31's fail-fast stands), 0 open. Every landed row was audited against the tree at HEAD. The five the audit found partial carry the review round's completing commit, andrelease/7's tags are pushed after merge. A refute-by-default verifier could not show either non-landed row required.Red-first table (130 rows)
hig[h].claude/rules/orchestration.mdtriage(a flow applies it)\.github/issue template[s]|with five HITL gate[s] .claude/skills/scaffold/shows)git \-continuedcommitand a continuedgh pr readypassed both Bash guardsbun.lockbandnpm-shrinkwrap.jsonpassed; the KB ask-gate's cases could be switched off silently"headers": "Authorization: Bearer ghp_…"→ exit 0cbk-convention.md§ ADR index sync → exit 0workstream:<slug>Sync notes preview (CHANGELOG [1.0.0])
Sync notes
.github/workflows/claude-review.ymland.github/workflows/claude.yml, your copies of blueprint's review templates. Keep your label names, turn cap, sizing, allowed tools and the filledREVIEW_LOGIN. Take the model and effort lines (claude.ymlnow runs--effort high), theRecord the resolved modelstep, the rebuilt "Assert the review posted" step, the fork guard,concurrencymoved into the job,defaults: run: shell: bash,runs-on: ubuntu-24.04, the step idsreviewandclaude, the prompt's restore section with its two new bracketed fills, the notes aboveclaude_args, the label step'sdispatch_ruleoutput with the two prompt lines that read it, the assertion's bold verdict marker and the prompt's**SKIPPED**docs-only line, and inclaude.ymlthe bot andskip-claudegates moved into the jobif:(the gate step is gone). Thepaths:filter replacespaths-ignore: put your own prose-only negations above the re-includes of.claude/**anddocs/adr/**. The kit sub-block now runsreview-trigger-fixture.pyandreview-assert-fixture.shagainst your filled.github/workflows/claude-review.yml, so a target that syncs the rules before merging this workflow is red until it does..github/workflows/ci.yml). It was never a kit copy, so there is nothing to merge: edit it to match the skeleton. That meansdefaults: run: shell: bash,runs-on: ubuntu-24.04, and everyuses:pinned to a full commit SHA with a trailing version comment, in place of a tag such asactions/checkout@v4..mcp.json, your copy of.mcp.json.example. Keep the servers your axes use. The kit ships no GitHub entry, hosted entries carry"type": "http", andtimeruns pinned throughuvx. Turn every literal credential into a${VAR}reference, export the variable before launchingclaude, and commit the file. The kit sub-block now checks the committed file's shape: a url entry without a type, an unpinnednpx,uvxorbunxserver, or a literal credential turns it red..claude/rules/orchestration.mdand.claude/rules/orchestration-reference.md. Keep your posture row and your dated applied instances; take the kit's text for its own kit-shipped observations, whose cost figures are re-derived. Re-point any citation of§ The three surfaces + resolution orderto§ The dispatch surfaces + resolution order..claude/rules/tooling.md. Keep your filled stack sections; take § MCP configuration's merge rule and § Automated review on the git host..claude/settings.json. Keep your own hooks' registrations, rewritten in exec form, and re-add their names and tiers to_comment_hooks. Then diff the hook event keys against your pre-merge copy (§ Syncing the kit)..claude/rules/cbk-conventions.mdand.claude/rules/cbk-conventions-reference.md. Keep your filled values and your project sub-block; take everything in the block above it. Incbk-conventions.md, take § Branch naming's rule that any issue a PR closes, cascade or not, puts its key in the branch (the bare issue number on thegithub-issuesaxis;<type>/<short-slug>only for work no issue tracks) with its Quick reference row, and the reworded substring trap in the CI-skip rule's section, which names all five skip tokens and GitHub's skip trailer and cites GitHub's page. The kit sub-block pins both, so a copy that keeps the old wording is red..claude/rules/pr-review.mdand.claude/rules/pr-review-reference.md. Keep your roster entries; take the floor, the sweep and the invariants..claude/rules/simplification.mdand.claude/rules/workflows.md. Take the kit's text; keep a filled § Cost+scope-explicit..claude/rules/knowledge-backend.md. It is now two halves. On thenotionaxis, take both,knowledge-backend.mdand.claude/rules/knowledge-backend-reference.md, and move each filled value under its heading. On thenoneaxis, where the rule was deleted, add neither half..claude/commands/finish-procedure.mdand.claude/skills/rough-in/references/finish-procedure.md. Their citations now point at the conventions: a copy whoseCONTRIBUTING.mdcitation was filled by hand takes the kit's text..claude/skills/adr-new/SKILL.md. Take the kit's text: it reads the index's rows in both forms, and its## Refines vs Supersedes vs Extendssection is now## Relation grains, a pointer to.claude/rules/cbk-conventions-reference.md§ ADR relation grains. Repoint any citation of "adr-new § Refines vs Supersedes" in your filled rules, reviewers or docs atadr-new§ Relation grains or § ADR relation grains; an ADR that cites it stays as written, and the correction goes indocs/adr/corrections.md..claude/hooks/format-on-edit.shand.claude/hooks/analyze-on-edit.sh. Keep your filled case arms; take the exec-formRegister:stanza and the exit 2 on a formatter failure or an analyzer error (exit 0 hid it in the debug log), and informat-on-edit.shthe skip floor, which now names.claude/workflows/(an analyzer only reads, soanalyze-on-edit.shhas no floor)..github/workflows/adr-immutability-check.yml. Re-copy it from the kit, then restore only your comments and the pinned job name: the job body is replaced. It reads a:(glob)pathspec, diffs three-dot from the merge base with--no-renames, checks out blobless, runs undershell: bashon a named runner image, and fails closed on any git error..claude/workflows/agent-cost.py,.claude/workflows/tests/agent-cost-fixture.sh,.claude/workflows/finish-ab/finish-ab.js,.claude/workflows/tests/finish-ab-shape.mjs,.claude/workflows/finish-ab/run-arms-headless.py,.claude/workflows/tests/run-arms-headless-fixture.sh,.claude/workflows/review-sweep.js,.claude/workflows/tests/review-sweep-accounting.mjs,.claude/hooks/protect-immutable-adrs.shand.claude/hooks/lib/resolve-path.sh. For each one your target already carries, diff your copy against v1.0.0's. Where the kit's application differs from yours, take the kit's, and keep only your project's own fills. From this sync on, v1.0.0 is these files' merge base, not your install (§ Syncing the kit)..claude/hooks/lib/resolve-path.sh, which is sourced, never run. A.gitignorewith an unanchoredlib/line hides it fromgit addwithout a word: add!/.claude/hooks/lib/after that line, and after the sync commits check thatgit ls-files .claude/hooks/lib/names the helper (if it does not,git check-ignore -v .claude/hooks/lib/resolve-path.shnames the line that hides it). Add the new fixtures under.claude/workflows/tests/, each anaddrow: the three hook fixturesprotected-paths-hook-fixture.sh,hook-guards-fixture.shandhook-payloads-fixture.sh;adr-ci-body-fixture.sh;review-assert-fixture.shandreview-trigger-fixture.py;run-verification-block-fixture.sh;extract-run-block.sh, which the ADR and review-assert fixtures source; and, unless your target already runs them ahead (above),.claude/workflows/finish-ab/run-arms-headless.pyandrun-arms-headless-fixture.sh, which are required, because the block runs the fixture and the fixture fails without the runner. The block runs every one of them but the sourced helper. Append the harness block from scaffold'sreferences/github-starter-templates.md§.gitignoreto your.gitignore, below every stack section, and state its pin assertions in the commit body; the kit's own.gitignoreis not in the drop-in set..claude/workflows/**from every repo-wide formatter and linter, forced for explicit paths (§ Syncing the kit). A wiredformat-on-edit.shmerges the new skip-floor line.pipefail. A target on mise adds the[task_config]block from blueprint'stemplates/tooling.mdstep 1 to itsmise.tomland pins mise ≥ 2026.7.15 wherever tasks run, CI's setup action included. Nothing in the kit checks a target'smise.toml.rust-toolchain.tomlis covered by Dependabot'srust-toolchainecosystem, and so is a container base-image tag; a target whose filled § Dependency settle-window lists either as uncovered corrects that section, which now names the three image gaps.agent-cost.py, which reads them at the standard 0.1x. A target that runsrun-arms-headless.pysets its config'sworktree_setupto its own per-worktree step, and adds any write-capable MCP server it wires beyond the kit'sgithub,linearandnotiontoWRITE_SERVERS, which makes the file amergerow in its sync table..claude/workflows/tests/run-verification-block.shis acopyrow with a third rail: it fails a target whose output lacksverification: project sub-block complete..claude/workflows/tests/run-verification-block-fixture.sh, which the block runs against the runner, is anaddrow. Wire the runner as a verification task yourcheckdepends on, per blueprint'stemplates/tooling.md. If you never stamped the bracketed manifest-and-lockfile globs incbk-conventions-reference.md'spaths:, stamp them now (the bootstrap checklist's disposition pass has the row), or the project sub-block is red. The block'sCLAUDE.mdcheck now fires in a target only oncedocs/cbk/blueprint.mdexists: from then on yourCLAUDE.mdmust mentioncbk-conventionsas a backticked path, never an@import.[the project's CI lockfile check — …]inprotect-lock-files.sh, and[the base branch's ruleset — pull requests only — where one exists]in both warnings ofprotect-main-branch.sh. The bootstrap checklist's disposition pass gains a Hook backstop slots row: replace each bracket with the real check or ruleset, or withnoneand the reason. The project sub-block refuses a slot left bracketed. A filled slot makes those two hooksmergerows in your sync table./finish,/intake,/enrich,/pr-respondandfinish-procedurecarrydisable-model-invocation: true: type them.context-builder-kit#N, because a bare number links to the target's own issue. The rewrite touches about thirteen files a target copies byte for byte, among them the hooks, the workflow scripts and their tests, the executor pair andadr-new: take the kit's side of every such hunk. A copy you patched by hand takes the kit's form. The kit sub-block's bare-citation check runs on the kit's own tree only; in yours a bare number is your own issue..github/ISSUE_TEMPLATE/cascade-meta.mdand.github/ISSUE_TEMPLATE/cascade-rough-in.mdare byte copies of scaffold'sreferences/issue-templates/, and both changed: re-copy them. The project sub-block is red while yourcascade-meta.mdstill cites the retired § Deferred meta-issues.jqcannot parse, and an ask-gate asks. A missingjqstill fails open, naming its backstop.protect-immutable-adrs.shsourceslib/resolve-path.sh, resolves the path lexically and physically, and denies an Edit, Write or MultiEdit to an existing numbered ADR in any checkout, whatever spelling, relative path or symlink names it; itsNot seen:line lists what it cannot see (a Bash write, a case-insensitive filesystem, a hand edit, a symlink swapped between check and write), which the ADR CI job refuses at the pull request. A target that customized its ADR hook takes the kit's hook (acopyrow) plus the helper (anaddrow), carries anything it still needs as a named exception (§ Syncing the kit), and checks the helper is tracked (New files, above).cdis judged where the shell is. A hook payload'scwdfollows the Bash tool'scd(the probe is recorded inrequire-repo-root-for-agents.sh'sTiming:paragraph). So the launch-root guard now denies an Agent, Task or Workflow dispatch made after acdinto a subdirectory, andprotect-main-branch.shjudges a commit against the checkout the shell is in. Return to the repository root, as its own command, before dispatching.anthropics/claude-code-actionrelease's Claude Code resolves them to.opusresolves to Opus 5.5 on the Anthropic API from Claude Code v2.1.280, andsonnetto Sonnet 5.5 from v2.1.284 (https://code.claude.com/docs/en/model-config§ Version history, read 2026-09-30). The action installs Claude Code 2.1.280 from its v1.0.232 and 2.1.284 from its v1.0.236 (src/entrypoints/run.ts,claudeCodeVersion, read 2026-09-30). A workflow pinned below v1.0.232 still reviews on the previous model underopus(Opus 5 from Claude Code v2.1.219), and below v1.0.236 on Sonnet 5 undersonnet: bump the pinned SHA.systemMessageon stdout as well as on stderr, which on exit 0 reaches only the debug log; the Stop hook gathers its notes the same way. The guards arecopyrows: take the kit's. A hook of your own follows § Hook authoring's fail-open bullet. Beside it, the main-branch guard allows the root commit of an unborn branch and reads a command continued with a backslash-newline, and the lock-file guard namesnpm-shrinkwrap.jsonandbun.lockb./finish,/intakeand framing look a workstream slug up there.blueprint.mdis immutable, so append the table — workstream, slug, layer — under its append-only## Amendments, from the slugs its slug gate confirmed. On thein-repo-markdownaxis, setup progress belongs on the roadmap's step-0 row, never in commits toblueprint.md.area:<slug>label per workstream; every Linear flow appliesworkstream:<slug>. Rename them, and create theenhancement,source:<name>,triageandappetite:small/medium/biglabels the flows apply.run-arms-headless.pyrefuses a config with an unknown key, a budget or continuation count out of range, or a dry run combined with resume, and records each arm inexecuted.jsonas{anon, cell, result}.finish-ab.jsrefuses an unknown argument or effort and an executed record whose cell does not match, and drops a judge whose scores do not cover every arm.review-sweep.jsrefuses malformed arguments: a list that is not a list, a bound that is not a whole number, averifyModelother thanopus,sonnetorhaiku. A saved config or call that relied on the old leniency fails loudly; fix it rather than the check.agent-cost.pysummed every transcript line, and Claude Code writes one response as several; it billed a fork again for the parent history its transcript opens with; and it priced every cache write at the 5-minute rate, where a 1-hour write bills at 2x. The kit's own dated figures were re-derived from their original transcripts: they ran 1.5x to about 2.2x high, all from the per-line count (orchestration-reference.md§ Applied instances). Re-run a figure of your own on its original transcripts before comparing it across releases. From Claude Code 2.1.278 a subagent's responses are mostly recorded mid-stream, so a figure read from a current transcript is a floor on output and cost; the reader names the rows it applies to.## Review gateblock has a "not run" form. An issue-less branch writes its/simplifyand toolkit linesnot run — <reason>; a cascade PR's docs-only diff still runs the floor (pr-review.md§ What NOT to flag).🤖 Generated with Claude Code