Skip to content

[harvest] Applying e92e9c4 to a target project: 11 defects and 6 smaller items, four with POC patches (you-are-hear #42 / PR #43) #58

Description

@j4th

Harvest report from the first target-project application of e92e9c4: you-are-hear issue #42, landed as PR #43 (chore/42-kit-sync-harvest-3, 67 commits, the kit's verification block adopted wholesale). 63 files were taken byte-identical, 38 merged, and the block now runs green there with two recorded exemptions (that project wires and registers the two advisory hooks instead of shipping them as exemplars).

The defects below were found by applying the kit, not by reading it. Four carry a POC patch already exercised in that repository; the rest are reproductions with a suggested fix. Every item is independently filable — say the word and I will split them into one issue each, as with #48–#53.

Update, 2026-09-07 — read the follow-up comment before acting on this body. PR #43 merged as 3f7548a. The flip's auto-review found two more kit-side defects (the verification block silently loses its back half on any project that wires the advisory hooks as § Hook authoring instructs; the Stop hook blocks the auto-review's own hand-off on claude-code-action's .claude-pr/ staging copy) and invalidated one claim in item 4 below — invoking a fixture from the project sub-block does not make it run, for the same reason. Item 2's POC also needs a writability probe before it is copied. The comment carries all of it with the evidence.

Confirmed, with a POC patch

1. detect-forked-agent-memory.sh: the hand-kept prune list is unmaintainable, and every ignore-driven replacement has three traps

The shipped list is -o -name node_modules -o -name target -o -name build -o -name _build -o -name dist -o -name .venv. On a Dart/Flutter monorepo it missed .dart_tool (the tool cache every package has) and carried five names that stack does not use. Reading the project's own ignore rules instead is the obvious fix, and it is where the traps are:

  • The prune list is spliced into find -path, which is a glob. An ignored directory named * (a legal Unix name) becomes -path './*', and * in -path crosses /, so it prunes the first top-level entry the walk reaches and the scan returns clean with empty stderr. Reproduced with an A/B control: with a genuine stray packages/foo/agent-memory present, the hook exits 2; create an ignored directory named * and it exits 0; remove it and it exits 2 again.
  • A project can ignore the tree the hook hunts for. agent-memory/ unanchored is the natural spelling for the local scope, and some projects ignore .claude/ wholesale. Either makes the hook's own skip list hide the fork.
  • --directory collapses a wholly-ignored subtree into its topmost ancestor, and in a repository with nothing committed yet that ancestor is packages/, which no rule names. Pruning it skips most of the tree.

POC (you-are-hear a83069d and fa6b0d3):

prunes=()
while IFS= read -r d; do
  d="${d%/}"
  [ -n "$d" ] || continue
  case "${d##*/}" in .claude|agent-memory|agent-memory-local) continue ;; esac
  prunes+=(-o -path "./$(printf '%s' "$d" | sed 's/[][*?\\]/\\&/g')")
done < <(
  git ls-files --others --ignored --exclude-standard --directory 2>/dev/null \
    | sed -n 's:/$::p' | git check-ignore --stdin 2>/dev/null || true
)

Seven scenarios, each in a fresh throwaway tree: stray tree 2; ignored directory named * 2 (0 before); unanchored agent-memory/ 2; .claude/ ignored wholesale 2; collapsed ancestor 2; stray under an explicitly ignored build directory 0 (pruned for speed, and not a directory agents are dispatched from); clean tree 0.

2. detect-forked-agent-memory.sh: find … 2>/dev/null turns a partial scan into a clean verdict

If find cannot read a subdirectory, its stderr is discarded and its exit status is never checked, so a fork under that directory is missed and the hook reports clean. Same class as the silent miss above. POC: capture stderr to a temp file and warn when it is non-empty, naming what could not be read (you-are-hear a83069d).

3. protect-lock-files.sh has no pubspec.lock arm

cbk-conventions-reference.md § Dependency settle-window says hand-edits to lock files are blocked by this hook, and the Dart/Flutter ecosystem's lock file is not in the case arm, so on that stack the rule was prose with no mechanism behind it. POC (6d05234), three lines:

-  uv.lock|pnpm-lock.yaml|package-lock.json|yarn.lock|Cargo.lock|Gemfile.lock|poetry.lock|composer.lock|mix.lock)
+  uv.lock|pnpm-lock.yaml|package-lock.json|yarn.lock|Cargo.lock|Gemfile.lock|poetry.lock|composer.lock|mix.lock|pubspec.lock)
+  - pubspec.lock         →  dart pub get / flutter pub get  (pub upgrade to move a resolution)

Worth considering a *.lock fallback arm with a generic remediation, so the next ecosystem is covered by construction rather than by an edit.

4. § Hook authoring puts the dry-run table in the PR body, so hook coverage is prose that nothing re-runs

The kit ships assertion-driven fixtures for agent-cost.py, review-sweep.js and finish-ab.js, and for the two highest-stakes surfaces (a HARD-DENY guard and a Stop-tier hook) it asks for a table of expected and observed exits in the PR body. On a real run that table is written once and never re-checked; three of the guards' branches turned out to have never been exercised at all.

POC: hook-payloads-fixture.sh (you-are-hear c55e964, 28494c3), invoked from the project sub-block of § Verification so it runs on every block run. It asserts, against throwaway git init trees under mktemp, every branch both headers document: deny for Agent, the legacy Task and Workflow; allow at the root and for a non-matching tool; the symlinked and trailing-slash spellings of the root; the cwd fallback from a controlled directory in both directions; the three fail-open paths (no checkout, no jq, no sourced helper); and for the detector, both memory scopes, the second stop under stop_hook_active, the ignore-driven prune and its three traps, the worktree exemption, a subdirectory project dir, no jq, a non-checkout, and a partial scan. Ten mutants fail it, including two that had slipped through an earlier version of the same fixture.

Two notes for the kit's own copy: state which branches are asserted by mutation rather than by payload (the guard's -z race branch is not payload-reachable), and never let a case depend on the directory the fixture is launched from. The first version of this fixture failed misleadingly when run from a subdirectory, for exactly that reason.

Confirmed, no patch offered (kit-side design calls)

5. The verification block's advisory-exemplar arm contradicts § Hook authoring for any project that wires them

§ Hook authoring says a project copies the advisory hooks, wires its analyzer and formatter into the case arms, and registers them. The block's registry loop then treats registration as a hard failure: case "$b" in format-on-edit.sh|analyze-on-edit.sh) [ "$n" -eq 0 ] || { echo "advisory exemplar $b is registered"; exit 1; }. A project that follows the instructions gets a permanently red block, which the block's own preamble calls the worst outcome. On the kit's tree the assertion is correct, so the repair is to make the arm project-conditional or move it into the project sub-block rather than flip the comparison.

6. agent-cost.py prices Fable 5.1 cache reads at 0.1x

cost += … + t['cr'] * pi * 0.10 + …, and the docstring says "cache reads at 0.1x input". The pricing page says: "Cache hits and refreshes on Claude Fable 5.1 and Claude Mythos 5.1 are priced at 0.025x the base input price. All other models use the standard 0.1x multiplier" (platform.claude.com/docs/en/build-with-claude/prompt-caching § Pricing, read 2026-09-07). Every Fable figure the script produced is therefore an upper bound, including the ratios quoted in orchestration-reference.md § Applied instances. The fixture exercises Opus and Sonnet only, both on the standard multiplier, so nothing catches it.

7. review-sweep.js: the path-hint match loses the directory boundary

d.pathHints are stripped of trailing slashes and matched with f.startsWith(h), so a hint of packages/yah_core matches a changed file in packages/yah_core_extra/. Demonstrated in node: startsWith true, f === h || f.startsWith(h + "/") false.

8. review-sweep.js: the roster agent() is the one call outside a parallel() thunk

agent() throws when a budget ceiling is reached. Every other call in the file is inside a thunk, where the runtime's catch turns a throw into null and the file's degrade path takes over. A throw at the roster call ends the run with no droppedCoverage, no gateLine and no partial result, which contradicts the file's own stated degrade-never-fail design.

9. finish-ab.js: arm labels are not validated for uniqueness, and hallucinations is dereferenced unguarded

anon ids are checked for uniqueness before dispatch and arm labels are not, although otherFile() depends on them differing; two arms labelled A throw inside the prompt builder after paying for the dispatch. Separately, wellFormed validates ranking but the flag accounting then reads j.hallucinations.filter(...) with no ?? [], unlike the arm-side code which documents "a result missing the fields is dropped and named, never dereferenced".

10. The test harnesses' parallel() stubs swallow every exception

try { return await t() } catch { return null } in review-sweep-accounting.mjs and t().catch(() => null) in finish-ab-shape.mjs convert a bug in the test's own mock into the "agent returned nothing" path. Demonstrated by a verifier: a scenario with one typo in its mock passes all three of its assertions while its gate line reports zero verifiers. The runtime resolves a failed agent to null rather than throwing, so the catch models no real failure mode.

11. The block's launch-root dry-runs discard the hook's own diagnostic

Both payload runs redirect stdout and stderr to /dev/null, so when the assertion fails the block prints its generic message and the hook's reason is gone. The neighbouring Stop-hook check deliberately keeps stderr and says so in its failure message.

Smaller items

  • /finish item 1 admits no meta-issue title. It accepts [<slug>:F<#>:R<#>], [<slug>:bug] and [<slug>:enh] with <slug> a workstream locked in the blueprint. cbk-conventions-reference.md § Contribution intake prescribes [<meta-tag>:…] for a cascade or tooling gap, and § Title-prefix scheme defines [<slug>:meta] and [<slug>:<meta-tag>:R<#>]. A tooling meta-issue therefore fails the executor's first precondition by construction. This report's own source issue is one: [cascade:meta] Sync .claude/ …, proceeded past deliberately at the plan gate.
  • framing/references/procedure.md:107 reads "See rough-in's the rough-in skill's references/plan-mode-prompts.md" — a duplication from the cross-skill citation rewrite. The verification block's reference-existence loop parses the <name> skill's form, so the intended text is "See the rough-in skill's …".
  • The hooks page contradicts the launch-root guard's Timing paragraph. The guard's header records a dated observation that the payload cwd is the session's launch directory and does not follow a shell cd. code.claude.com/docs/en/hooks § Reference scripts by path states the opposite: "cwd follows Claude: the cwd field in the hook's input JSON is the worktree root after Claude enters a worktree, and the new directory after Claude runs cd" (read 2026-09-07). The guard is correct under either reading, since it judges the cwd the payload carries, but the header should carry both and name the disagreement as a re-verify trigger.
  • rough-in-spec-template.md spells the second variant MEASUREMENT / SPIKE VARIANT with spaces around the slash, while consumers looking for the block tend to grep MEASUREMENT/SPIKE VARIANT. Worth settling on one spelling.
  • Test gaps in the shipped harnesses: finish-ab.js's unattributedFlags accounting has no scenario; the least-loaded-owner bound in review-sweep.js is only exercised with two converging reporters; agent-cost.py's minutes=None path, its empty-events transcript and its --json missing-argument exit are not in the fixture.
  • Two small redundancies: agent-cost.py builds a totals dict solely to destructure it, writing the same five-key tuple twice; finish-ab.js evaluates j.effort ?? "high" twice in one object literal. Separately, finish-ab.js and review-sweep.js now carry the same dispatch-retry-reindex shape with two different reindex idioms, one of them O(n²); worth converging and naming the pattern once.

What went right, for calibration

The sync itself was mechanical where the kit intended it to be: 63 files byte-identical, the executor pair re-spliced and byte-parallel in both directions, the block's project sub-block filled without touching the kit sub-block, and the two enumerated exemptions were the only red lines. The ## Review gate block, the bounded sweep with its runtime roster read and its gateLine, and the measurement-variant spec blocks all landed and were used in the same PR that installed them.

Activity

  1. added
    harvestHarvested from a real cascade run
    source:you-are-hearEvidence from the you-are-hear run (GitHub axis)
    on Sep 7, 2026
  2. j4th commented on Sep 7, 2026

    @j4th
    OwnerAuthor

    Update, 2026-09-07: final state, one correction to this report, and two more defects

    you-are-hear PR #43 merged as 3f7548a (72 commits; all four required checks green). The branch is gone, so the SHAs below are on that project's main. Flipping the PR to ready ran claude-review.yml, and the auto-review found two kit-side defects this report missed — one of which invalidates a claim I made in item 4.

    Correction to item 4: invoking a fixture from the project sub-block does not make it run

    Item 4's POC said the hook fixture is "invoked from the project sub-block of § Verification so it runs on every block run". That is false, and for a reason that generalises to every project. The block runs under bash -e. The advisory-exemplar assertion in the hook-registry loop is at line 131 of the 238-line extracted block, and it exits 1 on any project that wires and registers the two advisory hooks — the divergence item 5 describes, and the one § Hook authoring tells projects to make. Everything after line 131 therefore never executes.

    Measured on that repository: a raw bash -e run of the live-extracted block ends on advisory exemplar analyze-on-edit.sh is registered, and hook-payloads-fixture: ok appears zero times in its output.

    What is dead in that window is not marginal. Lines 132 to 238 carry the ${CLAUDE_PROJECT_DIR} placeholder-form check, the four tier tokens, the two-views paragraph check, the kit's own launch-root guard payload dry-runs, the Stop hook's live-tree check, the reviewers' ## Writing memory diff, the Reviewer-agent-memory row check, the P4 conventions checks (Licensing, .gitignore anchoring, the issue-less branch form, the lockfile counter-line), the .github starter and ADR-lint checks, the roadmap and executor term checks, the review-automation template checks, the cascade-events and Amendments checks, the LSP tool check, the framing.md index absence check, the context-budget print, both sentinels, and the entire project sub-block. Roughly 45% of the suite, including the two payload dry-runs the kit added precisely so hook behaviour is asserted on every run.

    So item 5 is not cosmetic and not only about a red line: a project that follows § Hook authoring loses the back half of its verification block silently. Three repair shapes, in the order I would consider them:

    1. Make the exemplar arm project-conditional (or move it into the project sub-block), which is what item 5 already proposed.
    2. Order the block so the project sub-block runs first, so a kit-side known-red never masks project checks.
    3. Collect failures instead of aborting: run each check, record red lines, exit non-zero at the end. This is the only shape where one deliberate exemption cannot hide an unrelated regression, and it fits the block's own stated standard that a permanently red line is a suite nobody runs.

    The local fix taken there was none of those, because the block is the kit's and the project did not want to edit a check out of it: the fixture now has a second runner, a mise task in the project's gate task dependency list, so it fires on every gate run and in CI (916f580). The block keeps its invocation for a fully-neutralized run. If the kit adopts shape 3, that second runner becomes redundant, which is the better outcome.

    New defect: the Stop hook blocks the auto-review's own hand-off in the claude-code-action container

    Reproduced by the reviewer, live, during the review of PR #43 — the run blocked on:

    BLOCKED: a reviewer memory tree exists outside the repository root:
      ./.claude-pr/.claude/agent-memory
    

    .claude-pr/ is claude-code-action's staging copy of the PR branch's tooling: it holds the branch's .claude/ including a copy of the committed root agent-memory/ tree, alongside a working tree checked out at the base. Nothing was dispatched from a subdirectory; there is no fork. The kit's own hook fires (its hard-coded prune list does not name .claude-pr), and the ignore-driven replacement in item 1 fires too unless the project ignores that path — rule 1 does not save it, because the directory needing the prune is .claude-pr, whose basename is none of the three protected names.

    This is kit-wide rather than project-specific: claude-review.yml is the kit's own blueprint template, so any project adopting both the Stop hook and the review workflow gets it. Two costs, worth weighing separately:

    1. The Stop tier blocks the hand-off of every session in that container. stop_hook_active caps it at one block, so a review still completes, but it burns a turn and the block is indistinguishable from a real one.
    2. The remediation text is actively wrong there. It says to "move each <reviewer>/ directory's files into the root tree … delete the forked tree", which, followed literally, deletes the harness's staging directory mid-run. A guard whose false-positive path instructs the agent to delete infrastructure is worse than a guard that misses.

    Two-part fix taken there (39ed311), both worth upstreaming:

    • An anchored /.claude-pr/ entry in .gitignore, after which the ignore-driven prune covers it with no hook change. For the kit that means adding the line to the scaffold's .gitignore starter in github-starter-templates.md, beside the harness-transient entries § .gitignore anchoring already prescribes.
    • The BLOCKED message now says that a path belonging to tooling rather than to a dispatched agent — a staging copy, a container's workspace, a vendored checkout — is not a fork, that nothing should be moved or deleted, and that the route is an anchored ignore entry.

    Correction to item 2's POC: the stderr capture needs a writability probe

    Item 2's POC (capture find's stderr, warn when non-empty) introduces its own silent miss if copied as written. With mktemp absent and the fallback path unwritable, the 2>"$scan_err" redirection fails before find runs: forks stays empty, [ -s "$scan_err" ] is false because the file was never created, and the hook exits 0 on a tree it never scanned. Probe the path first and degrade to /dev/null with a warning that the walk is unmonitored; and skip /dev/null in the cleanup, or every degraded stop also prints rm: cannot remove '/dev/null'. Both were caught by running the fixture from item 4 against the fix (694b8ed, cec420c).

    Related, and cheap: § Hook authoring's header shape asks for Blocked: / Allowed: / Timing: / Path: / Tier: but not for the hook's dependency set. That hook's header claimed "No jq dependency" while the rewrite had added git (the root and the prune list) and mktemp (the stderr capture). A Depends: line naming each dependency and what its absence costs — fail open, degrade unpruned, degrade unmonitored — would have caught it at authoring time.

    New smaller item: the ADR status-cell separator is pinned in two places a real project's index can already contradict

    docs/adr/README.md's starter and cbk-conventions-reference.md § ADR relation grains both pin Accepted · Refines ADR-0007 (D2). A project whose index predates that text uses whatever it used — there, Accepted — refines ADR-0003 (D3) — and adr-new merged from the kit would have written the second spelling into the same table, which is exactly the multi-surface drift § Multi-surface facts exists to prevent. Fixed there on both governing surfaces (36720d7), with adr-new now told to read the index before writing a row rather than assume a separator. For the kit, the smallest change is that instruction in adr-new: the separator is a project's existing convention, not the kit's to pin.

    Everything else in this report stands

    The auto-review checked the round-1 dispositions and the six refuted sweep findings against the code and agreed with them as written, so items 6 to 11 and the smaller items are unchanged. Items 1 and 3's POCs are unchanged and now merged.

  3. j4th commented on Sep 13, 2026

    @j4th
    OwnerAuthor

    Second application of e92e9c4 — echosphere (Linear axis), research pass only, 2026-09-12

    echosphere's ECH-39 is the same issue shape as you-are-hear #42 ([cascade:meta] Sync .claude/ to context-builder-kit harvest 3), authored 2026-09-07 before this report existed. Before executing it I ran this issue — body and follow-up comment — against the kit tree at e92e9c4 and the spec (six surface readers → two adversarial verifiers each → a completeness critic; load-bearing claims then re-run by hand). The run was aborted for a re-rough-in; nothing landed. Items 1–11, the smaller items and the comment's additions all reproduce at e92e9c4 as stated, with two calibrations: the MEASUREMENT / SPIKE VARIANT spelling has one spaced hit (rough-in-spec-template.md:256), no collapsed hit, and no mechanized consumer anywhere in .claude/; and the "empty-events transcript" test gap is asserted by agent-cost-fixture.sh:35 on the natural reading — it holds only for a transcript with zero parseable lines.

    What the second application adds, kit-side:

    1. The verification block's first bash -e abort on a target that does not ship .mcp.json.example is cbk-conventions-reference.md:493, not the advisory arm. absent grep -rn -i "…" .claude/ README.md .mcp.json.example — grep exits 2 on the missing path, absent() reads any non-1 exit as "not a clean miss" and exits 1, at extracted line ~21 of 228. Item 5's arm at :603 is never reached. Green on the kit only because the kit ships the file; the check needs a [ -f .mcp.json.example ] guard or the scaffold must make the file mandatory. A migration note beside it: the retired-vocabulary literal is split in this check (opinionate[d] profile), but the 88b1ede-era project sub-block shipped the un-split form (! grep -rn -i "opinionated profile\|opinionated_profile" .claude/, echosphere cbk-conventions.md:319); a project that re-homes its sub-block into the reference half per the split then trips its own check. Worth one sentence in the re-homing guidance.

    2. The bracket-truncation idiom trips spellcheckers. Under a typos config that includes hidden directories (echosphere's _typos.toml sets ignore-hidden = false deliberately, and forbids blanket excludes), cbk-conventions-reference.md fails with four errors: :506 onl (github-onl[y]), :516 summar, :588 chec (double-chec[k]), :594 doub (doub[t]). A target whose gate spellchecks .claude/ cannot take the file byte-identical and keep its gate green. Two shapes: ship an extend-ignore-re snippet for the idiom, or split on the leading letter ([g]ithub-only, [d]oubt) so the dictionary word stays intact — whether every remainder is clean under typos' dictionary needs a run.

    3. Item 7, sharpened, and a design gap behind it. The roster prompt at review-sweep.js:112 asks for hints "as bare paths (e.g. src/schema/)" — with the trailing slash — and :131 strips exactly that character before :141's bare startsWith, so a fully compliant roster read still yields boundary-less hints. Behind it: the pathHints form cannot express a segment anchor at all. echosphere's pre-e92e9c4 REVIEWER_TRIGGERS matched a crate family with /^echosphere-[^/]+\//; moving to the runtime roster read is a regression for that shape, and the only stable form is enumerating each directory with its slash (and a PR that touches only .claude/ cannot exercise the hint, so the "first run doubles as the roster check" only ever proves the name parsed).

    4. Item 9: hallucinations is dereferenced twice — finish-ab.js:159 and :161. A ?? [] at :159 alone leaves the crash live.

    5. N2 and the kit's own .gitignore:3-9. The kit tells target projects to commit .claude/agent-memory/ ("DELETE this line") — which is exactly the exit under which the .claude-pr/ false positive fires, since the staging copy then carries the committed tree. grep -rn claude-pr over the kit returns nothing, so the anchored-ignore remedy has no portable home other than github-starter-templates.md; and the hook's prune list is hard-coded, so on the shipped hook a .gitignore entry is inert anyway. The other exit is not documented either: a project choosing the local scope per its Surface-inventory row gets no ignore guidance for .claude/agent-memory-local/, so its reviewer trees end up untracked and unignored. One caveat for the version range: the .claude-pr/ reproduction was on claude-code-action v1.0.206; echosphere pins v1.0.193 and whether that version stages .claude-pr/ is unverified.

    6. The block's bash -e + absent() contract makes "recorded exemptions" unreachable by construction. The comment's repair shape 3 already says this; the sharpening from a second target is that any spec written to the block's own vocabulary ("fix or record every red line") is unsatisfiable, because a recorded red aborts before both sentinels. Shape 3 (collect, then exit non-zero) is the only one under which "recorded" means anything.

    Provenance: the full record — per-item reproduction, the echosphere-specific dispositions, and the assessor-vs-verifier adjudications — is on echosphere ECH-39 (Linear). Nothing above is a request to split; filing it against the existing items is the intent.

  4. j4th commented on Sep 13, 2026

    @j4th
    OwnerAuthor

    Second application of e92e9c4 — echosphere, executed (j4th/echosphere#37, 2026-09-13)

    The research pass above became ECH-40 (a re-rough-in of ECH-39) and was executed: branch chore/ech-40-kit-sync-harvest-3, 24 commits, draft PR j4th/echosphere#37. Method: a three-way merge, not hand-reconciliation — git merge-file with base = the kit at the install commit (88b1edeb), ours = the project, theirs = e92e9c4. Ten conflict hunks across seven files; every other project fill auto-resolved. The verification block runs literal-green under bash -e from the repo root (three sentinels, exit 0); the review floor (/simplify, pr-review-toolkit:review-pr) plus the copied review-sweep.js ran on the result. What the execution adds, kit-side — items 1–7 are defects or gaps in the kit tree at e92e9c4, 8–10 calibration:

    1. adr-conformance-reviewer.md:55 assumes the adr-starters register's id form. "A citation … already recorded wrong there is cited by its C- number" is true of scaffold/references/adr-starters/corrections.md (### C-NNN —, :27) and false for a target whose register predates the starter — echosphere's docs/adr/corrections.md numbers entries in prose under per-ADR sections ("entry 10"). The reviewer then looks for ids the file does not have. Fix: "by the register's own id form", or a bootstrap-checklist disposition row for the reviewer beside the rule files.

    2. The path-scoping callouts describe the shipped state after stamping. logging.md:8 and testing.md:11 read "As shipped it carries a placeholder: replace <ext> …" / "Replace <ext> with the project's test-file extension(s) …"; the bootstrap checklist rows (bootstrap_checklist_template.md:90–91) say stamped and never say to rewrite or delete the callout, and scaffold/SKILL.md:315's "no <ext> left" clause is scoped to the globs — so a stamped target carries a paragraph telling the reader to stamp it. On echosphere four review dimensions flagged it as rot (8 of the sweep's 12 unverified findings were this one item). Related: testing.md's set is <ext>-shaped, and a language with inline unit tests (Rust's #[cfg(test)]) has no test-file extension — the honest stamp there is **/tests/** alone, which every reviewer then flags as missing the inline modules. One sentence in the callout naming that case would settle it.

    3. adr-new/SKILL.md:73 — "A new ADR connects to an existing one through one of two relationships" sits under ## Refines vs Supersedes vs Extends (:69), which names three, and the section goes on to describe four grains (Supersedes:, Refines:, clause-scoped Supersedes: ADR-NNNN Dn, Extends:). Count rot from the grain additions — let the list be the count. Separately: Promotes: is a slot the skill fills (:32, :48, :58) and never defines in that section.

    4. cbk-conventions.md contradicts itself on a frame's ## Pre-flight checks table. The Quick reference row (:234) says "Adding a ## Pre-flight checks row to a frame | Append-only edit to the frame's ## Pre-flight checks table"; the Mutation discipline row for docs/cbk/frame-NN.md (:192) names ## Rough-in events as the only append-only table, "otherwise immutable post-commit". Framing writes the pre-flight table at frame creation and rough-in reads it; no skill appends to it. Either the frame row gains the table or the Quick-reference row goes.

    5. settings.json _comment_hooks, STOP tier: "exits 2 while a .claude/agent-memory/ tree exists anywhere but the root" — the hook hunts agent-memory and agent-memory-local alike (detect-forked-agent-memory.sh:15, :61, :70). The comment undercounts by one scope.

    6. Item 2 above, sharpened by the gate. A [default] extend-ignore-re of [A-Za-z-]+\[[A-Za-z]\] does exempt onl[y] / chec[k], and also exempts any misspelling that abuts a one-letter bracket anywhere in the tree — recieve[d] went unflagged under typos-cli 1.49.0 while bare recieve was caught. The right shape is a [type.<name>] table with extend-glob = ["cbk-conventions-reference.md"] carrying the rule alone; typos honors it (the reference file goes clean, a scratch file outside the glob stays red). The kit skills' bracket matches are index notation (R[i]/R[j]/R[N] and M[n] in rough-in/references/hitl-question-bank.md, M[i] in framing/references/procedure.md, IC-[m]) plus one ordinary inflection bracket (allow[s], scaffold/references/github-starter-templates.md:207) — typos flags none of them, so none needs an exemption. Wherever the kit tells a target how to spellcheck .claude/, this is the form to name.

    7. Nothing wires the block into a target's gate. The kit's own CI runs ADR immutability only, and a target that adopts the block as written runs it by hand (awk extraction, bash -e) — so the registry↔tier↔table checks, the ## Writing memory parity, the executor-pair diff and the always-loaded budget fire only when someone remembers. The extraction is two documented awk lines; the blueprint tooling template could ship it as a task the gate calls. A suggestion, not a defect — echosphere has not wired it yet either.

    8. Item 7 (roster hints), calibration evidence. With the roster line written as enumerated trailing-slash prefixes (echosphere-core/, … , firmware/, proto/), the copied review-sweep.js roster read parsed all four reviewers, dispatched the three cross-cutting ones, and reported the domain reviewer as not path-matched on a rules-and-config diff (.claude/, .gitignore, .mcp.json.example, _typos.toml, docs/adr/) that touches none of its enumerated crate prefixes — the runtime read works as designed; the boundary loss stands as reported. Echosphere pins the enumeration to Cargo.toml's [workspace] members with a check in its project sub-block, so the list is derived rather than believed.

    9. Item 1 (ignore-driven prune), the recorded residual reproduced independently. A fresh reviewer built build/nested/.claude/agent-memory/<x>/ under an ignored build/ and got exit 0 — the "stray under an explicitly ignored build directory → 0 (pruned for speed …)" scenario in the POC's own table. The kit's own hook at e92e9c4 prunes a hand-kept -name list (node_modules, target, build, _build, dist, .venv — :70–72), so the same tree is exit 0 there too; reproduced with both hooks. require-repo-root-for-agents.sh refuses Task/Agent/Workflow dispatch whose working directory is not the git top-level (fails open with a warning outside any checkout), so the residual is a tree placed under an ignored directory by hand or by tooling. In the ported version (you-are-hear's, echosphere's) the header sentence "Never prune a path that could BE or CONTAIN the tree this hook hunts for" promises more than its basename-only skip delivers; whichever version the kit adopts, the residual belongs in the header, not only in the POC's scenario table.

    10. What went right. Recording the install commit makes the next sync a merge: of the 48 copy rows whose project copy carried lines the kit lacks, 47 were byte-identical to the kit at 88b1edeb — pure kit drift, derivable from the base, no reconciliation at all. Echosphere now records the synced sha in its conventions preamble; the kit could tell every target to record the install commit and sync by merge-file. And the block's bash -e + absent() contract did what the comment's shape 3 predicted: every red was fixed, none recorded — including two of the target's own pre-split checks (the path-drift and "Deferred meta-issues" greps), which match the kit's skill content by construction and had to be retired rather than re-homed.

    Project-side findings (a hybrid status-cell example, a sed range that read a manifest to EOF, live backticks in a ported heredoc) were echosphere's own and are fixed on #37; none is the kit's.

  5. j4th commented on Sep 13, 2026

    @j4th
    OwnerAuthor

    Addendum — the host-side auto-review on the same PR (j4th/echosphere#37, 2026-09-13)

    Two more items from claude-review.yml (echosphere's is the blueprint template as provisioned; line refs are echosphere's .github/workflows/claude-review.yml). The ready-flip run — actions/runs/34785442101 — ended error_max_turns ("Reached maximum number of turns (60)") after 7m16s, having posted nothing but the action's "Claude encountered an error" comment. The PR carries 137 changed files, 88 of them byte-identical kit copies.

    1. The default path's turn cap and the prompt's every-file contract cannot both hold on a sync-sized diff, and the failure mode posts nothing. The prompt says "Begin by enumerating every changed file … and walk each one — 'thorough' means every changed file is examined and accounted for"; the default path runs --max-turns 60 with --disallowedTools Agent (single thread, :322–:323). 137 files > 60 turns by construction: the run spent 66 Bash calls reading the tree (git show HEAD:<file> one by one) and produced zero review text before the cap. Under the deliverable-trap rule the artifact is the review, and here there is none — not a partial summary, not the files it did cover. Two fixes, either or both: make the cap scale with the diff (a --max-turns derived from the changed-file count, or the claude-deep-review path — which keeps Agent — auto-selected above a file-count threshold instead of by hand-applied label); and have the prompt degrade explicitly ("if the diff exceeds N files, review the authored set the PR body names and say what was not read") so a capped run still posts what it has. Echosphere skipped the re-run with skip-claude on the operator's call — the two /finish fan-outs had already covered the authored files — but a first real diff of this shape should not need that judgment.

    2. origin/main is absent under the checkout the template configures, and the reviewer's recovery is unallowlisted. actions/checkout runs with fetch-depth: 0 but ref: ${{ github.event.pull_request.head.sha }} (:156–:157), so the base branch is never fetched: the reviewer's first git diff origin/main...HEAD -- mise.toml returned fatal: origin/main...HEAD: no merge base. It then tried git fetch origin main --depth=50, --deepen=100, --unshallow (twice) and git fetch origin main — all five refused ("This command requires approval": the allowlist is gh pr view/diff/comment/checks and gh run view/list, :323) — six of the sixty turns gone before it fell back to gh pr diff, which the prompt already names. Fix: either fetch the base in the checkout step (a second fetch of ${{ github.event.pull_request.base.ref }}, or fetch-depth: 0 without ref: and a checkout of the head sha after), or state in the prompt that gh pr diff is the only diff source and origin/main does not exist in this checkout. Allowlisting git fetch is the weaker fix — it spends network on what the checkout should have provided.

    Not a kit defect, for completeness: the Stop hook did not fire in the container — every detect-forked-agent-memory string in the log is the reviewer reading the hook file — so the anchored /.claude-pr/ ignore from the 2026-09-07 comment holds on a real run.

  6. j4th commented on Sep 13, 2026

    @j4th
    OwnerAuthor

    The _example_PostToolUse_* stanza silently disables the entire hooks block

    Found while applying e92e9c4 to echosphere (Claude Code 2.1.270, 2026-09-13). This is a defect in the kit's § Hook authoring convention itself, not in a project's fill — and its blast radius is every guard the kit ships.

    Symptom

    With the kit's .claude/settings.json as shipped, no hook in the file runs. All four tiers are inert: hard-deny (immutable ADRs, lock files, frozen corpus, commits on main, agent launch root), ask-gate (gh pr state), advisory (format-on-edit), and stop (forked agent-memory). The file is valid JSON, the scripts are healthy, the trust dialog is accepted, and nothing warns.

    A settings checker reports it as:

    PreToolUse/PermissionRequest hooks are declared outside "hooks" (at the top level or under another key)

    That wording is misleading — there are no PreToolUse entries outside hooks. The trigger is _example_PostToolUse_analyzer, a top-level key whose value is a hook-matcher object (matcher + hooks). The loader appears to reject the file's hook configuration wholesale when it finds a hook-shaped object outside hooks.

    Evidence — single-variable test

    Probe: git commit --dry-run -m x on main, which protect-main-branch.sh must block (exit 2).

    _example_PostToolUse_analyzer value Probe result
    object (as the kit ships it) not blocked — hook never invoked, no output
    same key, value converted to a string blocked — PreToolUse:Bash hook error: protect-main-branch.sh: BLOCKED: git commit on 'main'

    Nothing else changed; the hooks block was byte-identical across both. The change hot-reloaded mid-session, no restart. Fed a crafted payload directly, the script returns exit 2 with the right message in both states — so the scripts were never the problem, only whether the harness invoked them.

    I did not isolate whether the loader keys on the object's shape or on the event name in the key. The checker's message names PreToolUse/PermissionRequest while the offending key says PostToolUse, which points at shape rather than name — but that's inference, not a test.

    Why this is worse than a lint nit

    The kit's own thesis is that "hooks enforce non-negotiables more reliably than instructions" (workflows.md § Anti-patterns, cbk-conventions-reference.md § Mechanize the gates). Every project that adopts the kit and follows § Hook authoring gets that guarantee inverted: the guards are present, documented, registered, verified by the gate — and never run. There is no failure signal.

    Note also that cbk-conventions-reference.md's verification block pins the broken shape:

    ... and has("_example_PostToolUse_analyzer") and (has("_example_PostToolUse_formatter") | not) ...
    

    so mise run check stays green while the hooks are dead, and deleting the key to fix the hooks turns the gate red.

    Suggested fix

    § Hook authoring currently says:

    Exemplar stanzas ship commented out. An advisory hook the project must wire to its stack (a formatter, an analyzer) ships unregistered with its registration as an _example_PostToolUse_* key beside hooks

    JSON has no comments, so "commented out" is being approximated by a live object — which is exactly what breaks. Options, cheapest first:

    1. Ship the exemplar's value as a string holding the snippet. One-line change to the kit's settings.json; the has() assertion in the verification block still passes unmodified, since it is type-agnostic. Least churn, keeps the exemplar where an operator will find it.
    2. Move the exemplar out of settings.json — into analyze-on-edit.sh's header (which already documents the stanza) or a settings.example.json. Cleaner separation; requires editing the verification block.

    Either way § Hook authoring's bullet needs rewording, and it's worth a verification-block check that no top-level key in settings.json holds a hook-shaped object — that's the invariant, and it's one jq line.

    Reproduction

    # with the kit's settings.json as shipped, from a repo on main:
    git commit --dry-run -m probe        # expect BLOCKED; observe it run
    jq '.["_example_PostToolUse_analyzer"] = "…snippet as a string…"' \
      .claude/settings.json > /tmp/s && cp /tmp/s .claude/settings.json
    git commit --dry-run -m probe        # now BLOCKED, as intended
  7. reopened this on Sep 30, 2026
  8. j4th commented on Sep 30, 2026

    @j4th
    OwnerAuthor

    Residue after harvest 4 — 2026-09-29 audit

    Reopening. #59 closed this issue COMPLETED, but its own definition of done — every item landed or explicitly a non-goal — was not met. This comment is the residue, item by item. Each residue is now applied in crease-data/crease#281 (open), so the next mint can take it mechanically.

    How the audit was made

    This is the method proposed as the gate for the next mint.

    The result:

    • 32 of 37 items landed;
    • 4 landed only in part;
    • 1 declined item's promised mitigation never landed;
    • several landed items left a smaller residue, mostly introduced by the fix itself.

    Every closure in the kit was made from inside the kit: #54–#57, #59 and 1e10363. None was manual or from outside. #33 was audited the same way and is clean except for one dropped note, carried here as B.

    The residue, and where each fix is

    The table gives the crease commit (on crease-data/crease#281) and the test that proves each.

    # #58 item Residue at 74edf84 Fix crease commit · proof
    R1 body item 4 No durable behavioural coverage of the hooks' branches. The block has structural checks over all nine hooks plus about ten probes. protect-immutable-adrs.sh and the knowledge-backend ask-gate have zero behavioural cases; protect-main-branch.sh has no allow or master case; guard-pr-state.sh covers only merge; protect-lock-files.sh covers 2 of 10 arms. § Hook authoring › Verify by payload says the branches are asserted in hook-contract-fixture.sh, which holds structure only. Behavioural fixtures, one per hook family (below). § Verify by payload names them. f7c5659 hook-guards-fixture.sh: 50 cases, then 57 after a verification pass found three branches no case reached (bare gh pr merge, git commit with nothing after it, the $PWD fallback; a12e547). hook-payloads-fixture.sh (29 cases, ported from you-are-hear); protected-paths-hook-fixture.sh (57). The main-branch guard's payload-cwd precedence is pinned both ways, with two throwaway repos that disagree and each winning when it should; neither the kit's over-buffer probe nor its older tests pinned it. Proven against mutants of each hook: twelve in all. The fixture names this project's require-notion-ok.sh, so the kit copy renames it require-knowledge-backend-ok.sh.
    R2 smaller item 1 A meta's children, titled [<slug>:<meta-tag>:R<#>], still fail /finish Step 1's four-form title check, which contradicts the sentence around it. The child form is the fifth admitted form in finish.md, finish-command.md and both finish-procedure.md copies. b7625cb · the block's byte-parallel diffs
    R3 2026-09-07 comment (separator) adr-new/SKILL.md:52 still pins Accepted · Dn superseded by ADR-${NNNN} for the clause-scoped annotation, against line 51's "the separator is the index's convention". § ADR relation grains still pins ·. The conventions name the starter's form as the starter's; a target's index sets its own cell and separator. For the kit's adr-new:52: don't pin a form. Tell the executor to read the index's rows, which can hold more than one form. crease's holds two: the Status cell for a parent with no relation parenthetical (ADR-0011's row), and inside the Title parenthetical for one that has it (ADR-0016's row). 0f3fb98, 5fbe124
    R4 2026-09-13 comment item 7 Nothing in the templates wires the block into check. A target learns it only from § Verification › Run it or the runner's header. blueprint/references/templates/tooling.md: the minimum task set gains a verification task check depends on, with a rule that it is a leg of check. scaffold/references/bootstrap_checklist_template.md: a verification-matrix row that runs it and expects both sentinels. f2323e1, 80f6e57
    R5 2026-09-13 comment items 2 and 6 (declined) The non-goal promised "a one-line note in § Verification's preamble names the idiom's cost". No such note exists anywhere in the kit (grep spellcheck/typos/extend-ignore: 0). The note: an absent check brackets one letter so the grep never matches its own line, and a spellchecking target sees a typo. Exempt the idiom in that one file, never repo-wide: in typos, a [type.<name>] table with extend-glob naming cbk-conventions-reference.md and an extend-ignore-re for the idiom. The shape was checked against typos' docs/reference.md. 23f9dc5
    R6 body item 2 Landed, but the partial-scan warning has no fixture; hook-contract-fixture.sh probes only the degrade path. hook-payloads-fixture.sh carries a partial-scan case: an unreadable directory makes the warning print and the hook exit 0. covered by R1's fixture
    R7 2026-09-13 comment item 5, H3 The spec says github-starter-templates.md's .gitignore starter gained anchored /.claude-pr/. That file has no .gitignore section at all. github-starter-templates.md § .gitignore — the harness block: /.claude/agent-memory-local/, /.claude/worktrees/, /.claude-pr/, and the kit's bytecode under .claude/workflows/, each pin-asserted. f81c487 · git check-ignore pins in the commit body
    R8 body item 3 New in the fix: the *.lock fallback's remediation says "add its name to the hook's exemption comment". No such comment exists, and nothing would read one. "If this file is NOT a lock file, give it a case arm of its own above the *.lock) arm in this hook (an arm that exits 0), and say so in the PR." abe5970 (red) → 073c0b1 · hook-guards-fixture.sh asserts the new route and the absence of "exemption comment"
    R9 body item 6 Implied and not landed: § Applied instances quotes "a top:workhorse cost ratio of 3.2× pooled" (and agent-cost.py "3.2x to 3.6x") with no note that it was priced at the old 0.1× top-tier cache rate. Stated conditionally, because the run's model is not known from here. Claude Code 2.1.257, which made Fable 5.1 the default Fable, was published 2026-09-01T17:15Z, the day of the run. If the drafters ran on Fable 5.1 (0.025×) the figures are upper bounds, and the transcripts' model ids settle it. e2b88be
    R10 2026-09-13 comment item 3 "Enumerate directories; a segment anchor cannot be a prefix" lives only in a review-sweep.js code comment. The rules a project reads when writing its roster line never say it. pr-review.md's craft rule and pr-review-reference.md § Authoring: a family of directories has no prefix form, so enumerate each and add one when the family grows. 49420c0
    R11 body item 5 (follow-on) Once a target registers an advisory hook, the two-views check also goes red until cbk-conventions.md § Mutation discipline names it. The Register: header and § Hook authoring don't say so. "Wiring one is three edits, not one: the stanza into hooks.PostToolUse, the hook's name into the project sub-block's ADVISORY_WIRED, and the hook's name into cbk-conventions.md § Mutation discipline's two-views paragraph." In § Hook authoring and in the Register: header of both advisory exemplars. 546fc71
    R12 addendum item 11 Ask (b) landed, but the degrade clause's [N] comes with no way to size it against --max-turns 60. A comment above claude_args: --max-turns is spent per tool call, not per round trip (echosphere measured it: five Reads in one response still cost five turns). At one Read a file, keep about a third of the cap for gh, inline comments and the summary: N = 40 against 60, and the two move together. 713cc5c (crease's filled workflow also gains the clause it never had)
    R13 2026-09-13 comment item 9 The detector's rule 1 still says "Never prune a path that could BE or CONTAIN the tree", which the header's Residual line contradicts. "Never prune a directory whose own name is one of the three this rule protects — the two the hook hunts, agent-memory and agent-memory-local, and .claude, which holds them … A tree nested inside some OTHER ignored directory is pruned with it — the Residual above, a limit of this rule, not a promise it keeps." 48beaea, 0e536a5
    R14 several Kit-internal traceability slips, nothing to apply in a target: H3 claims a .gitignore starter that does not exist (R7); H4 says it "closes … item 6" against D31 and § Non-goals; the spec's cluster headings cite "the 2026-09-13 comment", and three comments carry that date; several items are covered only by a non-goal's wording, never by number. For the kit's next spec: cite comments by id, and trace every item by number. —
    B from #33 #33's review-layer calibration note (j4th/you-are-hear#43: the floor, the round-2 sweep and the flip's auto-review each caught what the previous layer missed, and "none of them was the layer that found the previous layer's bug") landed nowhere. pr-review-reference.md § Anti-patterns › "Folding one review layer into another". d06330c

    Two more residues are recorded, not applied:

    The hook fixtures (R1, R6, R8) are bash, need only git and jq, and run against throwaway trees, so they port as they are. The verification block in crease runs all three.

    The same final check swept all three repos for learnings with no kit home after the harvest-4 cutoff. What it found beyond #58 is filed as #70 (supply chain and shell safety in the templates), #71 (kit-owned code and a target's formatter scope), #72 (review-sweep), #73 (the verification block as a gate) and #74 (review and verification conventions). The rest go as comments on #60, #63, #66, #67, #68 and #69. Each post carries the crease commit and the test that proves it. #33 stays closed: its one dropped note is B above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    harvestHarvested from a real cascade runsource:you-are-hearEvidence from the you-are-hear run (GitHub axis)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions