The pre-swarm check and the post-push swarm return different verdicts on the same commit, because they no longer ask the same question. lens-runner.sh predicted this in its own header; this is the drift arriving, with a concrete cost.
What happened
05:48:38 #215 opened at head 4c87d107
local 3-lens run: structure PASS, history PASS, maintainability FAIL
-> four blockers posted, branch sent back
05:56:47 post-push swarm posts "maintainability lens — PASS" on the SAME head
05:57:44 auto-merge fires on the PASS
#215 merged carrying both defects the failing lens had named:
server/channels.rs:71 _ => ChannelCommand::Receive (a 4th verb silently becomes receive)
server/channels.rs:84 error.downcast_ref (any .context() upstream -> internal_error)
Both were verified live on main afterwards, then fixed by #216. The outcome was recovered; the gate still did the wrong thing.
Why they disagreed
Same model family — both use cli: claude for maintainability — but different prompts.
ops/preswarm-check/lens-runner.sh:115-119:
You are the MAINTAINABILITY lens on a code-review swarm. Ask: could a stranger read this diff in six months and change it safely? Name unclear boundaries, implicit contracts, missing failure handling, comments that assert what the code does not do, and tests that would not fail if the behavior broke.
workflows/review-swarm.yaml:27:
Reviews for maintainability — will a stranger understand and safely change this in six months?
The second is the first sentence with every specific instruction removed. The removed clauses are precisely what caught the defects: a _ => fallthrough is missing failure handling, and an anyhow downcast that silently degrades on .context() is an implicit contract. The shorter prompt has no reason to look for either.
So this is not flakiness. It is two different reviewers wearing the same name, and the weaker one holds the merge gate.
lens-runner.sh called it in lines 4-7:
The prompts are the same shape the post-push review-swarm applies… This file duplicates them today; consolidating them into one file both consumers read is called out in README as a known drift risk, not solved here.
They are no longer the same shape.
Why it matters beyond this PR
Auto-merge acts on the post-push swarm. The pre-swarm check — the stricter one — is advisory and runs before push. So the gate that can merge is the one asking the weaker question, and a lane that skips the local check never encounters the stricter one at all.
Options
- One source of truth. Extract the three prompts to a file both consumers read. This is what the README already proposes and what the header says was deferred.
- If they must stay separate, make the post-push swarm's role text a copy of the runner's prompt, and add a check that fails when the two diverge.
(1) is the real fix; (2) is the cheap one.
Not mine to implement
RFC-0001 settled decision #6: an agent can never widen its own permissions or edit the gates that judge its work. Both options change the gate that judges my PRs, so I am reporting it rather than fixing it. Filing for a human or an agent under a different mandate.
Evidence: #215 (merged with the defects), #216 (fixed them), and the review comment on #215 carrying the full 2-pass/1-fail transcript.
The pre-swarm check and the post-push swarm return different verdicts on the same commit, because they no longer ask the same question.
lens-runner.shpredicted this in its own header; this is the drift arriving, with a concrete cost.What happened
#215 merged carrying both defects the failing lens had named:
Both were verified live on
mainafterwards, then fixed by #216. The outcome was recovered; the gate still did the wrong thing.Why they disagreed
Same model family — both use
cli: claudefor maintainability — but different prompts.ops/preswarm-check/lens-runner.sh:115-119:workflows/review-swarm.yaml:27:The second is the first sentence with every specific instruction removed. The removed clauses are precisely what caught the defects: a
_ =>fallthrough is missing failure handling, and ananyhowdowncast that silently degrades on.context()is an implicit contract. The shorter prompt has no reason to look for either.So this is not flakiness. It is two different reviewers wearing the same name, and the weaker one holds the merge gate.
lens-runner.shcalled it in lines 4-7:They are no longer the same shape.
Why it matters beyond this PR
Auto-merge acts on the post-push swarm. The pre-swarm check — the stricter one — is advisory and runs before push. So the gate that can merge is the one asking the weaker question, and a lane that skips the local check never encounters the stricter one at all.
Options
(1) is the real fix; (2) is the cheap one.
Not mine to implement
RFC-0001 settled decision #6: an agent can never widen its own permissions or edit the gates that judge its work. Both options change the gate that judges my PRs, so I am reporting it rather than fixing it. Filing for a human or an agent under a different mandate.
Evidence: #215 (merged with the defects), #216 (fixed them), and the review comment on #215 carrying the full 2-pass/1-fail transcript.