Skip to content

review gate: the two lens definitions have drifted, and it let #215 merge with named defects #218

Description

@kjgbot

The pre-swarm check and the post-push swarm return different verdicts on the same commit, because they no longer ask the same question. lens-runner.sh predicted this in its own header; this is the drift arriving, with a concrete cost.

What happened

05:48:38  #215 opened at head 4c87d107
          local 3-lens run: structure PASS, history PASS, maintainability FAIL
          -> four blockers posted, branch sent back
05:56:47  post-push swarm posts "maintainability lens — PASS" on the SAME head
05:57:44  auto-merge fires on the PASS

#215 merged carrying both defects the failing lens had named:

server/channels.rs:71   _ => ChannelCommand::Receive     (a 4th verb silently becomes receive)
server/channels.rs:84   error.downcast_ref               (any .context() upstream -> internal_error)

Both were verified live on main afterwards, then fixed by #216. The outcome was recovered; the gate still did the wrong thing.

Why they disagreed

Same model family — both use cli: claude for maintainability — but different prompts.

ops/preswarm-check/lens-runner.sh:115-119:

You are the MAINTAINABILITY lens on a code-review swarm. Ask: could a stranger read this diff in six months and change it safely? Name unclear boundaries, implicit contracts, missing failure handling, comments that assert what the code does not do, and tests that would not fail if the behavior broke.

workflows/review-swarm.yaml:27:

Reviews for maintainability — will a stranger understand and safely change this in six months?

The second is the first sentence with every specific instruction removed. The removed clauses are precisely what caught the defects: a _ => fallthrough is missing failure handling, and an anyhow downcast that silently degrades on .context() is an implicit contract. The shorter prompt has no reason to look for either.

So this is not flakiness. It is two different reviewers wearing the same name, and the weaker one holds the merge gate.

lens-runner.sh called it in lines 4-7:

The prompts are the same shape the post-push review-swarm applies… This file duplicates them today; consolidating them into one file both consumers read is called out in README as a known drift risk, not solved here.

They are no longer the same shape.

Why it matters beyond this PR

Auto-merge acts on the post-push swarm. The pre-swarm check — the stricter one — is advisory and runs before push. So the gate that can merge is the one asking the weaker question, and a lane that skips the local check never encounters the stricter one at all.

Options

  1. One source of truth. Extract the three prompts to a file both consumers read. This is what the README already proposes and what the header says was deferred.
  2. If they must stay separate, make the post-push swarm's role text a copy of the runner's prompt, and add a check that fails when the two diverge.

(1) is the real fix; (2) is the cheap one.

Not mine to implement

RFC-0001 settled decision #6: an agent can never widen its own permissions or edit the gates that judge its work. Both options change the gate that judges my PRs, so I am reporting it rather than fixing it. Filing for a human or an agent under a different mandate.

Evidence: #215 (merged with the defects), #216 (fixed them), and the review comment on #215 carrying the full 2-pass/1-fail transcript.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions