Implemented
PR #14 added a supervisor-side adversarial role using a configured Ollama or OpenRouter model, bounded generated checks in separate OpenShell evaluation guests, provenance fields, a hard failure gate, and a capped findings penalty. Controller tests use fake provider responses and a fake fleet queue.
Remaining work
Run one bounded review with a real configured reviewer and OpenShell evaluation guest. Capture provider and model identity, prompt and response artifacts, generated test, test result, gate or penalty, and selection outcome. Verify that candidate-controlled state cannot alter the review record or score. A coding-agent review of proof PRs is separate from this runtime role.
Acceptance criteria
- Adversarial findings are isolated from candidate-controlled scoring state.
- A failing adversarial check blocks selection or produces an explicit penalty.
- Role, provider, model, prompt, and result provenance are recorded in the report.
Implemented
PR #14 added a supervisor-side adversarial role using a configured Ollama or OpenRouter model, bounded generated checks in separate OpenShell evaluation guests, provenance fields, a hard failure gate, and a capped findings penalty. Controller tests use fake provider responses and a fake fleet queue.
Remaining work
Run one bounded review with a real configured reviewer and OpenShell evaluation guest. Capture provider and model identity, prompt and response artifacts, generated test, test result, gate or penalty, and selection outcome. Verify that candidate-controlled state cannot alter the review record or score. A coding-agent review of proof PRs is separate from this runtime role.
Acceptance criteria