Skip to content

feat: code-owned run engine, enforced verification, real topology registry (v0.3.0) - #1

Merged
aditya0si merged 1 commit into
mainfrom
feat/code-owned-run-engine
Sep 18, 2026
Merged

aditya0si merged 1 commit into
mainfrom
feat/code-owned-run-engine

Conversation

@aditya0si

Copy link
Copy Markdown
Owner

What this changes

The run loop moves out of the prompt and into code.

The plugin entry point was 56 lines of config injection, and src/worktree.ts, src/cost.ts, src/state.ts and src/artifacts.ts were imported by nothing (repo-wide grep before this change). So the README's "DAG engine", "git worktree isolation", "cost tracking", "checkpoint/resume" and "Zod-validated artifact bus" were instructions to an LLM, not properties of the runtime. The six pattern files were never put in front of any model.

The rule now: the LLM proposes, the runtime disposes.

Engine (new)

File What it does
src/engine.ts plan validation (cycles, dangling deps, dup ids, unknown topologies), topological dispatch waves, concurrency cap, attempt accounting, per-round model-ladder escalation, dead-letter after maxRounds, budget enforcement that refuses dispatch at the cap
src/events.ts append-only events.jsonl with a sha256 hash chain; state.json is a derived snapshot; deriveSession() replays, verifyChain() refuses tampering
src/policy.ts single source of truth for topology names, model ladders, required checks, budget
src/guard.ts runtime role enforcement — read-only roles cannot call write tools
src/tools.ts teamwork_plan / _dispatch / _verify / _status / _resume, gated to orchestrating agents

Verification carries evidence

Checks now record cmd, exitCode, stdoutSha256, durationMs. The engine rejects a report when:

  • a programmatic/adversarial check has no cmd or no exitCode;
  • a check claims passed: false while its command exited 0 (or the reverse);
  • a PASS contains no executed check — "I read the diff and it looks right" is not a verification.

A rejected report does not count as a round.

Fixed

  • Frontmatter was inert. The whole markdown file (frontmatter included) was passed as the prompt string, so mode defaulted to all, permission: edit: deny on the verifier did nothing, and temperature was ignored. Now parsed into real AgentConfig keys. (The claim a reviewer tests first: "the verifier cannot write" was a paragraph, not a boundary.)
  • The plugin clobbered per-role models. config.agent[name] = { prompt } replaced what the installer wrote. Spread-merged now; a test asserts the model survives.
  • Three incompatible topology vocabularies. Files said long-proof/distributed-coding/document-review; the sentinel prompt and state.ts said proof/large-swarm/doc-review. Four of five routing names matched no file, and the pattern files never reached a model.
  • {{if eq .sessionId ""}} in the /teamwork command was literal text to the model (OpenCode expands only $ARGUMENTS, $1..$N, !cmd, @file), and $ARGUMENTS appeared twice. Flags are parsed in code now.
  • Shell-injection surface in the worktree manager: git was invoked via interpolated strings built from model-produced names. All calls use argv + shell: false, with name/session validation.
  • cleanupSession never cleaned up: it ran git worktree remove against the worktrees' parent directory, swallowed the failure, then deleted directories git still had registered.
  • dist/ didn't match exports (build:support emitted dist/src/*.js while exports pointed at dist/*.js), and SKILL.md wasn't in files.
  • Leaf agents get task: deny — a worker can no longer fan out its own swarm.

Tests

26 pass, 0 fail              # bun test
e2e passed                   # real git repo: worktrees, evidence, budget, dead-letter, resume, cleanup
plugin integration passed    # built bundle, real hook shapes: config, commands, tools, guard, compaction
smoke tests passed           # installer

CI runs all of them on linux/macos/windows plus a npm pack -> clean-install check. bun run verify is the single gate and prepublishOnly runs it.

Two bugs this branch's own tests caught while writing it: round double-counting in deriveSession (a FAIL emitted both verification.report and task.round), and --concurrency=3 leaking into the request text.

Docs

docs/self-improvement.md — how to improve the policy (topologies, model ladders, required checks) from the event log: attribution, typed JSON-Patch deltas, an adversarial falsifier on the hypothesis, shadow replay on a frozen benchmark, canary + rollback, and the escape-rate metric (verifier PASSes a human later rejects) that keeps the loop from just teaching the team to satisfy its own checker. Designed, not built — the log it needs is what this PR adds.

…istry (v0.3.0)

The run loop moves out of the prompt and into code. Previously worktree
isolation, the DAG, cost tracking, checkpoint/resume and schema validation
were instructions to an LLM: src/worktree.ts, src/cost.ts, src/state.ts and
src/artifacts.ts were imported by nothing, so none of those guarantees were
enforced by the runtime. The plugin entry point was 56 lines of config
injection.

Engine (new)
- src/engine.ts: plan validation (cycles, dangling deps, dup ids, unknown
  topologies), topological dispatch waves, concurrency cap, attempt
  accounting, per-round model ladder escalation, dead-letter after maxRounds,
  budget enforcement that refuses dispatch at the cap.
- src/events.ts: append-only events.jsonl with a sha256 hash chain.
  state.json is now a derived snapshot; deriveSession() replays a run,
  verifyChain() refuses a tampered log.
- Tools: teamwork_plan / _dispatch / _verify / _status / _resume, gated to
  orchestrating agents only.
- src/guard.ts: runtime role enforcement — read-only roles cannot call write
  tools even if the host ignores part of the permission block.
- src/policy.ts: single source of truth for topology names, model ladders,
  required checks and budget.

Verification
- Checks carry cmd + exitCode + stdoutSha256. A PASS with no executed check,
  or a check that contradicts its exit code, is rejected and does not count
  as a round. Fabricated PASSes were the weakest point of the design.

Fixed
- Agent frontmatter was inert: the whole markdown file (frontmatter included)
  was passed as `prompt`, so mode defaulted to "all", permission.edit:deny on
  the verifier did nothing, temperature was ignored. Now parsed and injected
  as real AgentConfig keys.
- The plugin clobbered the per-role models the installer writes. Config is
  spread-merged; a test asserts the model survives.
- Three incompatible topology vocabularies (files vs sentinel prompt vs
  state.ts); four of five routing names matched no file. Pattern files also
  never reached any model — orchestrating agents now get a generated index
  with absolute paths.
- {{if eq .sessionId ""}} in the /teamwork command was literal text (OpenCode
  expands only $ARGUMENTS, $1..$N, !`cmd`, @file) and $ARGUMENTS appeared
  twice. Flags are now parsed in code.
- Shell-injection surface: git was invoked via interpolated command strings
  built from model-produced names. All git calls use argv + shell:false, with
  name validation.
- cleanupSession ran `git worktree remove` against the worktrees' parent
  directory and then deleted directories git still had registered.
- dist/ layout did not match package.json exports (dist/src/*.js).
- SKILL.md was not in files.

Tests
- 26 unit tests (bun test)
- bun run e2e: the whole loop against a real git repo — worktrees, evidence
  capture, budget, dead-letter, resume, cleanup
- bun run integration: the built bundle driven with the hook shapes OpenCode
  sends — config, commands, tools, guard, compaction
- CI on linux/macos/windows runs all three + npm pack -> clean install

Also adds docs/self-improvement.md: the design for improving the policy
(topologies, model ladders, required checks) from the event log, with the
escape-rate guard metric and the failure modes to design against.
@aditya0si
aditya0si merged commit 281821a into main Sep 18, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant