feat: code-owned run engine, enforced verification, real topology registry (v0.3.0) - #1
Merged
Merged
Conversation
…istry (v0.3.0)
The run loop moves out of the prompt and into code. Previously worktree
isolation, the DAG, cost tracking, checkpoint/resume and schema validation
were instructions to an LLM: src/worktree.ts, src/cost.ts, src/state.ts and
src/artifacts.ts were imported by nothing, so none of those guarantees were
enforced by the runtime. The plugin entry point was 56 lines of config
injection.
Engine (new)
- src/engine.ts: plan validation (cycles, dangling deps, dup ids, unknown
topologies), topological dispatch waves, concurrency cap, attempt
accounting, per-round model ladder escalation, dead-letter after maxRounds,
budget enforcement that refuses dispatch at the cap.
- src/events.ts: append-only events.jsonl with a sha256 hash chain.
state.json is now a derived snapshot; deriveSession() replays a run,
verifyChain() refuses a tampered log.
- Tools: teamwork_plan / _dispatch / _verify / _status / _resume, gated to
orchestrating agents only.
- src/guard.ts: runtime role enforcement — read-only roles cannot call write
tools even if the host ignores part of the permission block.
- src/policy.ts: single source of truth for topology names, model ladders,
required checks and budget.
Verification
- Checks carry cmd + exitCode + stdoutSha256. A PASS with no executed check,
or a check that contradicts its exit code, is rejected and does not count
as a round. Fabricated PASSes were the weakest point of the design.
Fixed
- Agent frontmatter was inert: the whole markdown file (frontmatter included)
was passed as `prompt`, so mode defaulted to "all", permission.edit:deny on
the verifier did nothing, temperature was ignored. Now parsed and injected
as real AgentConfig keys.
- The plugin clobbered the per-role models the installer writes. Config is
spread-merged; a test asserts the model survives.
- Three incompatible topology vocabularies (files vs sentinel prompt vs
state.ts); four of five routing names matched no file. Pattern files also
never reached any model — orchestrating agents now get a generated index
with absolute paths.
- {{if eq .sessionId ""}} in the /teamwork command was literal text (OpenCode
expands only $ARGUMENTS, $1..$N, !`cmd`, @file) and $ARGUMENTS appeared
twice. Flags are now parsed in code.
- Shell-injection surface: git was invoked via interpolated command strings
built from model-produced names. All git calls use argv + shell:false, with
name validation.
- cleanupSession ran `git worktree remove` against the worktrees' parent
directory and then deleted directories git still had registered.
- dist/ layout did not match package.json exports (dist/src/*.js).
- SKILL.md was not in files.
Tests
- 26 unit tests (bun test)
- bun run e2e: the whole loop against a real git repo — worktrees, evidence
capture, budget, dead-letter, resume, cleanup
- bun run integration: the built bundle driven with the hook shapes OpenCode
sends — config, commands, tools, guard, compaction
- CI on linux/macos/windows runs all three + npm pack -> clean install
Also adds docs/self-improvement.md: the design for improving the policy
(topologies, model ladders, required checks) from the event log, with the
escape-rate guard metric and the failure modes to design against.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
The run loop moves out of the prompt and into code.
The plugin entry point was 56 lines of config injection, and
src/worktree.ts,src/cost.ts,src/state.tsandsrc/artifacts.tswere imported by nothing (repo-wide grep before this change). So the README's "DAG engine", "git worktree isolation", "cost tracking", "checkpoint/resume" and "Zod-validated artifact bus" were instructions to an LLM, not properties of the runtime. The six pattern files were never put in front of any model.The rule now: the LLM proposes, the runtime disposes.
Engine (new)
src/engine.tsmaxRounds, budget enforcement that refuses dispatch at the capsrc/events.tsevents.jsonlwith a sha256 hash chain;state.jsonis a derived snapshot;deriveSession()replays,verifyChain()refuses tamperingsrc/policy.tssrc/guard.tssrc/tools.tsteamwork_plan/_dispatch/_verify/_status/_resume, gated to orchestrating agentsVerification carries evidence
Checks now record
cmd,exitCode,stdoutSha256,durationMs. The engine rejects a report when:programmatic/adversarialcheck has nocmdor noexitCode;passed: falsewhile its command exited 0 (or the reverse);A rejected report does not count as a round.
Fixed
promptstring, somodedefaulted toall,permission: edit: denyon the verifier did nothing, andtemperaturewas ignored. Now parsed into realAgentConfigkeys. (The claim a reviewer tests first: "the verifier cannot write" was a paragraph, not a boundary.)config.agent[name] = { prompt }replaced what the installer wrote. Spread-merged now; a test asserts the model survives.long-proof/distributed-coding/document-review; the sentinel prompt andstate.tssaidproof/large-swarm/doc-review. Four of five routing names matched no file, and the pattern files never reached a model.{{if eq .sessionId ""}}in the/teamworkcommand was literal text to the model (OpenCode expands only$ARGUMENTS,$1..$N,!cmd,@file), and$ARGUMENTSappeared twice. Flags are parsed in code now.shell: false, with name/session validation.cleanupSessionnever cleaned up: it rangit worktree removeagainst the worktrees' parent directory, swallowed the failure, then deleted directories git still had registered.dist/didn't matchexports(build:supportemitteddist/src/*.jswhileexportspointed atdist/*.js), andSKILL.mdwasn't infiles.task: deny— a worker can no longer fan out its own swarm.Tests
CI runs all of them on linux/macos/windows plus a
npm pack-> clean-install check.bun run verifyis the single gate andprepublishOnlyruns it.Two bugs this branch's own tests caught while writing it: round double-counting in
deriveSession(a FAIL emitted bothverification.reportandtask.round), and--concurrency=3leaking into the request text.Docs
docs/self-improvement.md— how to improve the policy (topologies, model ladders, required checks) from the event log: attribution, typed JSON-Patch deltas, an adversarial falsifier on the hypothesis, shadow replay on a frozen benchmark, canary + rollback, and the escape-rate metric (verifier PASSes a human later rejects) that keeps the loop from just teaching the team to satisfy its own checker. Designed, not built — the log it needs is what this PR adds.