Repository navigation
feat: eight lifecycle methods from two MIT skill collections; release 0.15.0 - #42
Merged
Merged
Conversation
- New plugins/ess/skills/specifying/references/interview.md: ask the open decisions whose prerequisites are settled as one numbered round, each with the answer the agent would take; look facts up instead of asking for them; challenge a term against the entity the specification already declares; validate after each round; finish when no decision is open. - Headless runs do not hold the interview: each question and the answer taken go into an approval-record, per aep:planning section 4, and a decision with no evidence either way stays an UNMAPPED: marker. - specifying/SKILL.md points at it from "Starting a domain from nothing" and from the paragraph that turns open markers into questions; the planning skill's guardrail 7 points at it where a noun is undescribed. - New eval case specifying-interview-headless holds the headless half. Adapted, not copied, from skills/productivity/grilling and skills/engineering/domain-modeling in github.com/mattpocock/skills (MIT); no verbatim text.
- New activity skill plugins/aep/skills/diagnosing: build one command that drives the real code path, asserts the reported symptom, is deterministic and fast, and has already been run, before any hypothesis. Then minimise, rank 3-5 falsifiable hypotheses, probe one variable at a time with tagged debug output, and write the regression test at a seam that reproduces the real call pattern. - The red and the green run are recorded as test_result evidence against the owning story; a missing seam is filed as a draft story with an informed_by edge. - Registered in b10x:routing, the README plugin tree, website intro, structure and aep pages, and the checker's required-file list. - New eval case diagnosing-red-loop-first: a test command runs before the first write under src/, and both runs reach the store. Adapted, not copied, from skills/engineering/diagnosing-bugs in github.com/mattpocock/skills (MIT); its HITL shell template is not ported. No verbatim text.
- The adversary's attack table gains a row for a test that would still pass if every function it calls returned a default value, and a subsection naming the five shapes that do (no real assertion, absence only, self-referential expected value, constant pin, fixture asserting fixture) with a Rust example and a rewrite for each. The check runs before mutation because it is cheaper. - ess:testing-conformance applies the same check to a scenario: point the suite at a target that accepts every command and returns empty views; a scenario that still passes asserts no observed state, and gets a view assertion before any mutation is run. - New eval case adversary-tautological-test: the seeded tests/total.rs is rewritten, src/ is left alone, the suite runs after the rewrite. Adapted, not copied, from pstack/skills/principle-test-behavior-not- implementation in github.com/cursor/plugins (MIT). No verbatim text.
- wave.md: before a green unit merges, restate its story's claim as a condition, a measurement and a threshold, run the same command on the integration branch and on the unit's branch, and give exactly one verdict: VERIFIED, NOT VERIFIED or INCONCLUSIVE. - The routing table splits the green row on that verdict: VERIFIED merges, NOT VERIFIED goes back to the implementor with both outputs, INCONCLUSIVE does not merge and goes to a person after one tighter measurement. - The verdict is recorded as verification evidence naming both runs; the outputs live in the scratch directory the unit brief assigns, never /tmp. - New eval case wave-claim-verdict: a green unit whose claim does not hold gets a verification record and its branch is not merged. Adapted, not copied, from cursor-team-kit/skills/verify-this in github.com/cursor/plugins (MIT). No verbatim text.
- story-scoper.md: the returned Scope section gains a Safety fact line, the one fact the change is safe because of, with how far it was proved on a five-step ladder (stated, file:line, walked, ran, reproduced). A read-only scoper stops at 2 or 3 and marks the line unproven. Rule 6 also lists where a symbol search stops: API fields, columns, wire formats, flags, callers three hops away. - security-reviewer.md: before reporting a change as safe, prove that fact to step 4 (a test that calls the real code and fails if the fact is false) or report it unproven with the fact and the step. - New eval cases story-scoper-safety-fact and security-reviewer-safety-fact. Adapted, not copied, from pstack/skills/blast-radius in github.com/cursor/plugins (MIT). No verbatim text.
…ontal ones - decomposer.md: each story is a vertical slice through every layer the outcome needs, demonstrable alone and sized for one fresh session, with prefactoring drafted first. A wide mechanical change is drafted as expand, one migrate story per batch of callers, and contract, linked by depends_on; a final integrate-and-verify story where a batch cannot stay green alone. - plan-critic-design.md: a fifth defect, the horizontal slice (a set cut by layer so no item can be shown working end to end), with expand/migrate/ contract stated as not this defect. - New eval case decomposer-expand-migrate-contract. Adapted, not copied, from skills/engineering/to-tickets in github.com/mattpocock/skills (MIT). No verbatim text.
- New plugins/aep/skills/implementing/references/adversary-panel.md: on the operator's request, one adversary per available model family, each on the same worktree with the same brief and no sight of the others, each told a distinct test-file name. Each seat is its own review-result. Findings are sorted into act on, consider, noted and dismissed, marked with the families that raised them, and the unit is routed on act on. - With one family available the panel is one reviewer, and the report says so rather than presenting it as a cross-family review. - wave.md points at it from the adversary dispatch; the implementing skill's role list links it. - New eval case adversary-panel-one-family. Adapted, not copied, from pstack/skills/interrogate in github.com/cursor/plugins (MIT). No verbatim text.
- New plugins/b10x/skills/authoring-plugins/references/writing.md with five rules, each with a before/after taken from a skill in this repository: a description names each distinct trigger once; every step ends on a checkable completion criterion; reference only some branches need goes behind a pointer that says when to read it; state the target behaviour instead of only a prohibition; cut sentences and weak words the model already obeys. - authoring-plugins/SKILL.md section 2 points at it before any SKILL.md, procedure or AGENTS.md text is written or reviewed. - New eval case authoring-prohibition-only-rule. - evals/README.md: the case table lists all seventeen cases, and the coverage counts are recounted from the subject: fields (8 of 14 agents, 6 of 21 non-command skills), with the uncovered agents and skills named. Adapted, not copied, from skills/productivity/writing-for-agents in github.com/mattpocock/skills (MIT). The before/after quotations are from this repository's own skills. No verbatim text from the source.
- Workspace, all ten plugin manifests (Claude Code and Codex) and every Skill version line move from 0.14.17 to 0.15.0. - CHANGELOG.md gains the 0.15.0 section: the new aep:diagnosing skill, the specification interview, the tautology check, the claim verdict before merge, the safety fact, tracer slices with expand/migrate/contract, the cross-family adversary panel, and the agent-writing rules. - The checker test a_reference_beside_a_skill_is_in_that_skills_scope asserts against the live corpus; it now expects the seven cases that name aep:implementing instead of the original two. task check: exit 0 (fmt, clippy -D warnings, 139 tests, 17 eval cases, marketplace b10x with 5 plugins valid). agentplugins-check tools: every spelled command exists in aep 0.61.1, ess 0.37.0 and worktree 0.8.2; it still reports the verified.json pins for aep (0.60.0) and ess (0.35.0) as behind, which was already failing the scheduled Tools run on main before this change.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Eight methods from
github.com/mattpocock/skills(c55ee46) andgithub.com/cursor/plugins(ecc249f), both MIT, rewritten into theaep,essandb10xplugins. Each has an eval case beside it.approval-recordsplugins/ess/skills/specifying/references/interview.md, pointer fromaep:planningguardrail 7aep:diagnosingskill: a red-capable command before any hypothesis, runs recorded astest_resultplugins/aep/skills/diagnosing/, routing, README, websiteimplementing/references/adversary.md,ess:testing-conformanceimplementing/references/wave.mdSafety factline with a five-step proof ladder; security reviewer proves it to step 4story-scoper.md,security-reviewer.mdplanning/references/decomposer.md,plan-critic-design.mdreview-resultper seatimplementing/references/adversary-panel.mdauthoring-plugins/references/writing.mdVerbatim text: none. Each method is adapted in this repository's words, and each file names its source.
Eval corpus: 17 cases, up from 8. The 9 new
recorded/directories hold only a README with the live command, like the existing cases; no live run was made.Gates, run locally:
task check: exit 0 (fmt, clippy-D warnings, 139 tests, 17 eval cases, marketplaceb10xwith 5 plugins).agentplugins-check tools: every spelled command exists in aep 0.61.1, ess 0.37.0 and worktree 0.8.2. It still fails on theverified.jsonpins (aep 0.60.0, ess 0.35.0), as the scheduled Tools run onmainalready did on 2026-09-27; this PR does not move them.One checker test changed:
a_reference_beside_a_skill_is_in_that_skills_scopeasserts against the live corpus, and now expects the seven cases that nameaep:implementing.