Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
60 commits
Select commit Hold shift + click to select a range
683b26a
release: 0.10.0 — the coverage skill, and specify corrected by a real…
b10x-bot[bot] Sep 18, 2026
7e5b45a
feat(aep-drive): add a defensively-framed security-reviewer agent
b10x-bot[bot] Sep 18, 2026
9874295
chore: release 0.9.3
b10x-bot[bot] Sep 18, 2026
4ff6ede
chore: advance the shared documentation runtime and add the source check
b10x-bot[bot] Sep 23, 2026
bb2e053
Merge pull request #12 from beyond10x/docs/source-check-339b4b8
b10x-bot[bot] Sep 23, 2026
514ef9e
refactor!: retire ess-specify; route spec-driven work to the ESS plug…
b10x-bot[bot] Sep 23, 2026
cefcd53
refactor!: rename the marketplace b10x; list worktree by pin, retire …
b10x-bot[bot] Sep 24, 2026
e7a1cbe
chore: release 0.11.0
b10x-bot[bot] Sep 24, 2026
0f58d6f
Merge pull request #13 from beyond10x/refactor/b10x-catalog-worktree-pin
b10x-bot[bot] Sep 24, 2026
52673df
feat!: one-sentence onboarding through b10x setup
b10x-bot[bot] Sep 24, 2026
a6da593
chore: release 0.12.0
b10x-bot[bot] Sep 24, 2026
95bd079
Merge pull request #14 from beyond10x/feat/b10x-setup-onboarding
b10x-bot[bot] Sep 24, 2026
55983dc
feat!: one aep plugin, activity-named skills, and the concept enforced
b10x-bot[bot] Sep 24, 2026
20294ad
chore: release 0.13.0
b10x-bot[bot] Sep 24, 2026
8ac0ffa
Merge pull request #15 from beyond10x/docs/readme-agents-trim
b10x-bot[bot] Sep 24, 2026
d7a7400
feat(b10x): read installed skills before a restart; plan for this host
b10x-bot[bot] Sep 24, 2026
759af0d
chore: release 0.13.1
b10x-bot[bot] Sep 24, 2026
4e159f5
Merge pull request #16 from beyond10x/feat/b10x-skill-command
b10x-bot[bot] Sep 24, 2026
3c7d402
feat!: every plugin here; init and upgrade on each; CLIs by cargo or …
b10x-bot[bot] Sep 24, 2026
055f068
chore: release 0.14.0
b10x-bot[bot] Sep 24, 2026
f03fa47
Merge pull request #17 from beyond10x/feat/init-upgrade-all-local
b10x-bot[bot] Sep 24, 2026
8602ac7
fix(release): smoke-test the init skill; release 0.14.1
b10x-bot[bot] Sep 24, 2026
7574cb8
Merge pull request #18 from beyond10x/fix/release-smoke-test
b10x-bot[bot] Sep 24, 2026
0a4f4bb
fix: trial 2 findings; release 0.14.2
b10x-bot[bot] Sep 24, 2026
6db632a
Merge pull request #19 from beyond10x/fix/trial-2-findings
b10x-bot[bot] Sep 24, 2026
973eeb2
plan: record retire-ess-specify as implemented
b10x-bot[bot] Sep 25, 2026
871fde3
Merge pull request #20 from beyond10x/plan/retire-ess-specify-closeout
b10x-bot[bot] Sep 25, 2026
28510ff
fix: trial 3 findings; release 0.14.3
b10x-bot[bot] Sep 25, 2026
027714d
Merge pull request #21 from beyond10x/fix/trial-3-findings
b10x-bot[bot] Sep 25, 2026
a5331fb
fix: round-3 leftovers and adopter findings; release 0.14.4
b10x-bot[bot] Sep 25, 2026
6c8de4a
Merge pull request #22 from beyond10x/fix/round-3-leftovers
b10x-bot[bot] Sep 25, 2026
ba01d4a
docs: tidy the 0.14.x changelog; release notes from CHANGELOG.md
b10x-bot[bot] Sep 25, 2026
1b266ca
Merge pull request #23 from beyond10x/docs/changelog-release-notes
b10x-bot[bot] Sep 25, 2026
c027e47
docs: the trial-finding label is the issue ledger again
b10x-bot[bot] Sep 25, 2026
6f7e92e
Merge pull request #24 from beyond10x/docs/ledger-label
b10x-bot[bot] Sep 25, 2026
e17c6cc
fix: round-4 trial findings; release 0.14.5
b10x-bot[bot] Sep 25, 2026
5a009ed
Merge pull request #25 from beyond10x/fix/0-14-5
b10x-bot[bot] Sep 25, 2026
b1518f6
fix: skills follow aep 0.59.3 and ess 0.32.0; release 0.14.6
b10x-bot[bot] Sep 25, 2026
bcbf56c
Merge pull request #26 from beyond10x/fix/0-14-6
b10x-bot[bot] Sep 25, 2026
d69a158
plan: rebuild the one-sentence onboarding story
b10x-bot[bot] Sep 25, 2026
9e90f55
feat: prebuilt is the default install method; release 0.14.7
b10x-bot[bot] Sep 25, 2026
8355049
Merge pull request #27 from beyond10x/fix/prebuilt-default
b10x-bot[bot] Sep 25, 2026
7c5adb4
feat: b10x-harness in the catalog, prebuilt by target; release 0.14.8
b10x-bot[bot] Sep 25, 2026
17ccede
Merge pull request #28 from beyond10x/feat/harness-catalog
b10x-bot[bot] Sep 25, 2026
a4bf1d2
feat: trial definitions, full-package ESS trial and measured baseline…
b10x-bot[bot] Sep 25, 2026
bd58f1a
Merge pull request #29 from beyond10x/feat/trial-measure
b10x-bot[bot] Sep 25, 2026
7d6ec9a
fix: round-5 trial findings and first baseline; release 0.14.10
b10x-bot[bot] Sep 25, 2026
55caae7
Merge pull request #30 from beyond10x/fix/round-5
b10x-bot[bot] Sep 25, 2026
152b14b
feat: daily freshness check, verified releases and tooling pins; rele…
b10x-bot[bot] Sep 25, 2026
2f82cbc
Merge pull request #31 from beyond10x/feat/freshness-pins
b10x-bot[bot] Sep 25, 2026
b38f8d9
docs(ess): skills start from an existing specification and plan per host
b10x-bot[bot] Sep 26, 2026
3c7c007
fix(b10x): a one-host plan says what the other host's plan would do
b10x-bot[bot] Sep 26, 2026
0574d1f
Merge impl/ess-skill-fixes into integrate/ap-issues-32-33
b10x-bot[bot] Sep 26, 2026
98099fb
Merge impl/init-other-host-plan into integrate/ap-issues-32-33
b10x-bot[bot] Sep 26, 2026
1ae57c1
chore: release 0.14.12
b10x-bot[bot] Sep 26, 2026
40158ad
Merge pull request #34 from beyond10x/integrate/ap-issues-32-33
b10x-bot[bot] Sep 26, 2026
5e08555
feat(ess): a catalogue of specification-hardening techniques
b10x-bot[bot] Sep 26, 2026
7ddf54f
chore: release 0.14.13
b10x-bot[bot] Sep 26, 2026
615292c
Merge pull request #36 from beyond10x/feat/ess-hardening-catalogue
b10x-bot[bot] Sep 26, 2026
5b4d6af
Merge release/0.10.0 (superseded) keeping main's tree
b10x-bot[bot] Sep 26, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
28 changes: 11 additions & 17 deletions .agents/plugins/marketplace.json
Original file line number Diff line number Diff line change
@@ -1,14 +1,14 @@
{
"name": "beyond10x",
"name": "b10x",
"interface": {
"displayName": "Beyond10x"
},
"plugins": [
{
"name": "beyond10x",
"name": "b10x",
"source": {
"source": "local",
"path": "./plugins/beyond10x"
"path": "./plugins/b10x"
},
"policy": {
"installation": "AVAILABLE",
Expand All @@ -17,10 +17,10 @@
"category": "Productivity"
},
{
"name": "aep-plan",
"name": "aep",
"source": {
"source": "local",
"path": "./plugins/aep-plan"
"path": "./plugins/aep"
},
"policy": {
"installation": "AVAILABLE",
Expand All @@ -29,10 +29,10 @@
"category": "Productivity"
},
{
"name": "aep-drive",
"name": "connectors",
"source": {
"source": "local",
"path": "./plugins/aep-drive"
"path": "./plugins/connectors"
},
"policy": {
"installation": "AVAILABLE",
Expand All @@ -41,10 +41,10 @@
"category": "Productivity"
},
{
"name": "ess-specify",
"name": "worktree",
"source": {
"source": "local",
"path": "./plugins/ess-specify"
"path": "./plugins/worktree"
},
"policy": {
"installation": "AVAILABLE",
Expand All @@ -53,22 +53,16 @@
"category": "Productivity"
},
{
"name": "workspace-hygiene",
"name": "ess",
"source": {
"source": "local",
"path": "./plugins/workspace-hygiene"
"path": "./plugins/ess"
},
"policy": {
"installation": "AVAILABLE",
"authentication": "ON_INSTALL"
},
"category": "Productivity"
},
{
"name": "connectors",
"source": { "source": "local", "path": "./plugins/connectors" },
"policy": { "installation": "AVAILABLE", "authentication": "ON_INSTALL" },
"category": "Productivity"
}
]
}
180 changes: 180 additions & 0 deletions .agents/skills/improving-by-trial/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,180 @@
---
name: improving-by-trial
description: Improve this repository's plugins by trial — run fresh, isolated agents that get only a user's sentence and this repository's link, collect where they got stuck, fix what belongs here, file what belongs elsewhere, and repeat until the trial passes. Use when asked to trial, dogfood, test the onboarding or a plugin end to end, run a fresh agent against the plugins, or iterate on trial findings before a release.
---

# Improving the plugins by trial

A trial is a fresh agent with nothing but a user's sentence and a link. It shows what the skills
fail to say. The gate cannot show that: every check passes while an agent still guesses.

## 1. Build a sandbox from this checkout

```console
task trial:sandbox NAME=<name>
```

This creates `/var/tmp/b10x-trials-$USER/<name>/` (the Taskfile's `TRIALS`) with its own `home/`, a `b10x` built from this
checkout, and an `env` file. The `env` file sets `HOME` to the sandbox and `B10X_MARKETPLACE` to
this checkout, so `b10x init` installs the plugins as they are in the working tree, before any
release. It also unsets `ANTHROPIC_API_KEY`, `TMPDIR` and `TMPPREFIX`. A defined trial
(`task trial:run TRIAL=<name>`, § 2) builds its own sandbox and fixture; for an ad-hoc one, put the
fixture it needs (a small service, a `TODO.md`, an OpenAPI document) under `work/`, and commit it
there with `git` when the trial needs history or a remote.

To trial the released version instead, remove `home/.local/bin/b10x` and unset `B10X_MARKETPLACE`
in `env`; the agent then follows `SETUP.md` from the release.

## 2. Run it headless, never as a sub-agent

The round's trials are defined in `trials/`, one directory each:

| file | holds |
|---|---|
| `trials/<name>/trial.yaml` | `name`, `kind`, `prompt`, optional `dir` (the run's subdirectory of `work/`), `fixture` (a directory beside it), `remote` (a bare `origin` in the sandbox), `seeded`, `setup` (shell lines run from the sandbox root with its `env` before the run), `measures`, and `outputs` (name → path below the run's directory) |
| `trials/<name>/fixture/` | the small service or backlog the trial starts from |
| `trials/baseline.json` | the measures of the last accepted run of each trial |

`task check` validates every definition. Run one by name:

```console
task trial:run TRIAL=<name>
```

This builds a fresh sandbox (`trial:sandbox`), copies the fixture into `work/<dir>` and commits it
there (`agentplugins-check trial-prepare`), runs the definition's `setup` (installing the plugins
under test with `b10x init … --out plan.json` and `b10x setup apply --plan plan.json --yes`, or
seeding an older release), then runs the prompt. An ad-hoc trial still runs in an existing sandbox:

```console
task trial:run NAME=<name> PROMPT='<the user sentence>' [DIR=<subdirectory of work>] [SEEDED=true]
```

The task copies the operator's credentials in (mode 600), runs `claude -p` with the sandbox's
`env`, `--strict-mcp-config` (no MCP server, including the account's claude.ai connectors) and
stream-json output into `run.jsonl`, and deletes the credentials when it finishes. The sandbox
`PATH` has `cargo` and `go`, and `go` is an allowed tool, so a trial can build and test an
implementation.

**Never run a trial as a sub-agent of the working session.** A sub-agent gets the session's own
agents and plugins whatever `HOME` its shell uses. In trial 3 the aep critics that ran were this
machine's, not the plugin's under test.

The prompt is what a user would type, plus the link when the trial is about finding the
repository. Ask the agent to end by listing the skills and agents it used, quoting any text that
confused it, and pasting the final command output verbatim. Give it no other hints.

## 3. Only an isolated run counts

`trial:run` checks isolation before it measures:

```console
agentplugins-check trial-isolation <sandbox>/run.jsonl --sandbox <sandbox> --version <version>
```

It fails when a plugin loaded from outside the sandbox home and the checkout, when a plugin is not
in the sandbox's own registry as it was before the run, when a plugin is not at the version under
test, or when the run used an agent or skill from a plugin the sandbox did not load. One that was
only offered (claude.ai account skills reach sub-agent sessions) is printed as a note. A failing run
is discarded, not interpreted.

**A sandbox under the home directory is not isolated.** Claude Code reads `CLAUDE.md`,
`.claude/CLAUDE.md` and `.claude/settings*.json` in every directory above its working directory. In
trial 3, sandboxes under `~/.cache` loaded the operator's `~/.claude/CLAUDE.md` (one agent followed
its commit rules) and a `.claude/settings.local.json` an earlier trial left in `~/.cache`, which
enabled `aep@b10x`. `trial:sandbox` refuses to build below any of those files; that is why `TRIALS`
is outside `$HOME`. Neither leak shows in the `init` event, so the placement is the only guard. An upgrade trial seeds older plugins on
purpose (`seeded: true` in its definition, or `SEEDED=true`), so plugins must match what the sandbox
was seeded with; a registry that did not exist before the run is read after it.

## 4. Read the run

`trial:run` ends with the numbers, one line each:

```console
agentplugins-check trial-report <sandbox>/run.jsonl --trial <name> [--baseline trials/baseline.json] [--write-baseline]
```

| measure | read from |
|---|---|
| `tool_calls` | `tool_use` blocks in the run |
| `validate` | the last `ess specify validate` the run ran: `valid`, not valid, or not run |
| `synthesis` | `N scenario(s) … M refusal(s)` in the last `ess verify conform synthesize` output |
| `unmapped` | `UNMAPPED:` markers in the YAML files the run wrote, read from disk, not from its prose |
| `outputs` | which of the definition's `outputs` exist (a directory counts when it is not empty) |
| `go_test` | passed, failed and skipped tests of the last `go test` (`-v` or `-json`); a package that does not build counts as a failure |

A trial reports the measures its definition lists; without `--trial`, an ad-hoc run gets every
measure but `outputs`. With `--baseline` it exits 1 when a measure got worse than the trial's entry:
validate stops passing, refusals or `go test` failures go up, a synthesis or `go test` that ran no
longer runs, an output the baseline had is missing, or tool calls rise by more than 50%. Other
changes print and pass. `--write-baseline` records the run as the trial's entry; do that for an
isolated run the round accepts, and commit `trials/baseline.json` with the fixes it led to.

The numbers say what happened, not why. From `run.jsonl`, also collect:

| collect | how |
|---|---|
| stuck points | every refusal and error in `tool_result`s, verbatim |
| guesses | what the final report says it inferred or could not find |
| waste | calls repeated, files read twice, commands that failed and were retried unchanged |
| the outcome | the verbatim `validate` / `generate` / plan output it pasted |

Before writing a finding into a skill, reproduce each claimed behaviour with the released CLI. A
trial agent's explanation of a refusal is a hypothesis.

## 5. Triage every finding to one owner

| the cause | where the fix goes |
|---|---|
| a skill, agent, `SETUP.md`, the catalog or the `b10x` CLI | here: fix it on a branch, with a test for CLI behaviour |
| a product CLI or language (`ess`, `aep`, `worktree`) | an issue in that repository, created through the bot (`AGENTS.md`); the skill here documents the workaround until it is fixed |
| both | both: the workaround here, the issue there, each naming the other |

The issue body has the trial name, the exact command and output, the expected behaviour and the
workaround the skill now documents, and ends with the line *Found by an agentplugins trial*. The
bot creates it with the label `trial-finding` (`"labels": ["trial-finding"]` in the request). The
label is the ledger:

```console
gh search issues --owner beyond10x --label trial-finding --state open # still broken
gh search issues --owner beyond10x --label trial-finding --state closed # fixed: released yet?
```

Check both lists before each round. Delete a workaround from the skills once its issue is fixed
and the fix is in a release, not when the issue closes.

## 6. Fix, gate, re-run

1. Fix on a managed worktree branch.
2. Run `task check`, and `agentplugins-check tools` when a CLI command or the ESS example changed.
3. Rebuild the sandbox (`task trial:sandbox`) and re-run the trial that found the problem. It
passes when the finding does not recur. New findings go back to step 5.
4. Release when every trial of the round passes.

Each round runs every trial in `trials/`: 4 ESS trials (`ess-new`, a new specification;
`ess-retrofit`, an existing service; `ess-pipeline`, generation plus a synthesized suite;
`ess-full-package`, every output plus a Go implementation held to the synthesized suite), and
`aep-backlog`, `worktree-onboarding` and `upgrade-seeded`. Change the domains and fixtures each
round so the agents cannot copy the previous answer from the skills; a changed trial starts a new
baseline entry.

### Every product release is re-verified

`verified.json` names, per CLI (`aep`, `ess`, `worktree`), the release the skills were last
verified against. The daily `agentplugins-check tools` run fails with one line per CLI whose newest
release is newer. Then:

1. Run `agentplugins-check tools` and fix every command it reports.
2. Run an ESS trial round (at least `ess-full-package`) against the new release.
3. Set the CLI to the new release in `verified.json` in the same pull request.

## 7. Clean up

`trial:run` deletes the credentials it copied. When the round is released, check that no
credentials are left in any sandbox and remove the sandboxes:

```console
ls /var/tmp/b10x-trials-$USER/*/home/.claude/.credentials.json
rm -rf /var/tmp/b10x-trials-$USER/<name>
```
44 changes: 22 additions & 22 deletions .claude-plugin/marketplace.json
Original file line number Diff line number Diff line change
@@ -1,36 +1,36 @@
{
"name": "beyond10x",
"owner": { "name": "beyond10x" },
"name": "b10x",
"owner": {
"name": "beyond10x"
},
"metadata": {
"description": "Beyond10x agent plugins: the plugins this repository carries, and product plugins pinned at their own releases."
},
"plugins": [
{
"name": "beyond10x",
"source": "./plugins/beyond10x",
"description": "Navigate Beyond10x resources and create portable Codex and Claude Code plugins."
"name": "b10x",
"source": "./plugins/b10x",
"description": "Set up, upgrade and check the Beyond10x plugins and binaries, route work to them, and create portable plugins."
},
{
"name": "aep-plan",
"source": "./plugins/aep-plan",
"description": "Govern AEP planning artifacts, decomposition, review, and reverse engineering."
"name": "aep",
"source": "./plugins/aep",
"description": "Plan governed work in the AEP artifact store and deliver it in reviewed waves: decomposition, plan critique, reverse engineering, story scoping, implementation and adversarial review."
},
{
"name": "aep-drive",
"source": "./plugins/aep-drive",
"description": "Coordinate AEP development waves, story scoping, implementation, and adversarial review."
},
{
"name": "ess-specify",
"source": "./plugins/ess-specify",
"description": "Validate ESS models and guide deterministic schema and OpenAPI projections."
"name": "connectors",
"source": "./plugins/connectors",
"description": "Set up, inspect, and invoke governed integrations through the connectors CLI."
},
{
"name": "workspace-hygiene",
"source": "./plugins/workspace-hygiene",
"description": "Manage isolated Git worktrees with explicit ownership and safe cleanup proof."
"name": "worktree",
"source": "./plugins/worktree",
"description": "Create, lease, finish and safely clean isolated Git worktrees through the worktree CLI."
},
{
"name": "connectors",
"source": "./plugins/connectors",
"description": "Set up, inspect, and invoke governed integrations through the connectors CLI."
"name": "ess",
"source": "./plugins/ess",
"description": "Write, retrofit, validate and project ESS specifications, and hold implementations to them with conformance suites."
}
]
}
1 change: 1 addition & 0 deletions .claude/skills/improving-by-trial
Loading
Loading