Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
{
"at": "2026-09-29T09:39:27Z",
"actor": "human:timo",
"artifact": "story:investigating-skill",
"kind": "story",
"revision": 5,
"change": {
"change": "evidence",
"kind": "test_result",
"source": "task check",
"reference": "feat/aep-investigating"
}
}
65 changes: 65 additions & 0 deletions .engineering/planning/story/investigating-skill.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
---
format: aep.planning-md/3
id: story:investigating-skill
kind: story
status: implemented
title: aep:investigating — evidence-based investigation of a live incident
summary: 'One activity skill with ten technique references for investigating what cannot be re-run: capture before remediation, sourced timeline, onset, peer differential, negative claims, ship state, hypothesis ledger, checking the check, complete reads, post-incident.'
scope:
- confidence: cited
path: crates/agentplugins-check/src/main.rs
- confidence: cited
path: plugins/aep/skills/investigating/SKILL.md
- confidence: cited
path: plugins/aep/skills/investigating/references/techniques.md
revision: 6
transitions:
- {from: "draft", to: "proposed", at: "2026-09-29T09:39:27Z", actor: "human:timo", revision: 4}
- {from: "proposed", to: "active", at: "2026-09-29T09:39:27Z", actor: "human:timo", revision: 5}
- {from: "active", to: "implemented", at: "2026-09-29T09:39:27Z", actor: "human:timo", revision: 6, decided_on: {"recorded":{"test_result":1}}}
---
# Story: aep:investigating

## Outcome

An agent asked to investigate a production incident, an outage or a "what happened / when did it
start / has it shipped" question works from evidence it can cite, preserves state before anyone
remediates, and labels every claim verified or inferred.

## Context

`aep:diagnosing` covers a defect that can be reproduced: it builds a red-capable loop first. A live
incident cannot be re-run, and the evidence disappears on remediation. On 2026-09-29 an operator's
hung-process report was remediated by a restart about four minutes after it was posted; the pod was
gone within a minute and no thread or lock state survived, so the root cause cannot now be
determined. The ten techniques come from failures of that kind recorded in an operator's knowledge
store over 2026-08 and 2026-09, each generalised with no system or customer named.

Placement: R2 permits no plugin without a product and CLI, and R3 names activities in `-ing` form,
so the ten techniques are one activity skill with a reference catalogue, as `ess:hardening` does.

## Acceptance

- `plugins/aep/skills/investigating/SKILL.md` exists, carries the `**Skill version**` line, and
names ten techniques, each with its procedure in `references/techniques.md`.
- `agentplugins-check`'s file list for `aep` includes the skill and its references; `task check`
exits 0.
- `aep:diagnosing` and `b10x:routing` route a live incident to `aep:investigating` and a
reproducible defect to `aep:diagnosing`.
- The README tree, `website/docs/plugins/aep.md`, `website/docs/structure.md`,
`website/docs/intro.md` and `evals/README.md` name the skill.

## Out of Scope

- An eval case with a recorded transcript. `evals/README.md` lists the skill as uncovered.
- An agent role. The skill runs in the main session.
- Any tooling for a specific platform (Kubernetes, a particular metrics store). The techniques name
the kind of instrument, with common examples.

## Ambiguities

None.

## Open Questions

None.
19 changes: 19 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,24 @@
# Changelog

## [0.19.0] — 2026-09-29

A new skill investigates what cannot be re-run: a production incident, an outage, or a question
about a running system. `aep:diagnosing` builds a failing command first. `aep:investigating`
keeps the evidence and cites where every fact came from.

- New skill: `aep:investigating`, with ten techniques in `references/techniques.md`: capture
before remediation, a sourced UTC timeline, onset from state, peer differential, negative
claims, ship state, a hypothesis ledger, checking the check, complete reads, and the
post-incident report. Each technique names the failure it was written from. The investigation
is an `incident-report` artifact, observations are `health_observation` evidence, and each
follow-up is a draft story.
- `aep:diagnosing` and `b10x:routing` send an incident that cannot be re-run to
`aep:investigating`.
- No eval case covers the new skill yet; `evals/README.md` lists it with the uncovered ones.
- `verified.json` still pins aep 0.64.0 and ess 0.39.0. aep 0.65.0 and ess 0.42.0 are not yet
re-verified: `agentplugins-check tools` reports `aep plan artifact divergences`, spelled in
`aep:planning`, as not a command of aep 0.65.0.

## [0.18.0] — 2026-09-28

A second tutorial continues the first: the ESS tutorial's library gets an AEP plan, four critics
Expand Down
4 changes: 2 additions & 2 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ resolver = "2"
members = ["crates/agentplugins-check", "crates/b10x"]

[workspace.package]
version = "0.18.0"
version = "0.19.0"
edition = "2021"
rust-version = "1.85"
license = "Apache-2.0"
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ What each plugin ships:
- skills: [`init`](plugins/ess/skills/init/SKILL.md) · [`upgrade`](plugins/ess/skills/upgrade/SKILL.md) · [`hardening`](plugins/ess/skills/hardening/SKILL.md) · [`retrofitting`](plugins/ess/skills/retrofitting/SKILL.md) · [`specifying`](plugins/ess/skills/specifying/SKILL.md) · [`testing-conformance`](plugins/ess/skills/testing-conformance/SKILL.md)
- agents: [`author`](plugins/ess/agents/author.md) · [`conformance`](plugins/ess/agents/conformance.md) · [`retrofitter`](plugins/ess/agents/retrofitter.md)
- [`aep`](plugins/aep/) · [docs](website/docs/plugins/aep.md)
- skills: [`init`](plugins/aep/skills/init/SKILL.md) · [`upgrade`](plugins/aep/skills/upgrade/SKILL.md) · [`decompose`](plugins/aep/skills/decompose/SKILL.md) · [`diagnosing`](plugins/aep/skills/diagnosing/SKILL.md) · [`drive`](plugins/aep/skills/drive/SKILL.md) · [`implementing`](plugins/aep/skills/implementing/SKILL.md) · [`migrating`](plugins/aep/skills/migrating/SKILL.md) · [`planning`](plugins/aep/skills/planning/SKILL.md) · [`review-plan`](plugins/aep/skills/review-plan/SKILL.md) · [`wave`](plugins/aep/skills/wave/SKILL.md)
- skills: [`init`](plugins/aep/skills/init/SKILL.md) · [`upgrade`](plugins/aep/skills/upgrade/SKILL.md) · [`decompose`](plugins/aep/skills/decompose/SKILL.md) · [`diagnosing`](plugins/aep/skills/diagnosing/SKILL.md) · [`drive`](plugins/aep/skills/drive/SKILL.md) · [`implementing`](plugins/aep/skills/implementing/SKILL.md) · [`investigating`](plugins/aep/skills/investigating/SKILL.md) · [`migrating`](plugins/aep/skills/migrating/SKILL.md) · [`planning`](plugins/aep/skills/planning/SKILL.md) · [`review-plan`](plugins/aep/skills/review-plan/SKILL.md) · [`wave`](plugins/aep/skills/wave/SKILL.md)
- agents: [`adversary`](plugins/aep/agents/adversary.md) · [`decomposer`](plugins/aep/agents/decomposer.md) · [`implementor`](plugins/aep/agents/implementor.md) · [`plan-critic-acceptance`](plugins/aep/agents/plan-critic-acceptance.md) · [`plan-critic-design`](plugins/aep/agents/plan-critic-design.md) · [`plan-critic-parallel-safety`](plugins/aep/agents/plan-critic-parallel-safety.md) · [`plan-critic-scope`](plugins/aep/agents/plan-critic-scope.md) · [`plan-reviewer`](plugins/aep/agents/plan-reviewer.md) · [`reverse-engineer`](plugins/aep/agents/reverse-engineer.md) · [`security-reviewer`](plugins/aep/agents/security-reviewer.md) · [`story-scoper`](plugins/aep/agents/story-scoper.md)
- [`worktree`](plugins/worktree/) · [docs](website/docs/plugins/worktree.md)
- skills: [`init`](plugins/worktree/skills/init/SKILL.md) · [`upgrade`](plugins/worktree/skills/upgrade/SKILL.md) · [`cleanup`](plugins/worktree/skills/cleanup/SKILL.md) · [`managing-worktrees`](plugins/worktree/skills/managing-worktrees/SKILL.md)
Expand Down
2 changes: 2 additions & 0 deletions crates/agentplugins-check/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,8 @@ const PLUGINS: &[(&str, &[&str])] = &[
"skills/implementing/SKILL.md",
"skills/implementing/references/drive.md",
"skills/diagnosing/SKILL.md",
"skills/investigating/SKILL.md",
"skills/investigating/references/techniques.md",
"skills/wave/SKILL.md",
"skills/drive/SKILL.md",
"skills/review-plan/SKILL.md",
Expand Down
4 changes: 2 additions & 2 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ A case is four things and no others:
Every cell in the middle column is that case's `subject:` in full, read from its `case.yaml`, and
those fields — not this table and not any prose elsewhere — are the source of truth for what the
corpus covers: counted from them the seventeen cases name **8 of this repository's 14 agents** and
**6 of its 21 skills** (the 6 command skills are not counted).
**6 of its 22 skills** (the 6 command skills are not counted).

| Case | `subject:` agents and skills | The claim it holds the subject to |
|---|---|---|
Expand All @@ -40,7 +40,7 @@ corpus covers: counted from them the seventeen cases name **8 of this repository
| `authoring-prohibition-only-rule` | `b10x:authoring-plugins` | a wording review read the seeded skill, edited nothing, and named the prohibition with a positive target |

No case names the agents `aep:implementor`, `aep:plan-reviewer`, `aep:reverse-engineer`,
`ess:author`, `ess:conformance` or `ess:retrofitter`, nor the activity skills `aep:migrating`,
`ess:author`, `ess:conformance` or `ess:retrofitter`, nor the activity skills `aep:migrating`, `aep:investigating`,
`b10x:routing`, `ess:retrofitting`, `ess:testing-conformance`, `ess:hardening` or
`worktree:managing-worktrees`, nor any `init` or `upgrade` skill; a change to one of them turns no
row red. The `aep` ones are the remaining scope of `story:plugin-eval-cases` in
Expand Down
2 changes: 1 addition & 1 deletion plugins/aep/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
"name": "aep",
"displayName": "AEP",
"description": "Plan governed work in the AEP artifact store and deliver it in reviewed waves: decomposition, plan critique, reverse engineering, story scoping, implementation and adversarial review.",
"version": "0.18.0",
"version": "0.19.0",
"author": {
"name": "Beyond10x"
},
Expand Down
2 changes: 1 addition & 1 deletion plugins/aep/.codex-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "aep",
"version": "0.18.0",
"version": "0.19.0",
"description": "Plan governed work in the AEP artifact store and deliver it in reviewed waves.",
"author": {
"name": "Beyond10x"
Expand Down
4 changes: 2 additions & 2 deletions plugins/aep/skills/diagnosing/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
---
name: diagnosing
description: Diagnose a hard bug or a performance regression by building a red-capable feedback loop before any hypothesis, then ranked falsifiable hypotheses, one-variable probes, a regression test at the right seam, and evidence recorded in the AEP store. Use when the user says diagnose, debug, "why is this failing", "this is slow", or reports something broken, throwing, flaky or slower than before. Not for a failing CI job whose cause is already named in its log, and not for raising conformance coverage, which is `ess:testing-conformance`.
description: Diagnose a hard bug or a performance regression by building a red-capable feedback loop before any hypothesis, then ranked falsifiable hypotheses, one-variable probes, a regression test at the right seam, and evidence recorded in the AEP store. Use when the user says diagnose, debug, "why is this failing", "this is slow", or reports something broken, throwing, flaky or slower than before. Not for a production incident that cannot be re-run, which is `aep:investigating`; not for a failing CI job whose cause is already named in its log; and not for raising conformance coverage, which is `ess:testing-conformance`.
---

**Skill version 0.18.0** — the version in `.claude-plugin/plugin.json`.
**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`.

# Diagnosing a failure

Expand Down
2 changes: 1 addition & 1 deletion plugins/aep/skills/implementing/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ name: implementing
description: Implement accepted AEP work, in one of two modes. A wave picks the stories that can be implemented at once, proposes the wave for approval, dispatches one implementor per story into its own worktree, sends each result to the adversary and merges what goes green. A drive hands one story to a governed `metaharness aep drive` run and reports the run id. Use when the operator asks to implement, build or deliver planned stories, to pick or start the next wave, to implement several stories in parallel or fan out across sub-agents, to drive a story or start a governed run, or asks why a wave's rules are instructions and a drive's are enforced. A wave proposes first and stops; a drive starts one run and reports; neither moves an artifact itself.
---

**Skill version 0.18.0** — the version in `.claude-plugin/plugin.json`; a wave's stage-1 proposal quotes it.
**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`; a wave's stage-1 proposal quotes it.

# Implementing accepted work

Expand Down
109 changes: 109 additions & 0 deletions plugins/aep/skills/investigating/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
---
name: investigating
description: >-
Investigate a production incident, an outage or a question about a running system from evidence that can be cited — capture state before anyone remediates, build a sourced UTC timeline, date an onset from an instrument that can see a negative, compare against a healthy peer, and label every claim verified or inferred, with a catalogue of ten techniques. Use when the user says investigate, "what happened", "when did this start", "is it still happening", "has this shipped", "is this deployed", reports an alert, an outage, a hung or crashing process or a customer-visible failure, or asks for an incident report or a postmortem. Not for a defect that can be reproduced on demand, which is `aep:diagnosing`; not for checking a change before it merges, which is `aep:implementing`.
---

**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`.

# Investigating a live system

`aep:diagnosing` starts by building a command that goes red on the bug. A live incident has no such
command: it happened once, on a system that keeps moving, and the evidence is destroyed by the
restart that ends it. Here the work is **keeping the evidence and saying where each fact came
from**, because nothing can be re-run later to check it.

Redact every secret and every personal identifier (a phone number, an email address, an account
holder's name) in what you show and what you write: `<REDACTED>` in its place.

## The rule that makes any of this count

**Every specific is a claim, and every claim carries its source.** A time, a count, a version, a
duration, a name and a "nothing happened" are each claims. Next to each one, write the command whose
output it came from, or the file and line. A claim you did not read from a command or a file is
labelled **inferred**, or it is not written.

Two consequences:

- **An absence is a claim like any other.** "No alert fired", "no tag exists" and "nobody asked"
need the query that returned nothing, run against the system that owns the answer — technique 5.
- **An unverified detail that makes the finding look bigger is the one to distrust first.** That is
the direction plausible-but-wrong details run in.

## The catalogue

| # | technique | the question | what it prevents |
|---|---|---|---|
| 1 | capture before remediation | what state exists now that a restart will destroy? | a root cause that becomes unknowable because the hung process was restarted before anyone looked |
| 2 | sourced timeline | what happened, in which order, according to what? | an effect placed before its cause by a time-zone slip; a timeline nobody can re-check |
| 3 | onset from state | when did this really start? | an onset dated from the first delivered alert, weeks after the condition began |
| 4 | peer differential | does this signal also appear on something healthy? | a routine warning mistaken for the cause because it was the loudest line in the suspect's log |
| 5 | negative claims | does this really not exist / not happen? | "no fix exists" while the fix was merged; a delta scan read as an inventory |
| 6 | ship state | is the change actually running there? | a ticket status, a version range or an image tag name read as deployment |
| 7 | hypothesis ledger | which claims are observed and which are read off code? | a cause, a frequency or a severity asserted from reading a code path |
| 8 | check the check | what does this green probe, gate or scanner actually cover? | "Ready" or "exit 0" cited as evidence about something the check never looked at |
| 9 | complete reads | did the read return everything? | a truncated read treated as the whole file, and written back |
| 10 | post-incident | what was the impact, why was it not detected, and what was lost? | follow-ups that live only in a chat thread |

Procedures, with the commands and the failure each one was written from:
[references/techniques.md](references/techniques.md).

## The order

1. **Is it still happening?** If yes, run **technique 1 first**, before any other reading and before
anyone restarts, redeploys, drains or kills anything. Remediation is the operator's decision;
capture takes seconds and is yours. If someone has already remediated, say so, and record what
was lost.
2. **Open the record.** In a repository with an AEP store:

```console
$ aep plan artifact new incident-report <slug> --title "<the symptom, in the reporter's words>"
```

Outside one, a markdown file in the project's incident location. Either way the capture from step
1 goes beside it.
3. **Timeline** (2), then **onset** (3). The timeline is the spine every later step writes onto.
4. **Peer differential** (4) on every anomaly the timeline turns up, before it becomes a hypothesis.
5. **Hypothesis ledger** (7): 3–5 ranked candidate causes, each with the observation that would
refute it. Test with read-only probes. When a hypothesis leads to code that can be exercised,
hand that part to `aep:diagnosing`.
6. **Techniques 5, 6, 8 and 9** whenever a claim of their kind is about to be written: an absence, a
ship state, a green check, or a read that feeds a decision.
7. **Post-incident** (10) once the system is stable.

Record each load-bearing observation against the incident as it is made:

```console
$ aep plan artifact evidence incident-report:<slug> --kind health_observation \
--source "<the command>" --ref "<the capture file or a URL>" --at <instant, UTC>
```

`metric_observation` for a metric series, `deployment_result` for what a deployment actually ran.

## Boundaries

- **Read-only by default.** A read of a live system (logs, metrics, process state, an API `GET`)
is yours to run. A restart, a rollback, a scale change, a configuration write, a post to a shared
channel or a ticket write is the operator's: propose it, with the evidence, and wait.
- A debugger attached to a live process pauses it. Say so, and prefer a non-pausing read
(`/proc/<pid>/task/*/stack`, a runtime's own dump signal) when the process is still serving.
- Never put a secret or a personal identifier into a query that is logged, a file that is committed
or a message that is posted.

## Reporting

Verdict first: **resolved / ongoing / unknown**, and **root cause known / hypothesis / unknown**.
Then the timeline table, each row with its source. Then the ledger: verified claims, inferred claims,
refuted hypotheses. Then what evidence was lost, and the detection gap. Then the follow-ups, each one
an artifact:

```console
$ aep plan artifact new story <slug> --title "<the gap>" --relate informed_by:incident-report:<slug>
```

## Next

- A hypothesis points at code that can be exercised: `aep:diagnosing`.
- A follow-up story is ready to build: `aep:implementing`.
- The incident needs a written postmortem: technique 10, then
`aep plan artifact new postmortem <slug> --title "<title>" --relate derived_from:incident-report:<slug>`.
Loading
Loading