diff --git a/.engineering/evidence/story/investigating-skill/20260929T093927Z-000-8cc052d59502.json b/.engineering/evidence/story/investigating-skill/20260929T093927Z-000-8cc052d59502.json new file mode 100644 index 0000000..e3a444f --- /dev/null +++ b/.engineering/evidence/story/investigating-skill/20260929T093927Z-000-8cc052d59502.json @@ -0,0 +1,13 @@ +{ + "at": "2026-09-29T09:39:27Z", + "actor": "human:timo", + "artifact": "story:investigating-skill", + "kind": "story", + "revision": 5, + "change": { + "change": "evidence", + "kind": "test_result", + "source": "task check", + "reference": "feat/aep-investigating" + } +} diff --git a/.engineering/planning/story/investigating-skill.md b/.engineering/planning/story/investigating-skill.md new file mode 100644 index 0000000..ce9644e --- /dev/null +++ b/.engineering/planning/story/investigating-skill.md @@ -0,0 +1,65 @@ +--- +format: aep.planning-md/3 +id: story:investigating-skill +kind: story +status: implemented +title: aep:investigating — evidence-based investigation of a live incident +summary: 'One activity skill with ten technique references for investigating what cannot be re-run: capture before remediation, sourced timeline, onset, peer differential, negative claims, ship state, hypothesis ledger, checking the check, complete reads, post-incident.' +scope: +- confidence: cited + path: crates/agentplugins-check/src/main.rs +- confidence: cited + path: plugins/aep/skills/investigating/SKILL.md +- confidence: cited + path: plugins/aep/skills/investigating/references/techniques.md +revision: 6 +transitions: +- {from: "draft", to: "proposed", at: "2026-09-29T09:39:27Z", actor: "human:timo", revision: 4} +- {from: "proposed", to: "active", at: "2026-09-29T09:39:27Z", actor: "human:timo", revision: 5} +- {from: "active", to: "implemented", at: "2026-09-29T09:39:27Z", actor: "human:timo", revision: 6, decided_on: {"recorded":{"test_result":1}}} +--- +# Story: aep:investigating + +## Outcome + +An agent asked to investigate a production incident, an outage or a "what happened / when did it +start / has it shipped" question works from evidence it can cite, preserves state before anyone +remediates, and labels every claim verified or inferred. + +## Context + +`aep:diagnosing` covers a defect that can be reproduced: it builds a red-capable loop first. A live +incident cannot be re-run, and the evidence disappears on remediation. On 2026-09-29 an operator's +hung-process report was remediated by a restart about four minutes after it was posted; the pod was +gone within a minute and no thread or lock state survived, so the root cause cannot now be +determined. The ten techniques come from failures of that kind recorded in an operator's knowledge +store over 2026-08 and 2026-09, each generalised with no system or customer named. + +Placement: R2 permits no plugin without a product and CLI, and R3 names activities in `-ing` form, +so the ten techniques are one activity skill with a reference catalogue, as `ess:hardening` does. + +## Acceptance + +- `plugins/aep/skills/investigating/SKILL.md` exists, carries the `**Skill version**` line, and + names ten techniques, each with its procedure in `references/techniques.md`. +- `agentplugins-check`'s file list for `aep` includes the skill and its references; `task check` + exits 0. +- `aep:diagnosing` and `b10x:routing` route a live incident to `aep:investigating` and a + reproducible defect to `aep:diagnosing`. +- The README tree, `website/docs/plugins/aep.md`, `website/docs/structure.md`, + `website/docs/intro.md` and `evals/README.md` name the skill. + +## Out of Scope + +- An eval case with a recorded transcript. `evals/README.md` lists the skill as uncovered. +- An agent role. The skill runs in the main session. +- Any tooling for a specific platform (Kubernetes, a particular metrics store). The techniques name + the kind of instrument, with common examples. + +## Ambiguities + +None. + +## Open Questions + +None. diff --git a/CHANGELOG.md b/CHANGELOG.md index ce7455e..4e87503 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,24 @@ # Changelog +## [0.19.0] — 2026-09-29 + +A new skill investigates what cannot be re-run: a production incident, an outage, or a question +about a running system. `aep:diagnosing` builds a failing command first. `aep:investigating` +keeps the evidence and cites where every fact came from. + +- New skill: `aep:investigating`, with ten techniques in `references/techniques.md`: capture + before remediation, a sourced UTC timeline, onset from state, peer differential, negative + claims, ship state, a hypothesis ledger, checking the check, complete reads, and the + post-incident report. Each technique names the failure it was written from. The investigation + is an `incident-report` artifact, observations are `health_observation` evidence, and each + follow-up is a draft story. +- `aep:diagnosing` and `b10x:routing` send an incident that cannot be re-run to + `aep:investigating`. +- No eval case covers the new skill yet; `evals/README.md` lists it with the uncovered ones. +- `verified.json` still pins aep 0.64.0 and ess 0.39.0. aep 0.65.0 and ess 0.42.0 are not yet + re-verified: `agentplugins-check tools` reports `aep plan artifact divergences`, spelled in + `aep:planning`, as not a command of aep 0.65.0. + ## [0.18.0] — 2026-09-28 A second tutorial continues the first: the ESS tutorial's library gets an AEP plan, four critics diff --git a/Cargo.lock b/Cargo.lock index 1e7ead1..b67ef37 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -4,7 +4,7 @@ version = 4 [[package]] name = "agentplugins-check" -version = "0.18.0" +version = "0.19.0" dependencies = [ "clap", "serde", @@ -76,7 +76,7 @@ dependencies = [ [[package]] name = "b10x" -version = "0.18.0" +version = "0.19.0" dependencies = [ "clap", "serde", diff --git a/Cargo.toml b/Cargo.toml index 5ce4e99..8f44224 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -3,7 +3,7 @@ resolver = "2" members = ["crates/agentplugins-check", "crates/b10x"] [workspace.package] -version = "0.18.0" +version = "0.19.0" edition = "2021" rust-version = "1.85" license = "Apache-2.0" diff --git a/README.md b/README.md index b4a22ed..0a27dfc 100644 --- a/README.md +++ b/README.md @@ -20,7 +20,7 @@ What each plugin ships: - skills: [`init`](plugins/ess/skills/init/SKILL.md) · [`upgrade`](plugins/ess/skills/upgrade/SKILL.md) · [`hardening`](plugins/ess/skills/hardening/SKILL.md) · [`retrofitting`](plugins/ess/skills/retrofitting/SKILL.md) · [`specifying`](plugins/ess/skills/specifying/SKILL.md) · [`testing-conformance`](plugins/ess/skills/testing-conformance/SKILL.md) - agents: [`author`](plugins/ess/agents/author.md) · [`conformance`](plugins/ess/agents/conformance.md) · [`retrofitter`](plugins/ess/agents/retrofitter.md) - [`aep`](plugins/aep/) · [docs](website/docs/plugins/aep.md) - - skills: [`init`](plugins/aep/skills/init/SKILL.md) · [`upgrade`](plugins/aep/skills/upgrade/SKILL.md) · [`decompose`](plugins/aep/skills/decompose/SKILL.md) · [`diagnosing`](plugins/aep/skills/diagnosing/SKILL.md) · [`drive`](plugins/aep/skills/drive/SKILL.md) · [`implementing`](plugins/aep/skills/implementing/SKILL.md) · [`migrating`](plugins/aep/skills/migrating/SKILL.md) · [`planning`](plugins/aep/skills/planning/SKILL.md) · [`review-plan`](plugins/aep/skills/review-plan/SKILL.md) · [`wave`](plugins/aep/skills/wave/SKILL.md) + - skills: [`init`](plugins/aep/skills/init/SKILL.md) · [`upgrade`](plugins/aep/skills/upgrade/SKILL.md) · [`decompose`](plugins/aep/skills/decompose/SKILL.md) · [`diagnosing`](plugins/aep/skills/diagnosing/SKILL.md) · [`drive`](plugins/aep/skills/drive/SKILL.md) · [`implementing`](plugins/aep/skills/implementing/SKILL.md) · [`investigating`](plugins/aep/skills/investigating/SKILL.md) · [`migrating`](plugins/aep/skills/migrating/SKILL.md) · [`planning`](plugins/aep/skills/planning/SKILL.md) · [`review-plan`](plugins/aep/skills/review-plan/SKILL.md) · [`wave`](plugins/aep/skills/wave/SKILL.md) - agents: [`adversary`](plugins/aep/agents/adversary.md) · [`decomposer`](plugins/aep/agents/decomposer.md) · [`implementor`](plugins/aep/agents/implementor.md) · [`plan-critic-acceptance`](plugins/aep/agents/plan-critic-acceptance.md) · [`plan-critic-design`](plugins/aep/agents/plan-critic-design.md) · [`plan-critic-parallel-safety`](plugins/aep/agents/plan-critic-parallel-safety.md) · [`plan-critic-scope`](plugins/aep/agents/plan-critic-scope.md) · [`plan-reviewer`](plugins/aep/agents/plan-reviewer.md) · [`reverse-engineer`](plugins/aep/agents/reverse-engineer.md) · [`security-reviewer`](plugins/aep/agents/security-reviewer.md) · [`story-scoper`](plugins/aep/agents/story-scoper.md) - [`worktree`](plugins/worktree/) · [docs](website/docs/plugins/worktree.md) - skills: [`init`](plugins/worktree/skills/init/SKILL.md) · [`upgrade`](plugins/worktree/skills/upgrade/SKILL.md) · [`cleanup`](plugins/worktree/skills/cleanup/SKILL.md) · [`managing-worktrees`](plugins/worktree/skills/managing-worktrees/SKILL.md) diff --git a/crates/agentplugins-check/src/main.rs b/crates/agentplugins-check/src/main.rs index bc0f448..02eadd1 100644 --- a/crates/agentplugins-check/src/main.rs +++ b/crates/agentplugins-check/src/main.rs @@ -48,6 +48,8 @@ const PLUGINS: &[(&str, &[&str])] = &[ "skills/implementing/SKILL.md", "skills/implementing/references/drive.md", "skills/diagnosing/SKILL.md", + "skills/investigating/SKILL.md", + "skills/investigating/references/techniques.md", "skills/wave/SKILL.md", "skills/drive/SKILL.md", "skills/review-plan/SKILL.md", diff --git a/evals/README.md b/evals/README.md index 53714bf..609dee7 100644 --- a/evals/README.md +++ b/evals/README.md @@ -17,7 +17,7 @@ A case is four things and no others: Every cell in the middle column is that case's `subject:` in full, read from its `case.yaml`, and those fields — not this table and not any prose elsewhere — are the source of truth for what the corpus covers: counted from them the seventeen cases name **8 of this repository's 14 agents** and -**6 of its 21 skills** (the 6 command skills are not counted). +**6 of its 22 skills** (the 6 command skills are not counted). | Case | `subject:` agents and skills | The claim it holds the subject to | |---|---|---| @@ -40,7 +40,7 @@ corpus covers: counted from them the seventeen cases name **8 of this repository | `authoring-prohibition-only-rule` | `b10x:authoring-plugins` | a wording review read the seeded skill, edited nothing, and named the prohibition with a positive target | No case names the agents `aep:implementor`, `aep:plan-reviewer`, `aep:reverse-engineer`, -`ess:author`, `ess:conformance` or `ess:retrofitter`, nor the activity skills `aep:migrating`, +`ess:author`, `ess:conformance` or `ess:retrofitter`, nor the activity skills `aep:migrating`, `aep:investigating`, `b10x:routing`, `ess:retrofitting`, `ess:testing-conformance`, `ess:hardening` or `worktree:managing-worktrees`, nor any `init` or `upgrade` skill; a change to one of them turns no row red. The `aep` ones are the remaining scope of `story:plugin-eval-cases` in diff --git a/plugins/aep/.claude-plugin/plugin.json b/plugins/aep/.claude-plugin/plugin.json index c8bfb3d..25f1d9c 100644 --- a/plugins/aep/.claude-plugin/plugin.json +++ b/plugins/aep/.claude-plugin/plugin.json @@ -2,7 +2,7 @@ "name": "aep", "displayName": "AEP", "description": "Plan governed work in the AEP artifact store and deliver it in reviewed waves: decomposition, plan critique, reverse engineering, story scoping, implementation and adversarial review.", - "version": "0.18.0", + "version": "0.19.0", "author": { "name": "Beyond10x" }, diff --git a/plugins/aep/.codex-plugin/plugin.json b/plugins/aep/.codex-plugin/plugin.json index 64c3b24..5d023a0 100644 --- a/plugins/aep/.codex-plugin/plugin.json +++ b/plugins/aep/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "aep", - "version": "0.18.0", + "version": "0.19.0", "description": "Plan governed work in the AEP artifact store and deliver it in reviewed waves.", "author": { "name": "Beyond10x" diff --git a/plugins/aep/skills/diagnosing/SKILL.md b/plugins/aep/skills/diagnosing/SKILL.md index e646c4f..048355c 100644 --- a/plugins/aep/skills/diagnosing/SKILL.md +++ b/plugins/aep/skills/diagnosing/SKILL.md @@ -1,9 +1,9 @@ --- name: diagnosing -description: Diagnose a hard bug or a performance regression by building a red-capable feedback loop before any hypothesis, then ranked falsifiable hypotheses, one-variable probes, a regression test at the right seam, and evidence recorded in the AEP store. Use when the user says diagnose, debug, "why is this failing", "this is slow", or reports something broken, throwing, flaky or slower than before. Not for a failing CI job whose cause is already named in its log, and not for raising conformance coverage, which is `ess:testing-conformance`. +description: Diagnose a hard bug or a performance regression by building a red-capable feedback loop before any hypothesis, then ranked falsifiable hypotheses, one-variable probes, a regression test at the right seam, and evidence recorded in the AEP store. Use when the user says diagnose, debug, "why is this failing", "this is slow", or reports something broken, throwing, flaky or slower than before. Not for a production incident that cannot be re-run, which is `aep:investigating`; not for a failing CI job whose cause is already named in its log; and not for raising conformance coverage, which is `ess:testing-conformance`. --- -**Skill version 0.18.0** — the version in `.claude-plugin/plugin.json`. +**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`. # Diagnosing a failure diff --git a/plugins/aep/skills/implementing/SKILL.md b/plugins/aep/skills/implementing/SKILL.md index ba5e690..b87068f 100644 --- a/plugins/aep/skills/implementing/SKILL.md +++ b/plugins/aep/skills/implementing/SKILL.md @@ -3,7 +3,7 @@ name: implementing description: Implement accepted AEP work, in one of two modes. A wave picks the stories that can be implemented at once, proposes the wave for approval, dispatches one implementor per story into its own worktree, sends each result to the adversary and merges what goes green. A drive hands one story to a governed `metaharness aep drive` run and reports the run id. Use when the operator asks to implement, build or deliver planned stories, to pick or start the next wave, to implement several stories in parallel or fan out across sub-agents, to drive a story or start a governed run, or asks why a wave's rules are instructions and a drive's are enforced. A wave proposes first and stops; a drive starts one run and reports; neither moves an artifact itself. --- -**Skill version 0.18.0** — the version in `.claude-plugin/plugin.json`; a wave's stage-1 proposal quotes it. +**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`; a wave's stage-1 proposal quotes it. # Implementing accepted work diff --git a/plugins/aep/skills/investigating/SKILL.md b/plugins/aep/skills/investigating/SKILL.md new file mode 100644 index 0000000..c73c9e1 --- /dev/null +++ b/plugins/aep/skills/investigating/SKILL.md @@ -0,0 +1,109 @@ +--- +name: investigating +description: >- + Investigate a production incident, an outage or a question about a running system from evidence that can be cited — capture state before anyone remediates, build a sourced UTC timeline, date an onset from an instrument that can see a negative, compare against a healthy peer, and label every claim verified or inferred, with a catalogue of ten techniques. Use when the user says investigate, "what happened", "when did this start", "is it still happening", "has this shipped", "is this deployed", reports an alert, an outage, a hung or crashing process or a customer-visible failure, or asks for an incident report or a postmortem. Not for a defect that can be reproduced on demand, which is `aep:diagnosing`; not for checking a change before it merges, which is `aep:implementing`. +--- + +**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`. + +# Investigating a live system + +`aep:diagnosing` starts by building a command that goes red on the bug. A live incident has no such +command: it happened once, on a system that keeps moving, and the evidence is destroyed by the +restart that ends it. Here the work is **keeping the evidence and saying where each fact came +from**, because nothing can be re-run later to check it. + +Redact every secret and every personal identifier (a phone number, an email address, an account +holder's name) in what you show and what you write: `` in its place. + +## The rule that makes any of this count + +**Every specific is a claim, and every claim carries its source.** A time, a count, a version, a +duration, a name and a "nothing happened" are each claims. Next to each one, write the command whose +output it came from, or the file and line. A claim you did not read from a command or a file is +labelled **inferred**, or it is not written. + +Two consequences: + +- **An absence is a claim like any other.** "No alert fired", "no tag exists" and "nobody asked" + need the query that returned nothing, run against the system that owns the answer — technique 5. +- **An unverified detail that makes the finding look bigger is the one to distrust first.** That is + the direction plausible-but-wrong details run in. + +## The catalogue + +| # | technique | the question | what it prevents | +|---|---|---|---| +| 1 | capture before remediation | what state exists now that a restart will destroy? | a root cause that becomes unknowable because the hung process was restarted before anyone looked | +| 2 | sourced timeline | what happened, in which order, according to what? | an effect placed before its cause by a time-zone slip; a timeline nobody can re-check | +| 3 | onset from state | when did this really start? | an onset dated from the first delivered alert, weeks after the condition began | +| 4 | peer differential | does this signal also appear on something healthy? | a routine warning mistaken for the cause because it was the loudest line in the suspect's log | +| 5 | negative claims | does this really not exist / not happen? | "no fix exists" while the fix was merged; a delta scan read as an inventory | +| 6 | ship state | is the change actually running there? | a ticket status, a version range or an image tag name read as deployment | +| 7 | hypothesis ledger | which claims are observed and which are read off code? | a cause, a frequency or a severity asserted from reading a code path | +| 8 | check the check | what does this green probe, gate or scanner actually cover? | "Ready" or "exit 0" cited as evidence about something the check never looked at | +| 9 | complete reads | did the read return everything? | a truncated read treated as the whole file, and written back | +| 10 | post-incident | what was the impact, why was it not detected, and what was lost? | follow-ups that live only in a chat thread | + +Procedures, with the commands and the failure each one was written from: +[references/techniques.md](references/techniques.md). + +## The order + +1. **Is it still happening?** If yes, run **technique 1 first**, before any other reading and before + anyone restarts, redeploys, drains or kills anything. Remediation is the operator's decision; + capture takes seconds and is yours. If someone has already remediated, say so, and record what + was lost. +2. **Open the record.** In a repository with an AEP store: + + ```console + $ aep plan artifact new incident-report --title "" + ``` + + Outside one, a markdown file in the project's incident location. Either way the capture from step + 1 goes beside it. +3. **Timeline** (2), then **onset** (3). The timeline is the spine every later step writes onto. +4. **Peer differential** (4) on every anomaly the timeline turns up, before it becomes a hypothesis. +5. **Hypothesis ledger** (7): 3–5 ranked candidate causes, each with the observation that would + refute it. Test with read-only probes. When a hypothesis leads to code that can be exercised, + hand that part to `aep:diagnosing`. +6. **Techniques 5, 6, 8 and 9** whenever a claim of their kind is about to be written: an absence, a + ship state, a green check, or a read that feeds a decision. +7. **Post-incident** (10) once the system is stable. + +Record each load-bearing observation against the incident as it is made: + +```console +$ aep plan artifact evidence incident-report: --kind health_observation \ + --source "" --ref "" --at +``` + +`metric_observation` for a metric series, `deployment_result` for what a deployment actually ran. + +## Boundaries + +- **Read-only by default.** A read of a live system (logs, metrics, process state, an API `GET`) + is yours to run. A restart, a rollback, a scale change, a configuration write, a post to a shared + channel or a ticket write is the operator's: propose it, with the evidence, and wait. +- A debugger attached to a live process pauses it. Say so, and prefer a non-pausing read + (`/proc//task/*/stack`, a runtime's own dump signal) when the process is still serving. +- Never put a secret or a personal identifier into a query that is logged, a file that is committed + or a message that is posted. + +## Reporting + +Verdict first: **resolved / ongoing / unknown**, and **root cause known / hypothesis / unknown**. +Then the timeline table, each row with its source. Then the ledger: verified claims, inferred claims, +refuted hypotheses. Then what evidence was lost, and the detection gap. Then the follow-ups, each one +an artifact: + +```console +$ aep plan artifact new story --title "" --relate informed_by:incident-report: +``` + +## Next + +- A hypothesis points at code that can be exercised: `aep:diagnosing`. +- A follow-up story is ready to build: `aep:implementing`. +- The incident needs a written postmortem: technique 10, then + `aep plan artifact new postmortem --title "" --relate derived_from:incident-report:<slug>`. diff --git a/plugins/aep/skills/investigating/references/techniques.md b/plugins/aep/skills/investigating/references/techniques.md new file mode 100644 index 0000000..883b0dc --- /dev/null +++ b/plugins/aep/skills/investigating/references/techniques.md @@ -0,0 +1,268 @@ +# Investigation techniques + +The procedures behind the catalogue in [`SKILL.md`](../SKILL.md). Each technique names the failure +it was written from, what to run, and when it is done. The commands are examples for common +platforms (Linux, Kubernetes, Prometheus-style metrics, Git). Use the equivalent instrument where the +system differs, and name which one you used. + +## 1. Capture before remediation + +**Written from:** a process stopped serving calls while its health probe stayed green. The report +reached the operator and the pod was restarted about four minutes later. It was gone within a +minute, and with it every thread and lock state. The cause of the hang cannot now be determined. + +**When:** the faulty process is still running, whether hung, leaking, spinning or misrouting. + +Write everything to one capture file (`capture-<UTC timestamp>.txt` beside the incident record), +**in this order**, cheapest and most perishable first. Give each step a timeout, so one hung read +does not use up the window. + +1. **Identity and placement:** process id, host or pod, node, address, image and digest, start + time. (`kubectl get pod <pod> -o wide`, and `-o jsonpath` for + `.status.containerStatuses[*].imageID`.) +2. **Per-thread state, without pausing the process:** + + ```console + $ for t in /proc/<pid>/task/*; do echo "$(cat $t/comm) $(cat $t/wchan)"; done | sort | uniq -c | sort -rn + $ cat /proc/<pid>/task/*/stack # needs privilege; skip if refused, and say so + $ grep -E 'State|Threads|VmRSS' /proc/<pid>/status + ``` + + Many threads parked on one `futex` or wait channel is the signature of a lock wait. That points + toward a deadlock and does not prove one. +3. **The runtime's own dump**, when it has one: a thread dump (`jstack <pid>`, `kill -3`), a + goroutine dump (`SIGQUIT`, or `/debug/pprof/goroutine?debug=2`), an application lock table + (a PBX's `core show locks`, a database's lock view). Check first whether the signal kills the + process. `SIGQUIT` ends a Go program. +4. **A full backtrace, only if the process is already lost for service:** + `gdb -p <pid> -batch -ex 'thread apply all bt'`. It pauses the process while it runs. If the + image carries no debugger, record that as a finding. +5. **Resource usage, against a healthy peer:** CPU, memory, open files, thread count + (`kubectl top pod --containers`, `ls /proc/<pid>/fd | wc -l`). +6. **Logs:** the last few thousand lines with timestamps, from every container in the unit + (`kubectl logs <pod> -c <container> --timestamps --since=<window>`), plus `--previous` when a + container has restarted. Logs in a central store survive the pod, so read them there later. The + ones only on the node do not. +7. **Orchestrator events and description:** `kubectl describe pod <pod>`, + `kubectl get events --sort-by=.lastTimestamp`. + +Then tell the operator the capture is done, with the file path, so remediation can go ahead. If +remediation happened first, write down what was not captured. That gap is part of the finding. + +**Done when:** the capture file exists and says which of the seven steps ran, which were refused, +and why. + +## 2. Sourced timeline + +**Written from:** an alert burst was placed two hours after the deploy that caused it, because one +source was read in local time and another in UTC. The real outage was nearly written off as +coincidence. + +One table, oldest first, **all times UTC and labelled so**: + +| time (UTC) | event | source | +|---|---|---| +| 07:09:08 | last per-request log line from the suspect | `kubectl logs … --timestamps` | + +- Convert every epoch explicitly (`date -u -d @<seconds>`). Chat and message-queue timestamps are + usually epochs, and a screenshot's wall clock is in somebody's local zone. Never copy a time out + of a person's or a bot's prose. Read it from the record. +- A row without a source is not a row. When a person reported something, the source is + "<person>, in <where>", and the content is their claim, not a fact. +- Include what *did not* happen when it matters, with the query that showed it, such as "probe + stayed Ready until 09:05". +- Give an announced time (a maintenance window, a scheduled deploy) its own row, labelled + *announced*. The actual event gets the time the system recorded, such as a merge time or a + process start time. These have differed by more than a day. + +**Done when:** every row has a source, and every time is UTC. + +## 3. Onset from state + +**Written from:** a condition was dated to the first day an alert about it was delivered. The +underlying counter showed it had been happening every 48 hours for more than three weeks before +that. Roughly nineteen occurrences had produced no alert. + +- **An alert, a notification or a digest cannot give you an onset.** It shows events that were + produced *and* delivered. Its silence says nothing about the condition. A rule may not have + existed yet, may not have been routed, or may have been rate-limited. +- Find the instrument that can see a negative. A counter (`…_restarts_total`), a state + (`…_last_terminated_reason`), a metric series, the repository for a merge, the cluster for a + deploy. +- **Query back past the apparent onset**, not up to it. If the condition is already present at the + start of your window, the onset is "before <window start>". Say that, and do not report the + window start as the onset. +- **Name the datasource before you accept a retention limit.** "Logs only go back 30 days" belongs + to one store. Another one in the same system may hold 49 days. +- If the stream and the state disagree, the difference is a finding in its own right: other alerts + of that class may be missing too. + +**Done when:** the onset is stated as "at or before <time>", with the instrument and the query window +named. + +## 4. Peer differential + +**Written from:** a hung node's log was full of one repeated warning, and it read as the cause. A +healthy peer in the same deployment logged 115 of the same warning in the same four hours. + +For every anomaly on the suspect, run the same read on at least one healthy peer (same version, +same role, same window) before it enters the ledger: + +| signal | suspect | healthy peer | counts as evidence? | +|---|---|---|---| +| warning X per 4 h | 120 | 115 | no | +| memory | 940 MiB | 176 MiB | yes | + +A signal that the peers share is background. If there is no healthy peer (a single instance, or all +instances affected), say so. The differential is then against the same instance's own earlier +window. + +**Done when:** every anomaly in the ledger has its peer reading beside it. + +## 5. Negative claims + +**Written from:** three separate failures. First, "no release names this fix" was written while two +tags carrying it had existed for seventeen minutes. The scanner that was read only reports what +moved since its last run. Second, "nobody has asked for this" was written about a request made +twice in the same chat thread, because only one reply of the thread was in the reader's slice. +Third, "no content" in two thousand captured messages turned out to be a capture that never read +the field the content lived in. + +Before writing that something does not exist, did not happen or was never asked: + +1. **Query the system that owns the answer, directly**, and say that you did: the tag list, the + merge request, the release page, the full thread. +2. **A delta is not an inventory.** Anything incremental (a scan past a watermark, a digest, a + "since last run" slice) can only say "not in this window". +3. **A reply needs its thread.** Read the whole thread, not the one message and its root. +4. **If a whole source looks empty or uniform, suspect the capture first.** Uniform emptiness from a + busy source usually means the reader is broken, not the source. +5. **A bot's or a person's "none exist" is a claim to check.** It is not a source. + +**Done when:** each absence in the report names the direct query that returned nothing. + +## 6. Ship state + +**Written from:** three failures. First, a declared version range was read as current behaviour +("a downgrade is queued for four environments"). When measured, the four ran four different +versions, and the new one had reached none of them, because the range resolves only when the +environment next syncs. Second, a ticket in "Backlog" was read as "no code exists" while three +merges and four tags carried the fix. Third, "deployed nowhere" was concluded by comparing tag +*names* while the fix was running. + +Four states, four systems. Never infer one from another: + +| state | system of record | the read | +|---|---|---| +| intended | the tracker | the ticket status. It records intent only | +| implemented | every repository that could carry it | merged changes that name it, in *all* candidate repositories | +| releasable | tags or releases | the tags that contain the merge (`git tag --contains <sha>`) | +| deployed | the running system | the digest the process actually runs, from the pod status and not the spec's tag | + +For "deployed", follow the provenance chain: running digest → the registry's tags on that digest → +keep only an immutable reference (a version tag or a commit SHA, never a moving name like +`latest`) → ancestry check against the fix commit (`git merge-base --is-ancestor <fix> <ref>`). +Use the registry's push time as a bound. An image pushed before the fix commit cannot contain it. + +Distinguish **deployed nowhere** (provenance resolved everywhere and absent) from **unknown** +(provenance did not resolve somewhere). Name which environment did not resolve. + +**Done when:** each ship claim names its state and the read that established it. + +## 7. Hypothesis ledger + +**Written from:** a review stated that a defect fired "on every run" by reading the branch that +reaches it. Runtime data showed at most once a day. The severity had been set from the code-read +number. + +Keep three lists, and put every claim in the report into exactly one of them: + +- **Verified:** observed, with the source (technique 2's rule). +- **Inferred:** follows from verified facts or from reading code, and has not been observed. It is + labelled *inferred* wherever it appears. +- **Refuted:** a hypothesis, the observation that refuted it, and its source. + +Rules: + +- **Code-path inference is a hypothesis.** It becomes a finding only with a runtime observation: + a log line, a counter or a trace of one instance end to end. +- **Before blaming a code path, find its callers.** "Who calls this, and does that caller run + here?" rules out many theories quickly. +- **Frequency and severity come from runtime data only.** +- **"This cannot happen" needs every form enumerated.** Loops, retries, watchers, timers and + recursion, each with its exit condition read. One keyword search does not show that something + cannot happen. +- **Quote a vague statement verbatim, and mark it "unclear".** Do not invent a mechanism that + makes it make sense. +- Write 3–5 ranked hypotheses, each with the observation that would refute it, before probing any. + +**Done when:** no claim in the report sits outside the three lists. + +## 8. Check the check + +**Written from:** two failures. First, a health probe reported Ready for two hours while the +process served no requests, so the platform never replaced it. Second, a drift gate exited 0 and +was cited as confirming a sentence it had never extracted, because it resolved references in +qualified form only and the sentence used the short form. + +Before citing a probe, a gate, a scanner or a monitor as evidence: + +1. **What does it read?** A readiness endpoint that answers from a sidecar says nothing about the + main process. A gate that iterates the entries it finds says nothing about the entries it did + not find. +2. **State its denominator.** "36 of 36 clean" when 56 are declared is a check over a subset. +3. **Is your claim in its scan set?** Look for your file, line or entity in the check's own + coverage output. Finding the same identifier somewhere else does not count. +4. **Has it ever gone red on this case?** Run it against a case whose answer you already know. + A check that has never failed on a known-bad case has measured nothing. +5. **Empty output is not "all clear"** until you know the check could not have crashed, been rate + limited or timed out silently. + +A probe that stayed green through the incident is a **detection gap**, and it goes into technique 10. + +**Done when:** every check cited as evidence has its scope stated beside it. + +## 9. Complete reads + +**Written from:** a file API returned the first 64 KiB of a 134 KB file with a `truncated` flag +nobody read. The content was edited and written back, deleting 1,793 lines, and reached a merge +request before a person caught it. + +A partial read that reports success looks exactly like a complete one, unless you check its length +or its hash: + +- Check the truncation or `has_more` flag on every response. +- Compare the byte count returned with the size the API reports. +- Before editing through an API, compare the hash of what you hold with the hash the API reports + (`git hash-object <file>` against the blob id). After writing, read back and compare again. +- For a paged API, know which end the first page comes from. With some APIs it depends on which + bound was sent. Page until the envelope says there is no more. +- A query capped at N rows (an SQL console, a log search) that returns exactly N is truncated until + shown otherwise. + +**Done when:** every read that fed a decision or a write has its completeness check recorded. + +## 10. Post-incident + +Once the system is stable, write the report from techniques 2 and 7. Do not write it from memory. +In an AEP store: + +```console +$ aep plan artifact new postmortem <slug> --title "<title>" --relate derived_from:incident-report:<slug> +``` + +Sections, each with sources: + +- **Impact:** who was affected, how many, for how long, and whose count it is (a reporter's + estimate is labelled as one). +- **Timeline:** the table from technique 2. +- **Cause:** from the verified list, or "unknown" plus the leading hypothesis marked *inferred*. +- **Detection gap:** why monitoring did not see it, or saw it late (technique 8), and the time from + onset to the first human report. +- **Evidence lost:** what technique 1 could not capture and why. +- **Follow-ups:** each one an artifact + (`aep plan artifact new story <slug> --title "<the gap>" --relate informed_by:incident-report:<slug>`), + not a sentence. The standard follow-ups are the detection gap, the capture gap (tooling missing + from the image, no dump signal) and the cause, if known. + +**Done when:** each follow-up exists as an artifact, and the report carries no unsourced specific. diff --git a/plugins/aep/skills/migrating/SKILL.md b/plugins/aep/skills/migrating/SKILL.md index 1aaadfd..8205570 100644 --- a/plugins/aep/skills/migrating/SKILL.md +++ b/plugins/aep/skills/migrating/SKILL.md @@ -3,7 +3,7 @@ name: migrating description: Migrate a repository's legacy work tracking — story trees, TODO.md, plan and issue documents — into the governed AEP planning store, without deleting or rewriting the sources. Use when the user asks to migrate, import, port or convert an existing backlog into AEP, when a repository is adopting AEP and already has work written down somewhere, or when a store has been adopted beside a legacy backlog nobody retired. Read it before creating the first artifact in a repository that already tracks work in markdown. --- -**Skill version 0.18.0** — the version in `.claude-plugin/plugin.json`. +**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`. # Migrating legacy tracking into the store diff --git a/plugins/aep/skills/planning/SKILL.md b/plugins/aep/skills/planning/SKILL.md index 9e6b716..2075b4b 100644 --- a/plugins/aep/skills/planning/SKILL.md +++ b/plugins/aep/skills/planning/SKILL.md @@ -3,7 +3,7 @@ name: planning description: Plan engineering work in a governed markdown artifact store — create, relate, move and validate epics, stories, tasks and initiatives through the `aep` CLI. Use when the user mentions planning, a backlog, an epic, a story, a task, decomposing or breaking down work, an artifact's status ("move this to active", "what is still in draft?", "why can't this be implemented?"), or when the project contains a `.engineering/planning/` directory. Use it at adoption too — the user asks to adopt AEP, to migrate from or replace the track plugin, to start a first backlog, or works in a repository with no `.engineering/` directory at all — because § 5 says how a first store is populated and it is worth nothing after one has been hand-written. Also use before editing any file under `.engineering/planning/`. --- -**Skill version 0.18.0** — the version in `.claude-plugin/plugin.json`. +**Skill version 0.19.0** — the version in `.claude-plugin/plugin.json`. # Planning in a governed artifact store diff --git a/plugins/b10x/.claude-plugin/plugin.json b/plugins/b10x/.claude-plugin/plugin.json index d5860c3..029b02a 100644 --- a/plugins/b10x/.claude-plugin/plugin.json +++ b/plugins/b10x/.claude-plugin/plugin.json @@ -2,7 +2,7 @@ "name": "b10x", "displayName": "Beyond10x", "description": "Set up, upgrade and check the Beyond10x plugins and binaries, route work to them, and create portable plugins.", - "version": "0.18.0", + "version": "0.19.0", "author": { "name": "Beyond10x" }, diff --git a/plugins/b10x/.codex-plugin/plugin.json b/plugins/b10x/.codex-plugin/plugin.json index 6051b9f..8813aca 100644 --- a/plugins/b10x/.codex-plugin/plugin.json +++ b/plugins/b10x/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "b10x", - "version": "0.18.0", + "version": "0.19.0", "description": "Set up, upgrade and check the Beyond10x plugins and binaries, route work to them, and create portable plugins.", "author": { "name": "Beyond10x" diff --git a/plugins/b10x/skills/routing/SKILL.md b/plugins/b10x/skills/routing/SKILL.md index b7e9ff4..ec733a8 100644 --- a/plugins/b10x/skills/routing/SKILL.md +++ b/plugins/b10x/skills/routing/SKILL.md @@ -32,6 +32,7 @@ Route the request; do not reproduce a specialist plugin's full workflow. | Move an existing backlog into the AEP store without losing its sources | `aep:migrating` | | Scope and deliver accepted development work through a reviewed wave | `aep:implementing` | | Diagnose a failing, flaky or slow behaviour through a red-capable loop | `aep:diagnosing` | +| Investigate a production incident, an outage, an onset or a ship state from cited evidence | `aep:investigating` | | Specify a system or API | `ess:specifying` | | Derive a specification for an existing system | `ess:retrofitting` | | Run or raise a conformance suite | `ess:testing-conformance` | diff --git a/plugins/connectors/.claude-plugin/plugin.json b/plugins/connectors/.claude-plugin/plugin.json index 8e69e9f..406eb2e 100644 --- a/plugins/connectors/.claude-plugin/plugin.json +++ b/plugins/connectors/.claude-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "connectors", - "version": "0.18.0", + "version": "0.19.0", "description": "Set up, inspect, and invoke governed integrations through the connectors CLI.", "author": { "name": "Beyond10x" }, "license": "Apache-2.0", diff --git a/plugins/connectors/.codex-plugin/plugin.json b/plugins/connectors/.codex-plugin/plugin.json index 54d03e8..101bee2 100644 --- a/plugins/connectors/.codex-plugin/plugin.json +++ b/plugins/connectors/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "connectors", - "version": "0.18.0", + "version": "0.19.0", "description": "Set up, inspect, and invoke governed integrations through the connectors CLI.", "author": { "name": "Beyond10x" }, "license": "Apache-2.0", diff --git a/plugins/ess/.claude-plugin/plugin.json b/plugins/ess/.claude-plugin/plugin.json index e29228c..c89e778 100644 --- a/plugins/ess/.claude-plugin/plugin.json +++ b/plugins/ess/.claude-plugin/plugin.json @@ -2,7 +2,7 @@ "name": "ess", "displayName": "ESS", "description": "Write, retrofit, validate and project Executable System Specifications, and hold implementations to them with conformance suites.", - "version": "0.18.0", + "version": "0.19.0", "author": { "name": "Beyond10x" }, diff --git a/plugins/ess/.codex-plugin/plugin.json b/plugins/ess/.codex-plugin/plugin.json index 4aef5ea..09ad533 100644 --- a/plugins/ess/.codex-plugin/plugin.json +++ b/plugins/ess/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "ess", - "version": "0.18.0", + "version": "0.19.0", "description": "Write, retrofit, validate and project Executable System Specifications, and hold implementations to them with conformance suites.", "author": { "name": "Beyond10x" diff --git a/plugins/worktree/.claude-plugin/plugin.json b/plugins/worktree/.claude-plugin/plugin.json index 8e080f5..426c724 100644 --- a/plugins/worktree/.claude-plugin/plugin.json +++ b/plugins/worktree/.claude-plugin/plugin.json @@ -2,7 +2,7 @@ "name": "worktree", "displayName": "Worktree", "description": "Create, lease, finish, audit and safely clean isolated Git worktrees through the worktree CLI.", - "version": "0.18.0", + "version": "0.19.0", "author": { "name": "Beyond10x" }, diff --git a/plugins/worktree/.codex-plugin/plugin.json b/plugins/worktree/.codex-plugin/plugin.json index 68883f3..75a7815 100644 --- a/plugins/worktree/.codex-plugin/plugin.json +++ b/plugins/worktree/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "worktree", - "version": "0.18.0", + "version": "0.19.0", "description": "Create, lease, finish, audit and safely clean isolated Git worktrees through the worktree CLI.", "author": { "name": "Beyond10x" diff --git a/website/docs/intro.md b/website/docs/intro.md index c06455b..9a3a52f 100644 --- a/website/docs/intro.md +++ b/website/docs/intro.md @@ -14,7 +14,7 @@ directly when the work is already clear. | Plugin | Use it for | Includes | |---|---|---| | [`b10x`](./plugins/b10x.md) | Setup, upgrades and navigation | `init`, `upgrade`, `routing` and `authoring-plugins` skills, the `b10x` binary, drift check | -| [`aep`](./plugins/aep.md) | Governed planning and delivery | `planning`, `migrating`, `implementing` and `diagnosing` skills; decomposer, plan critics, reverse engineer, story scoper, implementor, adversary, security reviewer | +| [`aep`](./plugins/aep.md) | Governed planning and delivery | `planning`, `migrating`, `implementing`, `diagnosing` and `investigating` skills; decomposer, plan critics, reverse engineer, story scoper, implementor, adversary, security reviewer | | [`ess`](./plugins/ess.md) | Executable System Specifications | `specifying`, `retrofitting`, `testing-conformance` and `hardening` skills; author, conformance, retrofitter agents | | [`worktree`](./plugins/worktree.md) | Git workspaces | managed worktrees, leases, recovery proof, and safe cleanup | | [`connectors`](./plugins/connectors.md) | Integrations | the `integrating` skill for the `connectors` CLI: set up providers and invoke governed integrations | diff --git a/website/docs/plugins/aep.md b/website/docs/plugins/aep.md index 8dcc6c4..c5e3c1f 100644 --- a/website/docs/plugins/aep.md +++ b/website/docs/plugins/aep.md @@ -77,6 +77,15 @@ that reproduces the real call pattern. The red and green runs are recorded as `t evidence against the owning story. When no such seam exists, the missing seam is filed as a draft story. +The `investigating` skill handles what cannot be re-run: a production incident, an outage, or a +question about a running system such as "when did this start" or "is the fix deployed". It +captures the process state before anyone restarts it, builds a UTC timeline where every row names +its source, dates an onset from an instrument that can see a negative, and checks each anomaly +against a healthy peer. Every claim is labelled verified or inferred. The investigation is an +`incident-report` artifact, its observations are `health_observation` evidence, and each +follow-up is a draft story. Ten techniques cover the work, from capture before remediation to the +postmortem. + Every role above, in both halves, is written once, as `references/<role>.md` of the skill that owns it. Claude Code runs it as a subagent through a thin `agents/<role>.md` adapter; Codex, which loads skills but not `agents/`, runs the same file directly. diff --git a/website/docs/structure.md b/website/docs/structure.md index 51247d3..4486ca7 100644 --- a/website/docs/structure.md +++ b/website/docs/structure.md @@ -24,7 +24,7 @@ tools` checks the skills against the newest CLI releases. A change that breaks a | plugin | lifecycle | activities | commands | agents | |---|---|---|---|---| | `b10x` | `init` (guided onboarding), `upgrade` | `routing`, `authoring-plugins` | — | — | -| `aep` | `init`, `upgrade` | `planning`, `migrating`, `implementing` (wave or drive mode), `diagnosing` | `wave`, `drive` (both hand off to `implementing`) · `review-plan`, `decompose` (both hand off to `planning`) | `planning`: decomposer, four plan critics, plan reviewer, reverse engineer · `implementing`: story scoper, implementor, adversary, security reviewer (each procedure in its skill's `references/`) | +| `aep` | `init`, `upgrade` | `planning`, `migrating`, `implementing` (wave or drive mode), `diagnosing`, `investigating` | `wave`, `drive` (both hand off to `implementing`) · `review-plan`, `decompose` (both hand off to `planning`) | `planning`: decomposer, four plan critics, plan reviewer, reverse engineer · `implementing`: story scoper, implementor, adversary, security reviewer (each procedure in its skill's `references/`) | | `ess` | `init`, `upgrade` | `specifying`, `retrofitting`, `testing-conformance`, `hardening` | — | `specifying`: author · `retrofitting`: retrofitter · `testing-conformance`: conformance | | `worktree` | `init`, `upgrade` | `managing-worktrees` | `cleanup` (hands off to `managing-worktrees`) | — | | `connectors` | `init`, `upgrade` | `integrating` | — | — |