Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions examples/prompt-lab/.claude/settings.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
{
"permissions": {
"allow": [
"mcp__relaycast__*"
]
}
}
2 changes: 2 additions & 0 deletions examples/prompt-lab/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
lab/

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Gitignore hides nested proof lab

Low Severity

The new lab/ ignore matches any directory named lab, including evidence/run/lab/. That conflicts with the more specific .lock.db rule, which only makes sense if the rest of the proof lab stays tracked. Regenerating the proof writes new hash-keyed work directories under that path, so those files can stay untracked and the committed evidence can go stale.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 90f56b6. Configure here.

evidence/run/lab/.lock.db
161 changes: 161 additions & 0 deletions examples/prompt-lab/PROOF.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,161 @@
# prompt-lab: design notes and proof

The Prompt Lab product brief as one relayflow. Prompt Lab is the workbench for
writing and fixing the prompts that draft home-health charts. The brief defines
two jobs and a shelf of fake patients. This flow covers all three:

| `job` | Brief section | What it does |
| --- | --- | --- |
| `new-agency` | Job 1 · Config-level / new agency | Sorts the agency's questions into sharing piles, writes first-pass prompts for agency-specific questions, runs them on shelf patients, and puts the output in an edit grid. Your first pass is saved as targets. Iteration runs on changed rows. Agency-specific prompts are committed (all / only / all except). Shared rows go to the question manager. |
| `fix` | Job 2 · Question-level refinement | Takes one question, from an issue or picked directly. Runs it on shelf patients, and your review of the grid is saved as gold. The iterator rewrites the prompt, Prompt QA checks it, and it re-runs and scores against gold (% worked, per patient). Marking it done makes it live. |
| `patient` | Set up test patient | Takes a gap brief from the manager queue. An agent writes a patient plan and you kick generate. Patient QA loops until it passes, then locks the patient onto the shelf. |

Each job file is its brief diagram, line by line:

- **System** boxes are either deterministic TypeScript over journaled reads, or model calls.
- **You** boxes are `f.human` gates, asked of `input.reviewer`.
- **Outcome** boxes are writes to the lab store.

## What maps to what

| Brief | Here |
| --- | --- |
| Apricot's Bank (`livePromptId`, global prompts) | `bank.json` in the lab directory, written only by [`store.ts`](store.ts) |
| The chart-filling engine nurses use | an `llm` step given the live prompt and the patient. Its answer must be on the agency's menu (the schema's `enum`) |
| Sharing piles: a deterministic lookup, never an agent | [`lib/piles.ts`](lib/piles.ts) |
| Computer-highlighted rows | [`lib/grid.ts`](lib/grid.ts) `highlights`: confidence below High, a mismatch pile, or a never-reviewed first-pass prompt |
| First pass persists as targets / gold | `store.ts record`, fed only from the grid you reviewed. AI output never writes gold directly |
| Iterator: rewrite-only | [`prompts.ts`](prompts.ts) `iterate`. Input is the changeset, patients, brief and existing prompt; output is new prompt text |
| Prompt QA: the brief plus shared guidelines | `promptQa`, looping with the iterator. Capped at 3 tries, then the run parks as `needs_human` |
| Done = live | `store.ts publish` sets `livePromptId`, as a compare-and-swap on the live prompt the run read: a retried publish is a no-op, and it never rolls back one published since. There is no promote step. A shared question warns which agencies it will change, and the re-run scores the new prompt on **every distinct menu** it is asked with ([`lib/piles.ts`](lib/piles.ts) `distinctMenus`), because done changes it for all of them |
| Shared frozen at config level | a changed shared or mismatch row becomes a `config-send` issue that carries its targets. It gets no rewrite at config level: Job 2 iterates from the live prompt with the full changeset |
| Test planner: never invents a patient | picks shelf patients of the run's visit type (checked deterministically); each hole becomes one gap brief per question × visit type |
| You do not approve the chart | the only patient gate is *kick generate*. Patient QA, plus a deterministic identifier check ([`lib/phi.ts`](lib/phi.ts)), locks it |

Every read and write of the lab is a journaled `f.run` step. Every store verb
is idempotent, so a retried step lands the lab in the same state. `write-new`
never overwrites an edit you made. Mutating verbs take an exclusive lab lock,
so two runs at once never lose each other's update. The lock is a SQLite
`BEGIN EXCLUSIVE` on `<lab>/.lock.db`. That's a kernel file lock, so the OS
releases it when its process dies, even from a SIGKILL.

A first-pass prompt is offered for commit only after it has run on a shelf
patient and you have reviewed its rows. A prompt whose question is a gap waits
as a draft until a patient covers it.

**Not built:** anything the brief lists under "Not at the start". Also not
built:

- Drafting a question brief with an agent. The flow reads a brief when one
exists in `briefs/<questionId>.md`.
- The Apricot patient-brief generator. It runs on live patients, so it belongs
in Apricot.
- The UI screens.

The flow is the job graph that sits under those screens.

## Run it

```sh
npm install
npm test # 23 unit tests over the deterministic parts and the store
node --experimental-strip-types store.ts ./my-lab seed fixtures
npx flows run prompt-lab.flow.ts --local-agent --input \
'{"job":"new-agency","reviewer":"<who answers the gates>","lab":"./my-lab","agency":"sunrise","visitType":"soc"}'
```

`reviewer` and `lab` are required, and there is no default person. The run
parks at each gate and prints the file to edit plus the `flows answer` /
`flows resume` commands. Edit the file, answer `yes`, resume. Answering `no`
stops the job as `declined` and keeps what was saved.

The fixtures are invented and reproduce the brief's own examples. The new
agency `sunrise` asks four questions:

- **wound-status**: shared with harbor and maple, with an identical menu.
- **mood**: a shared prompt, but sunrise adds "Agitated" to the menu, so it's a mismatch.
- **living-situation**: agency-specific, with no prompt yet.
- **ostomy-supplies**: agency-specific, with no shelf patient.

The shelf holds Pat, Jordan and Riley. The shared wound prompt contains a
deliberate flaw: it lets the referral overrule today's visit notes.

## Proof

[`prove.sh`](prove.sh) seeds a fresh lab and drives all three jobs through the
real kernel with real Claude calls (`--local-agent`). At each gate it acts as
the reviewer and applies the edit described in its comments, captured as a
diff. It captures every command with its output and exit code in
[`evidence/run/`](evidence/run/), and the final lab lands in
`evidence/run/lab/`.

```sh
./prove.sh evidence/run
```

The captured run (Claude Code 2.1.280, the adapter's default model). `prove.sh`
stops on the first exit it didn't expect, so reaching the final state means
every step below exited as shown:

| Run | Result | What happened |
| --- | --- | --- |
| Job 1 · `sunrise` / `soc` ([01](evidence/run/01-job1-run.txt), [04](evidence/run/04-resume.txt), [06](evidence/run/06-resume.txt)) | 28 steps, `success` | Piles came out as shared / mismatch / agency-specific. Both agency-specific questions got first-pass prompts that passed Prompt QA. The planner covered 3 questions from the `soc` shelf and queued `gap-ostomy-supplies-soc`. Gate 1: 9 rows, 7 highlighted. The reviewer raised Pat's wound confidence to High, rewrote the explanation and added a note ([02](evidence/run/02-reviewer-edit.diff)). That shared row went to the question manager with its target, and got no rewrite at config level. Gate 2 offered only `living-situation`, and it went live. `ostomy-supplies` had no shelf patient, so it was held as a draft. |
| Patient · `gap-ostomy-supplies-soc` ([07](evidence/run/07-patient-run.txt), [09](evidence/run/09-resume.txt)) | 14 steps, `success` | Plan, then kick generate. Patient QA sent the chart back before one passed (`llm-6` … `llm-12`). `roderick` locked onto the shelf, and the brief is marked `locked`. |
| Job 2 · the config-send issue ([10](evidence/run/10-job2-run.txt), [12](evidence/run/12-resume.txt), [14](evidence/run/14-resume.txt)) | 24 steps, `success` | Four shelf patients, including `roderick`. **The live prompt answered Pat "Healed / High", which is the brief's failure.** Gold came prefilled, with Pat's config target "Ongoing" carried over. The iterator's rewrite passed Prompt QA. The re-run scored **3 of 4 golded patients worked (75%)**: Pat is now "Ongoing" and worked; `roderick` did not (gold "Ongoing", new run "No wound"). The gate warned it would change harbor, maple and sunrise. Done made the new prompt live and closed the issue. |

Final state: [15-lab-state.txt](evidence/run/15-lab-state.txt), with the whole lab in `evidence/run/lab/`.

**Read the 75% with care.** `prove.sh` answers `yes` at every gate, so it
marked done at 75%. A reviewer would look at `roderick` first. His gold was the
old prompt's answer, accepted without review, and an ostomy patient's
peristomal skin damage may or may not be a "primary wound". That's exactly the
clinical call this gate exists for. The reviewer here is a script, not a
clinician.

`wound-status` has one menu across its agencies, so the per-menu re-run ran
with one menu. The multi-menu case, a mismatch question like `mood`, is covered
by the `distinctMenus` unit test only.

## Runtime findings (relayflows 2.0.29)

Four things in the runtime shaped this flow or its proof. Each workaround is commented
where it lives, and each has captured evidence in
[`evidence/runtime-findings/`](evidence/runtime-findings/):

1. **More than a few concurrent `f.llm` calls lose the run.** The queued
calls' 30 s leases expire before the worker takes them. The late completion
of a dead attempt is then refused ("Agent lease is already expired"), and
the CLI turns that refusal into a fatal `protocol_error`. Five parallel
calls passed and nine failed, with `--agent-capacity` 4 or 1.
[`runtime-parallel-llm-repro.flow.ts`](evidence/runtime-findings/runtime-parallel-llm-repro.flow.ts)
reproduces it with no Prompt Lab code. **Workaround:** model calls run
sequentially. Tracked in [#561](https://github.com/AgentWorkforce/flows/issues/561) and [#560](https://github.com/AgentWorkforce/flows/issues/560).
2. **A predicate `.gate(fn)` before an `f.human` can't be resumed.** The
verdict is read back from the `predicate-gates` stream with its keys
re-ordered. The lowered `<step>.gate` command no longer matches, and the
resume is refused as `run_admission_conflict`. The fix is one line in
`packages/sdk/src/authored-flow-executor.ts` `applyPredicateGate`: build the
literal from fixed fields, not from `JSON.stringify(record)`.
**Workaround:** the checks run in the body and fail through a journaled
failing step (`failStep`). **Fixed in [#558](https://github.com/AgentWorkforce/flows/pull/558).**
3. **`f.llm(prompt, { output })` fails when the reply is fenced JSON.** The
worker validates the raw reply. Sonnet sometimes wraps valid JSON in
```` ```json ```` anyway, and a failed run can't be resumed.
**Workaround:** text-form `f.llm`, then
[`lib/reply.ts`](lib/reply.ts) strips one fence and validates the schema,
with one bounded re-ask. The text form takes no `model`, so calls use the
Claude adapter's default model. **Fixed in [#558](https://github.com/AgentWorkforce/flows/pull/558).**

4. **A lease renewal that races a completion kills the run.** A step whose
child run journaled `success` was reported as
`lease_conflict: attempt has no active worker lease`. The CLI made that a
fatal `protocol_error`
([00-prove-attempt1-lease-conflict-after-success.txt](evidence/runtime-findings/00-prove-attempt1-lease-conflict-after-success.txt)).
It's intermittent: it happened twice in about 100 sequential calls ([second](evidence/runtime-findings/00-prove-attempt3-lease-conflict.txt)). There's
no workaround in the flow, so rerun. Tracked in [#560](https://github.com/AgentWorkforce/flows/issues/560).

Findings 1 and 4 are the same class of problem: late lease traffic becomes
fatal to the whole run instead of being ignored.

Local only for now: Cloud receives a single authored source, and this flow
imports sibling modules.
58 changes: 58 additions & 0 deletions examples/prompt-lab/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# Prompt Lab

Prompt Lab is your product brief as one runnable relayflow. It does the brief's
two jobs, plus making test patients. It stops and asks you wherever the brief
says **You**.

| Job | What it does | Where it asks you |
|---|---|---|
| `new-agency` | Sets up a new agency's prompts and runs them on fake patients | 1. Review the answers. 2. Approve what goes live. |
| `fix` | Fixes one prompt from an issue, then scores the fix | 1. Set the right answers. 2. Mark it done (it goes live). |
| `patient` | Makes a new fake patient for a coverage gap | 1. Start the build. |

## Setup (once)

You need Node 22.6+ and a signed-in [Claude Code](https://claude.com/claude-code).

```sh
npm install
node --experimental-strip-types store.ts ./lab seed fixtures
```

This creates `./lab`: a sample Bank with three agencies and three fake
patients.

## Run a job

```sh
npx flows run prompt-lab.flow.ts --local-agent --input \
'{"job":"new-agency","reviewer":"you","lab":"./lab","agency":"sunrise","visitType":"soc"}'
```

Other jobs use the same command with a different `--input`:

- Fix a question: `{"job":"fix","reviewer":"you","lab":"./lab","issueId":"<id from lab/queue/issues.json>"}`
- Make a patient: `{"job":"patient","reviewer":"you","lab":"./lab","briefId":"<id from lab/queue/patient-briefs.json>"}`

## When it stops for you

The run pauses and prints three things:

1. **The file to review** (for example `lab/work/…/grid.json`). Open it and
change any answer, confidence or explanation that's wrong.
2. **An answer command:** `npx flows answer … yes`. Use `no` to stop.
3. **A resume command:** `npx flows resume …`. Run it to continue.

Your first edits are saved as the right answers. The AI never overwrites them.

## Good to know

- **The fake Bank.** "Apricot" here is the local `lab` folder, and the chart
engine is a Claude call. Connecting it to the real Bank and engine is the
next step.
- **Shared prompts.** Marking a shared prompt done changes it for every agency
that uses it. The run warns you first.
- **Runs locally, not on Cloud yet.** Each job takes a few minutes, because AI
calls run one at a time for now.

How it was tested, with full evidence: [PROOF.md](PROOF.md).
80 changes: 80 additions & 0 deletions examples/prompt-lab/evidence/run/01-job1-run.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
$ npx flows run --no-observer-link --data-dir /private/tmp/claude-501/-Users-khaliqgant-Projects-AgentWorkforce-flows/df229951-f0e4-4753-990e-0f66533486d1/scratchpad/prove4-data --local-agent prompt-lab.flow.ts --input {"job":"new-agency","reviewer":"prompt-lab-reviewer","lab":"evidence/run/lab","agency":"sunrise","visitType":"soc"}
○ run-1 (deterministic) 0.00s
✓ run-1 (deterministic) 0.07s completionReason: success
○ llm-2 (llm) 0.00s
WAITING [worker_lease] Run "01M36DY3KREWBH8HTKEB6VK7KY" step "llm-2" (llm) is running under a worker lease until 1790143595437.
↻ llm-2 (llm) 4.06s
WAITING [worker_lease] Run "01M36DY3KREWBH8HTKEB6VK7KY" step "llm-2" (llm) is running under a worker lease until 1790143605451.
↻ llm-2 (llm) 14.10s
✓ llm-2 (llm) 21.23s completionReason: success
○ llm-3 (llm) 0.00s
WAITING [worker_lease] Run "01M36DYQW8KYMWYEHR5M0ZBKBN" step "llm-3" (llm) is running under a worker lease until 1790143616189.
↻ llm-3 (llm) 3.58s
✓ llm-3 (llm) 10.98s completionReason: success
○ run-4 (deterministic) 0.00s
✓ run-4 (deterministic) 0.08s completionReason: success
○ llm-5 (llm) 0.00s
WAITING [worker_lease] Run "01M36DZ2MX8GEADP9FEW024PQ3" step "llm-5" (llm) is running under a worker lease until 1790143627218.
↻ llm-5 (llm) 3.55s
WAITING [worker_lease] Run "01M36DZ2MX8GEADP9FEW024PQ3" step "llm-5" (llm) is running under a worker lease until 1790143637222.
↻ llm-5 (llm) 13.59s
✓ llm-5 (llm) 21.92s completionReason: success
○ llm-6 (llm) 0.00s
WAITING [worker_lease] Run "01M36DZQZJEEY12HHWFC2M79B4" step "llm-6" (llm) is running under a worker lease until 1790143649061.
↻ llm-6 (llm) 3.47s
✓ llm-6 (llm) 10.50s completionReason: success
○ run-7 (deterministic) 0.00s
✓ run-7 (deterministic) 0.10s completionReason: success
○ llm-8 (llm) 0.00s
WAITING [worker_lease] Run "01M36E04MZ7A2QC64XYRK1M0Q3" step "llm-8" (llm) is running under a worker lease until 1790143662042.
↻ llm-8 (llm) 5.85s
WAITING [worker_lease] Run "01M36E04MZ7A2QC64XYRK1M0Q3" step "llm-8" (llm) is running under a worker lease until 1790143672056.
↻ llm-8 (llm) 15.91s
✓ llm-8 (llm) 20.34s completionReason: success
○ run-9 (deterministic) 0.00s
✓ run-9 (deterministic) 0.10s completionReason: success
○ llm-10 (llm) 0.00s
WAITING [worker_lease] Run "01M36E0PJPCQ5QDWNA6XJBR104" step "llm-10" (llm) is running under a worker lease until 1790143680394.
↻ llm-10 (llm) 3.76s
✓ llm-10 (llm) 7.97s completionReason: success
○ llm-11 (llm) 0.00s
WAITING [worker_lease] Run "01M36E0Y6HJFHGJ8ZJASK5SC73" step "llm-11" (llm) is running under a worker lease until 1790143688196.
↻ llm-11 (llm) 3.60s
✓ llm-11 (llm) 7.40s completionReason: success
○ llm-12 (llm) 0.00s
WAITING [worker_lease] Run "01M36E15M53REC3MBPCZFRN8AM" step "llm-12" (llm) is running under a worker lease until 1790143695802.
↻ llm-12 (llm) 3.80s
✓ llm-12 (llm) 7.81s completionReason: success
○ llm-13 (llm) 0.00s
WAITING [worker_lease] Run "01M36E1CZPG4F2E1ZBC0R8FQC6" step "llm-13" (llm) is running under a worker lease until 1790143703337.
↻ llm-13 (llm) 3.53s
✓ llm-13 (llm) 7.22s completionReason: success
○ llm-14 (llm) 0.00s
WAITING [worker_lease] Run "01M36E1MBC1EH7CYPTB4J2E0JD" step "llm-14" (llm) is running under a worker lease until 1790143710880.
↻ llm-14 (llm) 3.85s
✓ llm-14 (llm) 7.86s completionReason: success
○ llm-15 (llm) 0.00s
WAITING [worker_lease] Run "01M36E1VWY2WGSBSAS2PJWT2Y6" step "llm-15" (llm) is running under a worker lease until 1790143718611.
↻ llm-15 (llm) 3.73s
✓ llm-15 (llm) 8.01s completionReason: success
○ llm-16 (llm) 0.00s
WAITING [worker_lease] Run "01M36E23SGV07GQER5VHWWW69S" step "llm-16" (llm) is running under a worker lease until 1790143726690.
↻ llm-16 (llm) 3.80s
✓ llm-16 (llm) 7.83s completionReason: success
○ llm-17 (llm) 0.00s
WAITING [worker_lease] Run "01M36E2B6R5V4HF07ZHZEWHFQE" step "llm-17" (llm) is running under a worker lease until 1790143734285.
↻ llm-17 (llm) 3.56s
✓ llm-17 (llm) 10.32s completionReason: success
○ llm-18 (llm) 0.00s
WAITING [worker_lease] Run "01M36E2NQFFXQKABB7S1RPM8X1" step "llm-18" (llm) is running under a worker lease until 1790143745059.
↻ llm-18 (llm) 4.01s
✓ llm-18 (llm) 8.22s completionReason: success
○ run-19 (deterministic) 0.00s
✓ run-19 (deterministic) 0.11s completionReason: success
○ human-20 (deterministic) 0.00s
⏸ human-20 (human) 0.00s
PARKED [run_parked] Run "01M36DXZJ9142JCT8RPDZA59W7" is waiting for prompt-lab-reviewer to answer human-20: "Config workbench · sunrise soc: 9 rows on 3 questions (7 highlighted, listed first); 1 gap brief(s) queued for the test patient manager.\nEdit target answer / confidence / explanation / notes in evidence/run/lab/work/sunrise-soc/35966ab5/grid.json. Your first pass persists as the target for each question × patient.\nyes = persist targets and run iteration on every changed row; no = stop without persisting."
Answer with: flows answer --data-dir /private/tmp/claude-501/-Users-khaliqgant-Projects-AgentWorkforce-flows/df229951-f0e4-4753-990e-0f66533486d1/scratchpad/prove4-data 01M36DXZJ9142JCT8RPDZA59W7 human-20 yes|no
Then continue with: flows resume --data-dir /private/tmp/claude-501/-Users-khaliqgant-Projects-AgentWorkforce-flows/df229951-f0e4-4753-990e-0f66533486d1/scratchpad/prove4-data --local-agent 01M36DXZJ9142JCT8RPDZA59W7
RUN 01M36DXZJ9142JCT8RPDZA59W7 parked
exit=3
Loading
Loading