Skip to content

feat(kernel,sdk,surface): f.human parks on wait.human; flows answer resumes it - #466

Merged
khaliqgant merged 4 commits into
mainfrom
feat/f-human
Sep 18, 2026
Merged

khaliqgant merged 4 commits into
mainfrom
feat/f-human

Conversation

@khaliqgant

@khaliqgant khaliqgant commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Closes #400.

f.human(question, { to }) used to throw unsupported_verb. It now parks the run durably and resumes with the answer.

What happens

const ok = await f.human(`Ship this?\n${summary}`, { to: 'khaliq' });
if (!ok) return f.done('declined');
$ flows run ask.flow.ts --input '{"plan":"v2"}'
○ run-1 (deterministic) 0.00s
✓ run-1 (deterministic) 0.01s completionReason: success
○ human-2 (deterministic) 0.00s
⏸ human-2 (human) 0.00s
PARKED [run_parked] Run "01M2T…" is waiting for khaliq to answer human-2: "Ship this?\nplan: v2"
Answer with: flows answer 01M2T… human-2 yes|no
Then continue with: flows resume 01M2T…
RUN 01M2T… parked            # exit 3

$ flows answer 01M2T… human-2 yes --note "looks good"
ANSWERED 01M2T… human-2 yes (looks good)
Continue with: flows resume 01M2T…

$ flows resume 01M2T…
✓ run-1 … ✓ human-2 … ✓ run-3 …
RUN 01M2T… completed (4 steps) completionReason: success

Lowering

  • Kernel — new verb step.wait {run_id, step_id, attempt, idempotency_key, wait_id, prompt, requested_of, options?, timeout_at_ms?}. Only the lease holder may call it (same check as step.complete). It appends wait.human for the running attempt — no step.completed, no semantic iteration charged — releases the lease, and drives: the step folds to NeedsHuman, the run parks. A reused wait_id is refused.
  • Kernel — event.emit closes wait.human when event_key equals the wait_id, with completionReason: human_responded and result = payload. That is the contract kernel/DESIGN.md §5 already documented for event.emit; it was not implemented. The step folds to Runnable and is dispatched as a fresh attempt.
  • Kernel — run_not_found now comes back from every verb for a run with no registry row and no journal file, instead of a journal_write_failed open error (a corrupt-but-present journal is still a journal failure).
  • SDK executor — f.human consumes an ordinal (human-N) like every authored operation. It reads the root journal for wait.completed{human_responded} under that id. None → throws AuthoredHumanParked; the durable root catches it, calls step.wait with the body's own wait id, and the signal propagates (exit 3, humanWait in the JSON report). Answer present → the boolean is lowered as a human-N deterministic step carrying {"human","to","answer","note?","answeredBy?"} on stdout, memoized under its admission key, so the value the author branches on is journal evidence the IPC verifier holds to the same standard as every other step. f.human is now Step<boolean> on the surface (a thenable like every other verb; was Promise<boolean>).
  • Bun standalone → Node body (the Cloud runtime path) — the park signal crosses the authenticated IPC frame (wait on the error frame, validated) and the Bun parent, which holds the lease, issues step.wait.
  • flows resume on a still-open question reports it (exit 3) instead of waiting 30s for a dispatch that cannot come.
  • flows answer <run> <wait> yes|no [--note] [--by <identity>] — checks the wait is open (refuses unknown / already answered / non-human-N ids as human_wait_unknown), then event.emits the answer contract { answer: boolean, note?, answeredBy?, at } (answeredBy = --by, else the OS user; Cloud's resumed sandbox passes the Cloud caller). Attaches no worker; prints the flows resume that continues.

Answer contract

event.emit(runId, waitId, { answer: boolean, note?: string, answeredBy?: string, at: ISO }). This is what flows answer sends and what Cloud's resumed sandbox sends via flows answer --by (companion: AgentWorkforce/cloud#3789 — answer route + resume applying the recorded answer). Anything else journaled as the result is refused on resume as human_answer_invalid.

Live vs documented

  • Live, proven here: local CLI cycle against a real daemon (tests/human-live.test.ts); Bun 1.4.0 standalone → Node body park/answer/resume (tests/authored-node-runtime.test.ts); kernel step.wait/event.emit cycle (server/tests.rs).
  • Documented, not in this PR: delivering the question to a channel (to is recorded, not routed — RFC covenant 3 / ops/BACKLOG.md); timeout (kernel never reads timeout_at_ms, DESIGN §1.4); answering from Cloud (AgentWorkforce/cloud#3789, which needs this release in the runtime pin first); f.dispatch still unsupported_verb.
  • Authority locally is the journal socket; answeredBy records who. Cloud enforces caller identity on its route.

Docs: docs/SURFACE.md §5 Human gates, kernel/DESIGN.md verb table, docs/EVENT-AWAIT.md stale facts, examples/social-post-pipeline now runs (and its done("canceled") → done("declined"), since canceled is a kernel outcome the body cannot declare).

Tests

  • kernel: 257 passed
  • sdk: 1938 passed (full suite with RELAYFLOWD_BIN + FLOWS_BUILD_BUN=bun 1.4.0), new: authored-human.test.ts (11), human-live.test.ts (3), cli-answer.test.ts (13), root + standalone cases.

🤖 Generated with Claude Code


Note

High Risk
Changes kernel wait/event semantics, lease lifecycle, and run open errors, plus a new human-answer trust boundary on the daemon socket—mistakes could strand runs or accept bad answers.

Overview
f.human is wired end-to-end instead of unsupported_verb: authored flows park on the kernel's wait.human, record answers in the journal, and resume with a memoized boolean step (human-N).

The kernel adds step.wait so the lease holder can park an attempt without completing it (journals wait.human, releases the lease, run status parked). event.emit now also matches open wait.human when event_key is the wait_id, requiring answeredBy, stamping at_ms from the server, and journaling attribution: client_asserted. Missing runs surface as run_not_found instead of a journal open error.

The SDK/CLI adds flows answer <run> <wait-id> yes|no (optional --note, --by), exit 3 reports with humanWait and answer/resume hints, and f.human returns Step<boolean> on the surface. The durable root turns AuthoredHumanParked into step.wait; after an answer, the body re-executes and lowers the decision as deterministic journal evidence. f.dispatch stays unsupported.

Docs and social-post-pipeline are updated to describe the live park/answer/resume loop; timeout and channel delivery for to remain future work.

Reviewed by Cursor Bugbot for commit 527c09f. Bugbot is set up for automated code reviews on this repo. Configure here.

Relayflow Lead and others added 2 commits September 18, 2026 05:29
…answer closes it

The authored body's f.human no longer throws unsupported_verb. It reads the
root journal for a wait.completed{human_responded} under its ordinal
(human-N); when none is recorded it throws AuthoredHumanParked, the durable
root turns that into the new step.wait verb, and the kernel journals
wait.human for the running attempt, releases the lease and parks the run
(needs_human, exit 3). event.emit keyed by the wait id closes a wait.human as
human_responded — the contract DESIGN.md already documented — and the root is
dispatched again; the resumed body finds the answer and lowers it as a
memoized human-N deterministic step so the boolean the author branches on is
journaled evidence. flows answer <run> <wait> yes|no [--note] records it.

Unknown runs now surface run_not_found from every verb instead of a journal
open failure.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@coderabbitai

coderabbitai Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 59 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: d729c609-8b94-4003-8bb7-2b199111236f

📥 Commits

Reviewing files that changed from the base of the PR and between a251bd6 and 527c09f.

📒 Files selected for processing (34)
  • docs/EVENT-AWAIT.md
  • docs/SURFACE.md
  • examples/social-post-pipeline/README.md
  • examples/social-post-pipeline/social-post-pipeline.flow.ts
  • kernel/DESIGN.md
  • kernel/relayflowd/src/engine.rs
  • kernel/relayflowd/src/engine/remote.rs
  • kernel/relayflowd/src/lib.rs
  • kernel/relayflowd/src/server.rs
  • kernel/relayflowd/src/server/protocol.rs
  • kernel/relayflowd/src/server/tests.rs
  • kernel/relayflowd/src/server/wire.rs
  • packages/sdk/src/authored-flow-error.ts
  • packages/sdk/src/authored-flow-executor.ts
  • packages/sdk/src/authored-human.ts
  • packages/sdk/src/authored-node-entry.ts
  • packages/sdk/src/authored-node-runner.ts
  • packages/sdk/src/authored-root.ts
  • packages/sdk/src/cli.ts
  • packages/sdk/src/cli/answer.ts
  • packages/sdk/src/cli/direct-run.ts
  • packages/sdk/src/cli/run.ts
  • packages/sdk/src/failure-kinds.ts
  • packages/sdk/src/journal-client.ts
  • packages/sdk/src/progress.ts
  • packages/sdk/src/protocol.ts
  • packages/sdk/tests/authored-flow.test.ts
  • packages/sdk/tests/authored-human.test.ts
  • packages/sdk/tests/authored-node-runtime.test.ts
  • packages/sdk/tests/authored-root.test.ts
  • packages/sdk/tests/cli-answer.test.ts
  • packages/sdk/tests/human-live.test.ts
  • packages/sdk/tests/journal-client-loopback.ts
  • packages/surface/src/context.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 4 potential issues.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

...(recorded.note === undefined ? {} : { note: recorded.note }),
...(recorded.answeredBy === undefined ? {} : { answeredBy: recorded.answeredBy }) };
const literal = `'${JSON.stringify(record).replaceAll("'", "'\\''")}'`;
await lowerDeterministic(id, `printf '%s' ${literal}`, false);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Named gates silently pass human answers

When f.human(...).gate(config) resolves, lowerDeterministic omits the operation's namedGate. The child skips the declared check, so rejected answers can continue the flow.

Learn more

Every surface Step<T> accepts a named gate, and AuthoredFlowOperation stores that gate in namedGate. The run, llm, and agent implementations pass this value into their lowering paths. The human implementation constructs its operation inline, so this call cannot access the stored value and starts the deterministic answer child without verification. The kernel then journals a successful child regardless of the declared gate.

Example: await f.human('Ship?', {to: 'owner'}).gate({type: 'regex_match', pattern: '"answer":true'}) receives false. The answer child still succeeds and the await returns false; the declared gate never produces gate_failed.

Recommended fix: Bind the human AuthoredFlowOperation to a variable, as the run, llm, and agent implementations do, and pass humanOp.namedGate to lowerDeterministic. Add tests for both passing and failing named gates on f.human.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Real bug — fixed in 527c09f. human() now hoists its operation (humanOp, the same pattern as runOp/llmOp) and passes humanOp.namedGate into lowerDeterministic, so .gate(config) on f.human lowers into the human-N step's verification and becomes the human-N.gate child exactly like every other authored step.

Test: tests/authored-human.test.ts "lowers a named gate on the answer into the human-N step, so a rejected answer can fail it" — asserts the lowered spec carries human-1.gate with the embedded pattern, and that a rejected answer under that gate fails the step (step_failed). Before the fix the gate never reached the spec.

Comment on lines +447 to +448
const recorded = await readHumanAnswer(journal, rootRunId, id);
if (recorded === undefined) throw new AuthoredHumanParked({ waitId: id, question, to }, rootRunId);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Human questions never reach recipients

When f.human lacks an answer, it only raises AuthoredHumanParked; no path delivers the question to to. Unattended runs remain parked until someone independently discovers and answers them.

Learn more

The repository constitution requires a declared human ask to be delivered to the recipient's channels while the run parks. This implementation reads the durable answer and emits a park signal, while driveRoot only journals wait.human. The newly added surface documentation also states that to is merely recorded and the local kit does not deliver it. A user who is not watching the originating command receives no indication that their decision is required.

Example: A scheduled flow asks f.human('Publish?', {to: 'khaliq'}) with no terminal attached. The run durably parks, but Khaliq receives no Slack, WhatsApp, Telegram, or iMessage request. The flow cannot continue until someone separately inspects the run and executes flows answer.

Recommended fix: Add a durable delivery path for newly journaled human waits. Route the prompt, evidence, run ID, and one-tap answer action to the declared recipient, with retries and deduplication keyed by the wait ID. Do not expose f.human as complete until at least one supported channel fulfills the RFC delivery contract.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and deliberate for this PR: to is recorded (wait.human.requested_of), not routed. This is stated in docs/SURFACE.md §5 Human gates: "to names who is asked and is recorded with the question; the local kit does not deliver it anywhere (that is the channel-delivery work in RFC covenant 3 and ops/BACKLOG.md)". The delivery surface today is the park report + flows answer locally, and on Cloud the run record + POST /api/v1/workflows/runs/<id>/answer (AgentWorkforce/cloud#3789), which is what a channel callback would hit.

Filed #468 for actual delivery (resolve to against declared tools, journaled send under the root so a resume does not re-send, one-tap answer → the Cloud route, and timeout), referencing RFC covenant 3. Leaving this PR to the durable wait/answer mechanics.

Comment on lines +398 to +406
} else if entry.entry_type == EntryType::WaitHuman {
let wait: WaitHumanPayload = serde_json::from_value(entry.payload.clone())?;
if wait.wait_id == event_key {
open.push((
wait.wait_id,
entry.step_id.clone(),
entry.attempt,
WaitCompletionReason::HumanResponded,
));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟥 Human gates lack recipient authorization

Any daemon client can answer any discovered wait through event.emit; the kernel never checks requested_of. Unauthorized approvals can resume sensitive flows as valid human decisions.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The daemon socket is the trust boundary for every verb, not just this one: a client that can event.emit can also run.start, run.cancel, step.complete on any run, or journal.read everything. The kernel does not authenticate clients; whatever reached the socket is trusted, which is the existing model for the local kit (same-user filesystem socket). Checking requested_of inside the kernel would be a check against a string the same untrusted client supplied, so it would add no authority.

On Cloud the authority check lives where identity exists: the answer route (AgentWorkforce/cloud#3789) requires session / cli:auth / workflow:invoke:write plus canAccessWorkflowRun, and refuses run-bound and sandbox tokens outright (human_answer_forbidden) so a run's own credentials — i.e. an agent inside it — cannot satisfy the gate. The sandbox then relays the authenticated identity via flows answer --by.

What I did tighten in the kernel (527c09f): event.emit now refuses a human response that lacks a non-empty answeredBy, so an unattributed answer can never close a wait.human; and the journal records the attribution as what it is (see the next thread). Kernel test step_wait_parks_the_attempt_and_a_human_answer_redispatches_it covers the refusal.

Comment thread packages/sdk/src/authored-human.ts Outdated
Comment on lines +39 to +42
if (typeof record !== 'object' || record === null || typeof record.answer !== 'boolean'
|| (record.note !== undefined && typeof record.note !== 'string')
|| (record.answeredBy !== undefined && typeof record.answeredBy !== 'string')
|| (record.at !== undefined && typeof record.at !== 'string')) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟨 Approval audit fields are caller-controlled

Non-CLI clients can set arbitrary answeredBy and at strings in event.emit. The journal then presents forged approval attribution as recorded evidence.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed both were client-asserted with nothing saying so. Smallest honest change (527c09f), on the kernel side where the journal is written:

  • at: the kernel now drops any client-supplied at and stamps at_ms in the wait.completed.result from its own clock — the same value as the entry's at_ms. The journal says when; the client never does. The SDK reads at_ms back and renders ISO in the lowered human-N record.
  • answeredBy: kept under that name (it is what flows answer --by / Cloud's route send and what humans read), but the kernel adds attribution: "client_asserted" to the result so the journal states that the identity was asserted by whoever held the socket, not verified by the kernel. On Cloud that holder is the resumed sandbox relaying an authenticated caller; locally it is the OS user. An empty/missing answeredBy is now refused.

I chose this over renaming to claimedBy because it keeps one field name across CLI, Cloud and journal, while making the epistemic status explicit in the record itself. Contract updated in authored-human.ts doc comment, SURFACE.md §5 and kernel/DESIGN.md event.emit row; kernel + SDK tests assert at absent, at_ms == entry clock, attribution present.

Relayflow Lead and others added 2 commits September 18, 2026 05:51
Cloud's answer route runs flows answer inside the resumed sandbox; the OS
user there is the sandbox, not the person who decided. --by records who.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nd timestamps answers

- A `.gate(config)` on f.human is now lowered into the human-N step's
  verification like every other authored step (it was dropped, so a declared
  gate was never evaluated).
- event.emit refuses a human response without a non-empty answeredBy,
  journals attribution: client_asserted (the socket authenticated the
  caller, not the kernel), drops any client at and stamps at_ms from the
  entry's own clock. The SDK contract follows: clients send
  { answer, note?, answeredBy }; the body reads at_ms/attribution back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Review swarm: maintainability

No fresh transcript was produced for run fca26a2a-49c7-4d13-ae57-0e00db3b62a6 (MISSING).

@github-actions

Copy link
Copy Markdown

Review swarm: history

No fresh transcript was produced for run fca26a2a-49c7-4d13-ae57-0e00db3b62a6 (MISSING).

@github-actions

Copy link
Copy Markdown

Review swarm: structure

No fresh transcript was produced for run fca26a2a-49c7-4d13-ae57-0e00db3b62a6 (MISSING).

@github-actions

Copy link
Copy Markdown

Review swarm: FAILED

  • maintainability: MISSING
  • history: MISSING
  • structure: MISSING

Cloud run: fca26a2a-49c7-4d13-ae57-0e00db3b62a6

@khaliqgant
khaliqgant merged commit 29edf0a into main Sep 18, 2026
9 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support durable human approval gates in authored TypeScript flows

1 participant