From 2193c01688b4a33c4a113f0d6210014596699c36 Mon Sep 17 00:00:00 2001 From: Relayflow Lead Date: Tue, 22 Sep 2026 22:07:34 -0700 Subject: [PATCH 1/3] fix(sdk): judge fenced llm replies by their content; resume past a predicate gate Two defects that stop an authored flow on relayflows 2.0.29. A structured f.llm reply wrapped in one markdown fence (```json ... ```) failed verification_failed, though the value inside matched the schema. Models add the fence despite the instruction not to, and a failed run cannot be resumed. The llm worker now parses the value inside exactly one surrounding fence; the schema still judges it, and prose or two fences still fail. A resumed body could not get past a predicate `.gate(fn)`: the recorded verdict comes back from the predicate-gates stream with sorted keys, so JSON.stringify of it lowered a different `.gate` command than the first run, and the admission key refused it (run_admission_conflict). The gate command is now built in a fixed field order. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/SURFACE.md | 4 +- .../mutation-1-fence-restored.txt | 5 ++ .../mutation-1-fence-reverted.txt | 9 ++++ .../mutation-2-predicate-restored.txt | 4 ++ .../mutation-2-predicate-reverted.txt | 10 ++++ .../suite-vs-main.txt | 37 +++++++++++++ packages/sdk/src/authored-flow-executor.ts | 11 +++- packages/sdk/src/llm-worker.ts | 13 ++++- .../tests/authored-agent-artifacts.test.ts | 53 +++++++++++++++++++ packages/sdk/tests/worker-transcript.test.ts | 45 +++++++++++++++- 10 files changed, 187 insertions(+), 4 deletions(-) create mode 100644 evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt create mode 100644 evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt create mode 100644 evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt create mode 100644 evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt create mode 100644 evidence/llm-fence-predicate-resume/suite-vs-main.txt diff --git a/docs/SURFACE.md b/docs/SURFACE.md index 4fc9f3565..557bfd82d 100644 --- a/docs/SURFACE.md +++ b/docs/SURFACE.md @@ -399,7 +399,9 @@ authentication probes, and exact `flows.json` model allow-list as agent steps; a declared `model` must be in that project's `models` array. A template call such as ``await f.llm`Summarize ${text}` `` returns text. The structured overload parses JSON and checks `output` before submitting a successful completion; -the kernel independently checks the schema before accepting the output. +the kernel independently checks the schema before accepting the output. A +reply that is exactly one markdown code fence around a value is judged by the +value inside it; prose around the JSON, or two fenced values, is still invalid. Invalid JSON or a schema mismatch completes with `verification_failed` and prevents downstream work. Retry and lease handling use the existing kernel policies; this overload introduces no separate retry contract. diff --git a/evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt b/evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt new file mode 100644 index 000000000..dbfcce816 --- /dev/null +++ b/evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt @@ -0,0 +1,5 @@ +# restored byte-for-byte (cmp exit 0) +$ npx vitest run tests/worker-transcript.test.ts -t "fenced reply" + ✓ tests/worker-transcript.test.ts (8 tests | 7 skipped) 373ms + ✓ the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value 373ms + Tests 1 passed | 7 skipped (8) diff --git a/evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt b/evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt new file mode 100644 index 000000000..dbca272d6 --- /dev/null +++ b/evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt @@ -0,0 +1,9 @@ +# mutation: llm-worker.ts parses result.stdout_tail without unfenced() +$ npx vitest run tests/worker-transcript.test.ts -t "fenced reply" + × the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value 403ms +⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ + FAIL tests/worker-transcript.test.ts > the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value +AssertionError: expected 'verification_failed' to be 'success' // Object.is equality +Expected: "success" +Received: "verification_failed" + Tests 1 failed | 7 skipped (8) diff --git a/evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt b/evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt new file mode 100644 index 000000000..8c594624e --- /dev/null +++ b/evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt @@ -0,0 +1,4 @@ +# restored byte-for-byte (cmp exit 0) +$ npx vitest run tests/authored-agent-artifacts.test.ts -t "sorted keys" + ✓ tests/authored-agent-artifacts.test.ts (5 tests | 4 skipped) 28ms + Tests 1 passed | 4 skipped (5) diff --git a/evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt b/evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt new file mode 100644 index 000000000..c51156333 --- /dev/null +++ b/evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt @@ -0,0 +1,10 @@ +# mutation: applyPredicateGate lowers JSON.stringify(record) instead of the canonical field order +$ npx vitest run tests/authored-agent-artifacts.test.ts -t "sorted keys" + × predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run 30ms + → expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality +⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ + FAIL tests/authored-agent-artifacts.test.ts > predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run +AssertionError: expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality +Expected: "printf '%s' '{"gate":"predicate","step":"run-1","verdict":"pass","because":"plan covers every question"}'" +Received: "printf '%s' '{"because":"plan covers every question","gate":"predicate","step":"run-1","verdict":"pass"}'" + Tests 1 failed | 4 skipped (5) diff --git a/evidence/llm-fence-predicate-resume/suite-vs-main.txt b/evidence/llm-fence-predicate-resume/suite-vs-main.txt new file mode 100644 index 000000000..f7f2bcf16 --- /dev/null +++ b/evidence/llm-fence-predicate-resume/suite-vs-main.txt @@ -0,0 +1,37 @@ +# Full SDK suite on this branch (packages/sdk: npm test), summary lines + ❯ tests/hosted-extension-isolation.test.ts (22 tests | 12 failed | 3 skipped) + ❯ tests/babysitter-native-extension.test.ts (41 tests | 41 skipped) + ❯ tests/hosted-base-snapshot.test.ts (18 tests | 15 failed | 1 skipped) + ❯ tests/authored-node-runtime.test.ts (14 tests | 14 skipped) + ❯ tests/mcp.test.ts (30 tests | 4 skipped) + ❯ tests/bundle.test.ts (26 tests | 1 failed) + ❯ tests/stop-process-group.test.ts (9 tests | 2 failed) + ❯ tests/hosted-extension-protocol.test.ts (24 tests | 7 failed | 1 skipped) + ❯ tests/worker-cli.test.ts (18 tests | 3 failed) + ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) + ❯ tests/cli-watch.test.ts (10 tests | 1 failed) + ❯ tests/canonical-software-factory.test.ts (3 tests | 3 failed) + ❯ tests/communication-mixed-resume.test.ts (1 test | 1 failed) + Test Files 13 failed | 166 passed | 3 skipped (182) + Tests 67 failed | 2753 passed | 72 skipped (2892) + +# The same 13 failing files on unmodified main 78cc5556 (git stash), summary lines +$ npx vitest run + ❯ tests/hosted-base-snapshot.test.ts (18 tests | 15 failed | 1 skipped) + ❯ tests/hosted-extension-isolation.test.ts (22 tests | 12 failed | 3 skipped) + ❯ tests/babysitter-native-extension.test.ts (41 tests | 41 skipped) + ❯ tests/authored-node-runtime.test.ts (14 tests | 14 skipped) + ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) + ❯ tests/canonical-software-factory.test.ts (3 tests | 3 failed) + ❯ tests/communication-mixed-resume.test.ts (1 test | 1 failed) + ❯ tests/mcp.test.ts (30 tests | 4 skipped) + ❯ tests/hosted-extension-protocol.test.ts (24 tests | 7 failed | 1 skipped) + Tests 60 failed | 114 passed | 64 skipped (238) + +# The four files that failed only in the full run, rerun on this branch in isolation (cwd packages/sdk) +$ npx vitest run tests/bundle.test.ts tests/stop-process-group.test.ts tests/worker-cli.test.ts tests/cli-watch.test.ts + ✓ tests/bundle.test.ts (26 tests) + ✓ tests/cli-watch.test.ts (10 tests) + ✓ tests/stop-process-group.test.ts (9 tests) + ✓ tests/worker-cli.test.ts (18 tests) + Tests 63 passed (63) diff --git a/packages/sdk/src/authored-flow-executor.ts b/packages/sdk/src/authored-flow-executor.ts index ebf9c28a1..afd32dd24 100644 --- a/packages/sdk/src/authored-flow-executor.ts +++ b/packages/sdk/src/authored-flow-executor.ts @@ -325,7 +325,16 @@ export async function executeAuthoredFlow( (await recordedVerdicts)?.set(id, record); } } - const literal = `'${JSON.stringify(record).replaceAll("'", "'\\''")}'`; + // Fixed field order, never the record as it came back: the journal stores + // stream messages with sorted keys, so JSON.stringify of a resumed record + // would lower a different command than the first run did, and the gate + // run's admission key would refuse it as bound to a different spec. + const canonical: PredicateRecord = { + gate: 'predicate', step: record.step, verdict: record.verdict, + ...(record.because === undefined ? {} : { because: record.because }), + ...(record.threw === undefined ? {} : { threw: record.threw }), + }; + const literal = `'${JSON.stringify(canonical).replaceAll("'", "'\\''")}'`; const command = record.verdict === 'pass' ? `printf '%s' ${literal}` : `printf '%s' ${literal} >&2; exit 1`; try { await observeStep(`${id}.gate`, 'deterministic', () => lowerDeterministic(`${id}.gate`, command, false), options.onProgress); diff --git a/packages/sdk/src/llm-worker.ts b/packages/sdk/src/llm-worker.ts index 35895079c..bb30d4029 100644 --- a/packages/sdk/src/llm-worker.ts +++ b/packages/sdk/src/llm-worker.ts @@ -67,7 +67,7 @@ export class LlmWorker extends EventEmitter { let detail = result.stderr_tail; if (reason === 'success' && schema !== undefined) { try { - output = JSON.parse(result.stdout_tail); + output = JSON.parse(unfenced(result.stdout_tail)); const invalid = jsonSchemaOutputError(schema, output); if (invalid !== undefined) { reason = 'verification_failed'; @@ -98,3 +98,14 @@ export class LlmWorker extends EventEmitter { }); } } + +/** + * A reply that is exactly one markdown code fence around a value, reduced to + * that value. Models add the fence despite the instruction not to; the schema + * still judges what is inside it. Anything else (prose around the JSON, two + * fences) is returned unchanged and fails the parse as before. + */ +export function unfenced(text: string): string { + const fence = /^\s*```[A-Za-z0-9_-]*[ \t]*\r?\n([\s\S]*?)\r?\n?```\s*$/.exec(text); + return fence === null ? text : fence[1]!; +} diff --git a/packages/sdk/tests/authored-agent-artifacts.test.ts b/packages/sdk/tests/authored-agent-artifacts.test.ts index d23297124..6a5578a38 100644 --- a/packages/sdk/tests/authored-agent-artifacts.test.ts +++ b/packages/sdk/tests/authored-agent-artifacts.test.ts @@ -247,4 +247,57 @@ describe('predicate verdicts recorded on the root run', () => { expect(gateSpecs).toHaveLength(2); for (const spec of gateSpecs) expect(spec).toContain('"verdict":"pass"'); }); + it('a resumed verdict read back with sorted keys lowers the same gate command as the first run', async () => { + path = sockPath(); + // The kernel returns stream messages with their keys sorted; the first run + // appended them in authoring order. Both runs must lower one command, or + // the gate run's admission key refuses the resume. + const stream: Record[] = []; + const commands: string[] = []; + let nextRun = 1; + const specs = new Map(); + server = startLoopback(path, { + hello: ctx => sendOk(ctx), + 'stream.read': (ctx, params) => { + const from = params.from_offset as number; + const page = params.stream === 'predicate-gates' ? stream.slice(from) : []; + const sorted = page.map(m => Object.fromEntries(Object.entries(m).sort(([a], [b]) => a.localeCompare(b)))); + sendResult(ctx, { messages: sorted.map(message => ({ message })), next_offset: from + page.length }); + }, + 'stream.append': (ctx, params) => { + if (params.stream === 'predicate-gates') stream.push(params.message as Record); + sendResult(ctx, { offset: stream.length }); + }, + 'run.start': (ctx, params) => { + const step = (params.spec as { steps: Array<{ id: string; command?: string }> }).steps[0]!; + if (step.id.endsWith('.gate')) commands.push(step.command!); + const runId = `sorted-run-${nextRun++}`; + specs.set(runId, step.id); + sendResult(ctx, { run_id: runId, status: 'completed', completion_reason: 'success', completed_steps: 1 }); + }, + 'journal.read': (ctx, params) => { + sendResult(ctx, { entries: [{ entry_type: 'step.completed', step_id: specs.get(params.run_id as string), payload: { + completionReason: 'success', disposition: 'step_done', output: { exit_code: 0, stdout_tail: 'x', stderr_tail: '' } } }] }); + }, + }); + const client = new JournalClient(path, { requestTimeoutMs: 2000 }); + await client.connect(); + await client.hello('predicate-sorted-test'); + try { + const handle = flow('predicate-sorted', async (f) => { + await f.run('echo a').gate(() => true, 'plan covers every question'); + f.done('success'); + }); + for (let run = 0; run < 2; run++) { + const result = await executeAuthoredFlow(handle, client, undefined, { rootRunId: 'root-sorted' }); + expect(result.completionReason).toBe('success'); + } + } finally { + client.close(); + } + expect(stream).toHaveLength(1); + expect(commands).toHaveLength(2); + expect(commands[1]).toBe(commands[0]); + expect(commands[0]).toContain('{"gate":"predicate","step":"run-1","verdict":"pass","because":"plan covers every question"}'); + }); }); diff --git a/packages/sdk/tests/worker-transcript.test.ts b/packages/sdk/tests/worker-transcript.test.ts index f745c69cc..06d931321 100644 --- a/packages/sdk/tests/worker-transcript.test.ts +++ b/packages/sdk/tests/worker-transcript.test.ts @@ -6,7 +6,7 @@ import { fileURLToPath } from 'node:url'; import { afterEach, describe, expect, it } from 'vitest'; import { LLM_ERROR_MAX_BYTES, TRANSCRIPT_DIGEST_MAX_BYTES, type TranscriptDigest } from '../src/agent-transcript.js'; import type { JournalClient } from '../src/journal-client.js'; -import { LlmWorker } from '../src/llm-worker.js'; +import { LlmWorker, unfenced } from '../src/llm-worker.js'; import type { Pins } from '../src/protocol.js'; import { AgentWorker } from '../src/worker.js'; import { runAgentCli } from '../src/worker-cli.js'; @@ -186,3 +186,46 @@ describe('the llm worker bounds and redacts its error and journals the digest', expect(Buffer.byteLength(JSON.stringify(payload.trajectory_tail), 'utf8')).toBeLessThan(16 * 1024); }); }); + +describe('the llm worker judges the value inside one markdown fence', () => { + async function completeWith(result: string): Promise { + const root = makeDirectory(); + const lines = fixtureLines.map(line => { + const frame = JSON.parse(line) as Record; + return frame.type === 'result' ? JSON.stringify({ ...frame, result }) : line; + }); + const claude = fakeClaude(root, lines); + const { client, completions } = stubClient(); + const worker = new LlmWorker(client, 'llm'); + const errors: unknown[] = []; + worker.on('error', error => errors.push(error)); + await worker.attach(); + (client as unknown as EventEmitter).emit('step.dispatch', { + run_id: 'run-f', step_id: 'llm-1', attempt: 1, step_type: 'llm', + spec: { cli: claude, prompt: 'answer', verification: { json_schema: { type: 'object', required: ['x'], properties: { x: { type: 'number' } } } } }, + pins, lease_id: 'lease', lease_deadline_ms: Date.now() + 30_000, idempotency_key: 'k', + }); + await worker.close(); + expect(errors).toEqual([]); + return completions[0]!; + } + + it('accepts a fenced reply whose content matches the schema, and journals the parsed value', async () => { + const completion = await completeWith('```json\n{"x": 4}\n```'); + expect(completion[4]).toBe('success'); + expect((completion[5] as { output: unknown }).output).toEqual({ x: 4 }); + }); + + it('still refuses prose around the JSON, and a fenced value that breaks the schema', async () => { + expect((await completeWith('Here it is: {"x": 4}'))[4]).toBe('verification_failed'); + expect((await completeWith('```json\n{"y": 4}\n```'))[4]).toBe('verification_failed'); + }); + + it('unfenced strips exactly one surrounding fence and nothing else', () => { + expect(unfenced('```json\n{"x":1}\n```')).toBe('{"x":1}'); + expect(unfenced(' ```\n[1]\n```\n')).toBe('[1]'); + expect(unfenced('{"x":1}')).toBe('{"x":1}'); + // Two fenced values are not one value: whatever is left must still fail JSON.parse. + expect(() => JSON.parse(unfenced('```json\n{}\n```\n```json\n{}\n```'))).toThrow(); + }); +}); From f2e535b386ecc5021c31e1b16e9eccabec110c52 Mon Sep 17 00:00:00 2001 From: Relayflow Lead Date: Tue, 22 Sep 2026 22:18:43 -0700 Subject: [PATCH 2/3] fix(sdk): accept tilde fences and longer closing runs in unfenced Addresses review: a fence opens with three or more backticks or tildes and closes with a run of the same character at least as long. Evidence is now regenerated by evidence/llm-fence-predicate-resume/verify.sh, which captures every literal command, including the restore and cmp. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/SURFACE.md | 6 ++- .../mutation-1-fence-restored.txt | 5 --- .../mutation-1-fence-reverted.txt | 9 ----- .../mutation-2-predicate-restored.txt | 4 -- .../mutation-2-predicate-reverted.txt | 10 ----- .../suite-vs-main.txt | 37 ----------------- evidence/llm-fence-predicate-resume/verify.sh | 40 +++++++++++++++++++ packages/sdk/src/llm-worker.ts | 11 +++-- packages/sdk/tests/worker-transcript.test.ts | 5 +++ 9 files changed, 56 insertions(+), 71 deletions(-) delete mode 100644 evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt delete mode 100644 evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt delete mode 100644 evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt delete mode 100644 evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt delete mode 100644 evidence/llm-fence-predicate-resume/suite-vs-main.txt create mode 100755 evidence/llm-fence-predicate-resume/verify.sh diff --git a/docs/SURFACE.md b/docs/SURFACE.md index 557bfd82d..41dbb6be7 100644 --- a/docs/SURFACE.md +++ b/docs/SURFACE.md @@ -400,8 +400,10 @@ a declared `model` must be in that project's `models` array. A template call such as ``await f.llm`Summarize ${text}` `` returns text. The structured overload parses JSON and checks `output` before submitting a successful completion; the kernel independently checks the schema before accepting the output. A -reply that is exactly one markdown code fence around a value is judged by the -value inside it; prose around the JSON, or two fenced values, is still invalid. +reply that is exactly one markdown code fence around a value (three or more +backticks or tildes, closed by a run of the same character at least as long) +is judged by the value inside it; prose around the JSON, or two fenced values, +is still invalid. Invalid JSON or a schema mismatch completes with `verification_failed` and prevents downstream work. Retry and lease handling use the existing kernel policies; this overload introduces no separate retry contract. diff --git a/evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt b/evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt deleted file mode 100644 index dbfcce816..000000000 --- a/evidence/llm-fence-predicate-resume/mutation-1-fence-restored.txt +++ /dev/null @@ -1,5 +0,0 @@ -# restored byte-for-byte (cmp exit 0) -$ npx vitest run tests/worker-transcript.test.ts -t "fenced reply" - ✓ tests/worker-transcript.test.ts (8 tests | 7 skipped) 373ms - ✓ the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value 373ms - Tests 1 passed | 7 skipped (8) diff --git a/evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt b/evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt deleted file mode 100644 index dbca272d6..000000000 --- a/evidence/llm-fence-predicate-resume/mutation-1-fence-reverted.txt +++ /dev/null @@ -1,9 +0,0 @@ -# mutation: llm-worker.ts parses result.stdout_tail without unfenced() -$ npx vitest run tests/worker-transcript.test.ts -t "fenced reply" - × the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value 403ms -⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ - FAIL tests/worker-transcript.test.ts > the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value -AssertionError: expected 'verification_failed' to be 'success' // Object.is equality -Expected: "success" -Received: "verification_failed" - Tests 1 failed | 7 skipped (8) diff --git a/evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt b/evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt deleted file mode 100644 index 8c594624e..000000000 --- a/evidence/llm-fence-predicate-resume/mutation-2-predicate-restored.txt +++ /dev/null @@ -1,4 +0,0 @@ -# restored byte-for-byte (cmp exit 0) -$ npx vitest run tests/authored-agent-artifacts.test.ts -t "sorted keys" - ✓ tests/authored-agent-artifacts.test.ts (5 tests | 4 skipped) 28ms - Tests 1 passed | 4 skipped (5) diff --git a/evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt b/evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt deleted file mode 100644 index c51156333..000000000 --- a/evidence/llm-fence-predicate-resume/mutation-2-predicate-reverted.txt +++ /dev/null @@ -1,10 +0,0 @@ -# mutation: applyPredicateGate lowers JSON.stringify(record) instead of the canonical field order -$ npx vitest run tests/authored-agent-artifacts.test.ts -t "sorted keys" - × predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run 30ms - → expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality -⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ - FAIL tests/authored-agent-artifacts.test.ts > predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run -AssertionError: expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality -Expected: "printf '%s' '{"gate":"predicate","step":"run-1","verdict":"pass","because":"plan covers every question"}'" -Received: "printf '%s' '{"because":"plan covers every question","gate":"predicate","step":"run-1","verdict":"pass"}'" - Tests 1 failed | 4 skipped (5) diff --git a/evidence/llm-fence-predicate-resume/suite-vs-main.txt b/evidence/llm-fence-predicate-resume/suite-vs-main.txt deleted file mode 100644 index f7f2bcf16..000000000 --- a/evidence/llm-fence-predicate-resume/suite-vs-main.txt +++ /dev/null @@ -1,37 +0,0 @@ -# Full SDK suite on this branch (packages/sdk: npm test), summary lines - ❯ tests/hosted-extension-isolation.test.ts (22 tests | 12 failed | 3 skipped) - ❯ tests/babysitter-native-extension.test.ts (41 tests | 41 skipped) - ❯ tests/hosted-base-snapshot.test.ts (18 tests | 15 failed | 1 skipped) - ❯ tests/authored-node-runtime.test.ts (14 tests | 14 skipped) - ❯ tests/mcp.test.ts (30 tests | 4 skipped) - ❯ tests/bundle.test.ts (26 tests | 1 failed) - ❯ tests/stop-process-group.test.ts (9 tests | 2 failed) - ❯ tests/hosted-extension-protocol.test.ts (24 tests | 7 failed | 1 skipped) - ❯ tests/worker-cli.test.ts (18 tests | 3 failed) - ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) - ❯ tests/cli-watch.test.ts (10 tests | 1 failed) - ❯ tests/canonical-software-factory.test.ts (3 tests | 3 failed) - ❯ tests/communication-mixed-resume.test.ts (1 test | 1 failed) - Test Files 13 failed | 166 passed | 3 skipped (182) - Tests 67 failed | 2753 passed | 72 skipped (2892) - -# The same 13 failing files on unmodified main 78cc5556 (git stash), summary lines -$ npx vitest run - ❯ tests/hosted-base-snapshot.test.ts (18 tests | 15 failed | 1 skipped) - ❯ tests/hosted-extension-isolation.test.ts (22 tests | 12 failed | 3 skipped) - ❯ tests/babysitter-native-extension.test.ts (41 tests | 41 skipped) - ❯ tests/authored-node-runtime.test.ts (14 tests | 14 skipped) - ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) - ❯ tests/canonical-software-factory.test.ts (3 tests | 3 failed) - ❯ tests/communication-mixed-resume.test.ts (1 test | 1 failed) - ❯ tests/mcp.test.ts (30 tests | 4 skipped) - ❯ tests/hosted-extension-protocol.test.ts (24 tests | 7 failed | 1 skipped) - Tests 60 failed | 114 passed | 64 skipped (238) - -# The four files that failed only in the full run, rerun on this branch in isolation (cwd packages/sdk) -$ npx vitest run tests/bundle.test.ts tests/stop-process-group.test.ts tests/worker-cli.test.ts tests/cli-watch.test.ts - ✓ tests/bundle.test.ts (26 tests) - ✓ tests/cli-watch.test.ts (10 tests) - ✓ tests/stop-process-group.test.ts (9 tests) - ✓ tests/worker-cli.test.ts (18 tests) - Tests 63 passed (63) diff --git a/evidence/llm-fence-predicate-resume/verify.sh b/evidence/llm-fence-predicate-resume/verify.sh new file mode 100755 index 000000000..a7062e5aa --- /dev/null +++ b/evidence/llm-fence-predicate-resume/verify.sh @@ -0,0 +1,40 @@ +#!/usr/bin/env bash +# Reproduces every claim in this directory. Run from packages/sdk: +# ../../evidence/llm-fence-predicate-resume/verify.sh +# Each step echoes the literal command, then its complete output and exit code. +set -u +OUT=../../evidence/llm-fence-predicate-resume +KEEP=$(mktemp -d) +FILES="tests/authored-node-runtime.test.ts tests/babysitter-native-extension.test.ts tests/bundle.test.ts tests/canonical-software-factory.test.ts tests/cli-watch.test.ts tests/communication-mixed-resume.test.ts tests/hosted-base-snapshot.test.ts tests/hosted-extension-isolation.test.ts tests/hosted-extension-protocol.test.ts tests/mcp.test.ts tests/stop-process-group.test.ts tests/stuck-run-triage.test.ts tests/worker-cli.test.ts" +show() { echo "\$ $*"; eval "$@" 2>&1; echo "exit=$?"; } + +# mutate +mutate() { + local name=$1 src=$2 expr=$3 test=$4 filter=$5 + { + show "cp $src $KEEP/$(basename "$src")" + show "sed -i.bak '$expr' $src && rm $src.bak" + show "git diff --stat -- $src" + show "npx vitest run $test -t '$filter'" + show "cp $KEEP/$(basename "$src") $src" + show "cmp $src $KEEP/$(basename "$src")" + show "git diff --stat -- $src" + show "npx vitest run $test -t '$filter'" + } > "$OUT/mutation-$name.txt" +} +mutate 1-fence src/llm-worker.ts 's/JSON.parse(unfenced(result.stdout_tail))/JSON.parse(result.stdout_tail)/' tests/worker-transcript.test.ts 'fenced reply' +mutate 2-predicate src/authored-flow-executor.ts 's/JSON.stringify(canonical)/JSON.stringify(record)/' tests/authored-agent-artifacts.test.ts 'sorted keys' + +# The 13 files that fail in the full suite here, with this change's two source +# files at main (the other changed files are tests outside this list), then as on this branch. +{ + show "cp src/llm-worker.ts src/authored-flow-executor.ts $KEEP/" + show "git show main:packages/sdk/src/llm-worker.ts > src/llm-worker.ts" + show "git show main:packages/sdk/src/authored-flow-executor.ts > src/authored-flow-executor.ts" + show "git diff --stat main -- src/llm-worker.ts src/authored-flow-executor.ts" + show "npx vitest run $FILES" + show "cp $KEEP/llm-worker.ts $KEEP/authored-flow-executor.ts src/" + show "cmp src/llm-worker.ts $KEEP/llm-worker.ts && cmp src/authored-flow-executor.ts $KEEP/authored-flow-executor.ts" +} > "$OUT/thirteen-files-at-main.txt" +show "npx vitest run $FILES" > "$OUT/thirteen-files-on-branch.txt" +show "npm test" > "$OUT/full-suite-on-branch.txt" diff --git a/packages/sdk/src/llm-worker.ts b/packages/sdk/src/llm-worker.ts index bb30d4029..a6271460b 100644 --- a/packages/sdk/src/llm-worker.ts +++ b/packages/sdk/src/llm-worker.ts @@ -102,10 +102,13 @@ export class LlmWorker extends EventEmitter { /** * A reply that is exactly one markdown code fence around a value, reduced to * that value. Models add the fence despite the instruction not to; the schema - * still judges what is inside it. Anything else (prose around the JSON, two - * fences) is returned unchanged and fails the parse as before. + * still judges what is inside it. A fence opens with three or more backticks + * or tildes and closes with a run of the same character at least as long. + * Anything else (prose around the JSON, two fences) is returned unchanged and + * fails the parse as before. */ export function unfenced(text: string): string { - const fence = /^\s*```[A-Za-z0-9_-]*[ \t]*\r?\n([\s\S]*?)\r?\n?```\s*$/.exec(text); - return fence === null ? text : fence[1]!; + const fence = /^\s*(?:(`{3,})[^`\n]*\n([\s\S]*?)\n?\1`*|(~{3,})[^\n]*\n([\s\S]*?)\n?\3~*)\s*$/.exec(text); + if (fence === null) return text; + return fence[2] ?? fence[4]!; } diff --git a/packages/sdk/tests/worker-transcript.test.ts b/packages/sdk/tests/worker-transcript.test.ts index 06d931321..f52ce8cf8 100644 --- a/packages/sdk/tests/worker-transcript.test.ts +++ b/packages/sdk/tests/worker-transcript.test.ts @@ -225,6 +225,11 @@ describe('the llm worker judges the value inside one markdown fence', () => { expect(unfenced('```json\n{"x":1}\n```')).toBe('{"x":1}'); expect(unfenced(' ```\n[1]\n```\n')).toBe('[1]'); expect(unfenced('{"x":1}')).toBe('{"x":1}'); + expect(unfenced('~~~json\n{"x":1}\n~~~')).toBe('{"x":1}'); + expect(unfenced('````json\n{"x":1}\n`````')).toBe('{"x":1}'); + // A closing run shorter than the opener, or of the other character, closes nothing. + expect(unfenced('````json\n{"x":1}\n```')).toBe('````json\n{"x":1}\n```'); + expect(unfenced('```json\n{"x":1}\n~~~')).toBe('```json\n{"x":1}\n~~~'); // Two fenced values are not one value: whatever is left must still fail JSON.parse. expect(() => JSON.parse(unfenced('```json\n{}\n```\n```json\n{}\n```'))).toThrow(); }); From 2d356ca7fcfb510b7d5a058f5fc773051490d24b Mon Sep 17 00:00:00 2001 From: Relayflow Lead Date: Tue, 22 Sep 2026 22:23:14 -0700 Subject: [PATCH 3/3] test(sdk): regenerate fence/predicate evidence with literal commands verify.sh captures every command it runs, including the restore and cmp, and the complete output of the 13 suspect files at main and on the branch and of the full suite. Co-Authored-By: Claude Opus 5.5 (1M context) --- .../full-suite-on-branch.txt | 1655 +++++++++++++++++ .../mutation-1-fence.txt | 58 + .../mutation-2-predicate.txt | 58 + .../thirteen-files-at-main.txt | 1209 ++++++++++++ .../thirteen-files-on-branch.txt | 1193 ++++++++++++ 5 files changed, 4173 insertions(+) create mode 100644 evidence/llm-fence-predicate-resume/full-suite-on-branch.txt create mode 100644 evidence/llm-fence-predicate-resume/mutation-1-fence.txt create mode 100644 evidence/llm-fence-predicate-resume/mutation-2-predicate.txt create mode 100644 evidence/llm-fence-predicate-resume/thirteen-files-at-main.txt create mode 100644 evidence/llm-fence-predicate-resume/thirteen-files-on-branch.txt diff --git a/evidence/llm-fence-predicate-resume/full-suite-on-branch.txt b/evidence/llm-fence-predicate-resume/full-suite-on-branch.txt new file mode 100644 index 000000000..530c099af --- /dev/null +++ b/evidence/llm-fence-predicate-resume/full-suite-on-branch.txt @@ -0,0 +1,1655 @@ +$ npm test + +> @relayflows/sdk@2.0.29 test +> sh scripts/test.sh + + +> @relayflows/sdk@2.0.29 test:prep +> ( cd ../../kernel && sh ../ops/cargo.sh build ) && ( [ ! -d ../../testdata/preflight ] || find ../../testdata/preflight -name '*-cli' -type f -exec chmod +x {} + ) + + Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.64s + +> @relayflows/sdk@2.0.29 typecheck +> tsc --noEmit && tsc -p tsconfig.type-tests.json + + +> @relayflows/sdk@2.0.29 build +> tsc && node scripts/make-cli-executable.mjs + + +> @relayflows/sdk@2.0.29 typecheck:tests +> tsc -p tsconfig.tests.json + + + RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk + +stdout | tests/live-kernel.test.ts +LIVE_KERNEL relayflowd=/Users/khaliqgant/.relayflows-toolchain/target/1166253295/debug/relayflowd +LIVE_KERNEL flows=/Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk/dist/cli.js + + ✓ tests/cloud-read.test.ts (41 tests) 38ms + ✓ tests/preflight.test.ts (67 tests) 98ms + ❯ tests/hosted-extension-isolation.test.ts (22 tests | 12 failed | 3 skipped) 187ms + × hosted extension capability isolation > resolves locked artifacts without importing extension top-level code 19ms + → expected [ { …(6) } ] to deeply equal [ { …(6) } ] + × hosted extension capability isolation > executes the exact capability-only handler for a queued receipt 6ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > executes the exact capability-only handler for a duplicate receipt 4ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > streams verified bytes when the live store is replaced and no writable staging path exists 16ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > mounts pinned private Surface bytes when the live package changes before launch 14ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > refuses oversized Surface files through the bounded descriptor reader 14ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + × hosted extension capability isolation > shields verified Surface files before async settlement 13ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > preserves a typed host refusal while disclosing only a fixed marker to the child 6ms + → expected Error: hosted extension isolation require… { code: '…' } to be Error: private Cloud policy detail { code: '…' } // Object.is equality + × hosted extension capability isolation > denies ambient credentials, host files, writes, network, subprocesses, and undeclared context verbs 8ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > enforces OS address-space and data bounds on native Buffer allocation 4ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + × hosted extension capability isolation > blocks extra handler fields and authority-bearing receipt fields at the parent port 5ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension capability isolation > constructs adapter authority with the captured freeze intrinsic 5ms + → hosted extension isolation requires Linux + ✓ tests/cloud-transcript-codex.test.ts (39 tests) 12ms + ✓ tests/plugin-extension.test.ts (90 tests) 520ms +(node:79026) ExperimentalWarning: SQLite is an experimental feature and might change at any time +(Use `node --trace-warnings ...` to show where the warning was created) + ✓ tests/observer-link.test.ts (39 tests) 142ms + ✓ tests/agent-transcript.test.ts (29 tests) 223ms + ✓ tests/authored-flow.test.ts (34 tests) 727ms + ❯ tests/babysitter-native-extension.test.ts (41 tests | 41 skipped) 379ms + ✓ tests/relay-cli-surface.test.ts (75 tests) 30ms + ✓ tests/cloud-run.test.ts (58 tests) 964ms + ✓ hosted v2 submission > submits exact authored UTF-8 bytes with source and pinned Surface authority 330ms + ✓ tests/cli-status.test.ts (27 tests) 2795ms + ✓ flows status > resolves the run with no arguments from inside a worker-spawned agent 2589ms + ✓ tests/daemon-lifecycle.test.ts (42 tests) 30ms + ✓ tests/cloud-sync.test.ts (40 tests) 4863ms + ✓ packWorkingTree > packs a Git checkout by ls-files semantics: tracked plus untracked, never ignored, .git or node_modules 2684ms + ✓ tests/run-state.test.ts (21 tests) 6ms + ✓ tests/step-failure-diagnostic.test.ts (25 tests) 2793ms + ✓ step failure diagnostic > surfaces command exit, stderr and replay hint through the CLI (json=false) 811ms + ✓ step failure diagnostic > surfaces command exit, stderr and replay hint through the CLI (json=true) 699ms + ✓ step failure diagnostic > exposes the middle failure through the CLI itself (json=false) 488ms + ✓ step failure diagnostic > exposes the middle failure through the CLI itself (json=true) 785ms + ✓ tests/cloud-deploy.test.ts (40 tests) 1339ms + ✓ deployToCloud > reports a missing or unloadable source as an input refusal (exit 2), before HTTP 308ms + ❯ tests/hosted-base-snapshot.test.ts (18 tests | 15 failed | 1 skipped) 978ms + × hosted base private snapshot > shadows inherited thenables on completed snapshot authority 11ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > locks the inherited then slot before authored code can schedule a replacement 4ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > keeps the buffered base bytes when the live source changes 1ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > keeps source digests sensitive when ambient Array.map is poisoned 1ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > defines source entries without consulting inherited numeric setters 2ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > uses the module-captured platform during source traversal 1ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > shadows native directory arrays before promise resolution can substitute them 1ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > excludes project node_modules from the admitted generation 2ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > refuses an oversized source file before buffering its contents 4ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > captures conversion and allocation across declarations and source snapshots 3ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > keeps declaration and reviewed-base reads bound after builtin export synchronization 2ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > reads through captured descriptors and charges admitted descriptor sizes 1ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > ignores inherited snapshot test hooks when production omits them 2ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > bounds a file that grows after its admitted size was checked 1ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > stops streaming project entries at the shared count limit 938ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + ✓ tests/cloud-connect.test.ts (24 tests) 3229ms + ✓ hosted verbs connect before they submit > flows run --cloud submits once the prompt connected the integration 2121ms + ✓ tests/journal-client.test.ts (17 tests) 73ms + ✓ tests/flow-extension-compose.test.ts (22 tests) 3803ms + ✓ composing flow extensions onto a base flow > composes two extensions in declaration order, and the order is the lockfile order 490ms + ✓ composing flow extensions onto a base flow > flows check reports the composition and keeps the composed triggers deliverable 980ms + ✓ tests/authored-root.test.ts (13 tests) 164ms + ❯ tests/authored-node-runtime.test.ts (14 tests | 14 skipped) 23ms + ✓ tests/close-pr-flow.test.ts (28 tests) 2104ms + ✓ close-pr journaled repair loop > executes the deterministic commit and force-push steps against a local Git remote, including a no-op repair 397ms + ✓ PR state parsing and shell boundaries > preserves gh output and handles exit status 0 300ms + ✓ PR state parsing and shell boundaries > preserves gh output and handles exit status 8 311ms + ✓ PR state parsing and shell boundaries > preserves gh output and handles exit status 2 377ms + ✓ PR state parsing and shell boundaries > preserves gh output and handles exit status 127 391ms + ✓ tests/cli.test.ts (70 tests) 14794ms + ✓ flows check CLI > binds a checked relative wrapper to the flow directory for worker execution 3332ms + ✓ flows check CLI > uses Codex login status and reports a rejected model as unavailable, not unauthenticated 325ms + ✓ flows check CLI > refuses a typo model before probing or contacting relayflowd 601ms + ✓ flows check CLI > runs a flow that declares no model against a flows.json that only names the cli 303ms + ✓ flows check CLI > keeps the project config path with the CLI clause when the step declares the model 391ms + ✓ flows check CLI > resolves a bare PATH-resolved claude with no declared model, in an isolated PATH 521ms + ✓ flows check CLI > distinguishes an allowlisted but inaccessible model from broken auth 498ms + ✓ flows check CLI > maps every input refusal path to its declared kind without raw exceptions 844ms + ✓ flows run/resume CLI over the journal protocol > parses run options, submits the kernel dialect, and exits 0 on success 534ms + ✓ flows run/resume CLI over the journal protocol > exits 1 and emits the declared completionReason for a failed run 486ms + ✓ flows run/resume CLI over the journal protocol > exits 3 and names the parked llm step 552ms + ✓ flows run/resume CLI over the journal protocol > reports a needs_human agent step as parked for human recovery 519ms + ✓ flows run/resume CLI over the journal protocol > classifies a typed hello refusal as a protocol error, not an unreachable daemon 532ms + ✓ flows run/resume CLI over the journal protocol > follows a dispatched worker step instead of reporting a protocol error 540ms + ✓ flows run/resume CLI over the journal protocol > bounds a worker wait by its lease and reports what it is waiting for 525ms + ✓ flows run/resume CLI over the journal protocol > resumes a parked run from snapshot step types without reading journal sequence one 500ms + ✓ flows run/resume CLI over the journal protocol > maps only run_not_found resumes to exit 2 1669ms + ✓ tests/validate.test.ts (68 tests) 35ms + ✓ tests/verb-field-lint.test.ts (96 tests) 225ms + ✓ tests/stop-process-group.test.ts (9 tests) 15724ms + ✓ every stop reaches the process group, not just the direct child > exits the run after an execution-timeout stop 1091ms + ✓ every stop reaches the process group, not just the direct child > exits the run after a protocol terminate stop 898ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after a protocol terminate stop 2324ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after an execution-timeout stop 2617ms + ✓ every stop reaches the process group, not just the direct child > holds the loop open long enough for the escalation to run 1139ms + ✓ a wrapper that exits with no execution deadline still drains > reports the wrapper result and reaps a grandchild holding its pipes 774ms + ✓ a wrapper that exits with no execution deadline still drains > reaps a SIGTERM-deaf grandchild holding its pipes 2112ms + ✓ a wrapper that exits with no execution deadline still drains > settles on its own deadline when an escaped holder withholds close 4490ms + ✓ tests/tick-source.test.ts (33 tests) 16ms + ❯ tests/mcp.test.ts (30 tests | 4 skipped) 9067ms + ✓ MCP preflight and transports > flows check refuses an undeclared server with exit 2 and no daemon 474ms + ✓ MCP preflight and transports > flows check reports a refusing server and leaves no PID 491ms + ✓ MCP preflight and transports > kills a SIGTERM-resistant silent child after a parent-owned handshake deadline 1323ms + ✓ MCP preflight and transports > reaps a SIGTERM-resistant descendant with inherit stdio before cleanup finishes 1137ms + ✓ MCP preflight and transports > reaps a SIGTERM-resistant descendant with ignore stdio before cleanup finishes 2078ms + ✓ MCP preflight and transports > reports malformed connection configuration as config_invalid 422ms + ✓ tests/authored-step-graph.test.ts (25 tests) 874ms + ✓ the authored step DAG > does not walk a long-running step's own polling chain to find its dependents' edges 328ms + ✓ tests/bundle.test.ts (26 tests) 6991ms + ✓ immutable bundles > verifies with --verify in any position and answers --json with one object 623ms + ✓ immutable bundles > builds and verifies the canonical YAML fixture through the compiled CLI 779ms + ✓ immutable bundles > emits the ephemeral warning on CLI stderr and uses the default output directory 507ms + ✓ immutable bundles > builds a standalone TS fixture twice with identical executable hashes 2268ms + ✓ immutable bundles > refuses invalid CLI arguments "--out" 367ms + ✓ immutable bundles > refuses invalid CLI arguments "--verify" 316ms + ✓ tests/authored-flow-lifecycle-executor.test.ts (27 tests) 570ms + ✓ tests/step-failure-excerpt.test.ts (42 tests) 102ms + ✓ tests/pr-review-post.test.ts (21 tests) 1901ms + ✓ tests/agent-relay-transport.test.ts (16 tests) 2581ms + ✓ Relay completion at the journal boundary > does not complete at readiness and journals exact output, receipt, and priced accounting 1016ms + ✓ Relay completion at the journal boundary > aborts polling on rejected renewal and never writes a stale completion 1006ms + ✓ tests/authored-flow-slack.test.ts (7 tests) 2053ms + ✓ authored Slack helper effects > replays after SIGKILL before confirm with the same token and one successful completion 557ms + ✓ authored Slack helper effects > replays after SIGKILL before complete with the same token and one successful completion 458ms + ✓ authored Slack helper effects > writes two files for two calls and supports dm, reply, and react 637ms + ✓ tests/tick-runner.test.ts (22 tests) 1728ms + ✓ CLI argument parsing refuses coercion rather than accepting it > refuses --interval-ms hex as an invocation error 313ms + ✓ CLI argument parsing refuses coercion rather than accepting it > refuses --interval-ms empty as an invocation error 326ms + ✓ tests/authored-step-index.test.ts (17 tests) 11ms +(node:89692) ExperimentalWarning: SQLite is an experimental feature and might change at any time +(Use `node --trace-warnings ...` to show where the warning was created) + ✓ tests/authored-agent-artifacts.test.ts (5 tests) 763ms + ✓ tests/cli-replay.test.ts (37 tests) 681ms + ✓ flows replay > --json is byte-identical across two CLI invocations (diff) 497ms + ✓ tests/authored-node-result.test.ts (39 tests) 10ms + ✓ tests/gate-contract.test.ts (20 tests) 107ms + ✓ tests/authored-human.test.ts (13 tests) 80ms + ✓ tests/worker-cli.test.ts (18 tests) 28772ms + ✓ registered CLI model defaults > passes the same priced Claude default to the real provider invocation 1569ms + ✓ step discovery environment > names the run, step, attempt and an absolute data dir for a direct agent spawn 596ms + ✓ step discovery environment > exports none of the four without a data dir, even when the worker inherited them 641ms + ✓ wrapper discovery environment > sets the four names from the dispatch and still refuses ambient values and other secrets 483ms + ✓ wrapper discovery environment > exports none of the four to a wrapper without a data dir, even when the worker inherited them 486ms + ✓ custom wrapper execution identity > passes an explicit safe environment at identification and execution 780ms + ✓ custom wrapper execution identity > refuses a wrapper symlink retarget before delivering private values 762ms + ✓ custom wrapper execution identity > bounds wrapper execution after acknowledgement 635ms + ✓ custom wrapper execution identity > bounds captured wrapper output 544ms + ✓ custom wrapper execution identity > refuses a duplicate execute protocol frame 490ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a conforming wrapper leaks a stdio pipe to a background helper 2414ms + ✓ custom wrapper execution bounds are reader-owned > resolves when the leaked helper inherits stderr only 2051ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a wrapper leaks a stdio pipe and exits before identifying 3456ms + ✓ custom wrapper execution bounds are reader-owned > journals a completionReason at the default bound when a wrapper leaks a stdio pipe 11457ms + ✓ custom wrapper execution bounds are reader-owned > accepts an execute token and an over-8KiB payload flushed in one write 451ms + ✓ custom wrapper execution bounds are reader-owned > accepts the same over-8KiB payload whether or not it coalesces with the execute token 1210ms + ✓ custom wrapper execution bounds are reader-owned > still bounds an un-terminated handshake buffer and names the bound 361ms + ✓ delivers the journaled memory pack to the real wrapper and excludes its charge from completion usage 386ms + ✓ tests/cloud-schedule.test.ts (17 tests) 2932ms + ✓ schedule lowering > marks a non-grid cron as Cloud-only rather than approximating it, with a silence budget from its own cadence 1227ms + ✓ flows check prints declared schedules > shows the lowering for a fixed interval and the Cloud-only note for a real cron 1296ms + ✓ tests/flow-executor-chain.test.ts (14 tests) 14672ms + ✓ flow executor LLM and output-binding chain > runs f.llm -> f.agent -> f.run with schema-verified journal output and the exact allowed model 1205ms + ✓ flow executor LLM and output-binding chain > runs a dollar-budgeted authored Claude agent with the same default used by preflight 612ms + ✓ flow executor LLM and output-binding chain > fails invalid LLM output before the next step: not JSON 500ms + ✓ flow executor LLM and output-binding chain > fails invalid LLM output before the next step: {"message":7} 582ms + ✓ flow executor LLM and output-binding chain > retains tagged-template text output 463ms + ✓ flow executor LLM and output-binding chain > preserves JSON values without promoting them to process wrappers: null 472ms + ✓ flow executor LLM and output-binding chain > preserves JSON values without promoting them to process wrappers: [1,2] 590ms + ✓ flow executor LLM and output-binding chain > preserves JSON values without promoting them to process wrappers: "hello" 700ms + ✓ flow executor LLM and output-binding chain > runs the exact authored flagship f.llm -> f.agent -> f.run path through the durable CLI root 2593ms + ✓ flow executor LLM and output-binding chain > resumes an interrupted durable authored root without replaying completed flagship effects 3839ms + ✓ flow executor LLM and output-binding chain > passes a declarative verified value through an agent into a deterministic artifact 791ms + ✓ flow executor LLM and output-binding chain > journals a missing optional field as a failure before the consuming command executes 415ms + ✓ flow executor LLM and output-binding chain > flows run consumes YAML bindings and resume reuses the original journal output 1840ms + ✓ tests/cli-hn-monitor.test.ts (16 tests) 93ms + ✓ tests/authored-run-failure-evidence.test.ts (9 tests) 1093ms + ✓ the child index after the process that wrote it is gone > still names every child, with its own run id, after a daemon restart 523ms + ✓ tests/worker-transcript.test.ts (8 tests) 1740ms + ✓ the llm worker bounds and redacts its error and journals the digest > cuts a 10 KiB stderr to 4 KiB, states the cut, and keeps the completion under the kernel cap 322ms + ✓ the llm worker judges the value inside one markdown fence > still refuses prose around the JSON, and a fenced value that breaks the schema 398ms + ✓ tests/wrapper-execution-duration.test.ts (7 tests) 12412ms + ✓ treats an explicit executionTimeoutMs of 0 as no deadline, not as the old fallback 428ms + ✓ keeps the handshake deadline independent of the removed execution deadline 10442ms + ✓ still lets a lease abort stop an unlimited wrapper before it produces output 523ms + ✓ tests/direct-input.test.ts (6 tests) 9235ms + ✓ direct .flow.ts input through the built CLI and live runtime > returns exit 3 for an authored human handoff and persists its outcome 924ms + ✓ direct .flow.ts input through the built CLI and live runtime > returns exit 1 for an authored step_failed verdict and persists its outcome 948ms + ✓ direct .flow.ts input through the built CLI and live runtime > executes inline and file JSON input through relayflowd 2444ms + ✓ direct .flow.ts input through the built CLI and live runtime > refuses missing and malformed input before contacting relayflowd 2989ms + ✓ direct .flow.ts input through the built CLI and live runtime > does not run the authored body before daemon availability 992ms + ✓ direct .flow.ts input through the built CLI and live runtime > refuses oversized file input before contacting relayflowd 937ms +(node:91656) Warning: Transcript tail for run-9/analyze attempt 1 (stdout) could not be written; the step continues without it: EACCES: permission denied, mkdir '/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/transcript-tail-amUxYR/runs/run-9/steps' +(Use `node --trace-warnings ...` to show where the warning was created) + ✓ tests/transcript-tail.test.ts (11 tests) 1357ms + ✓ direct agent spawn > tees stdout and stderr into tail files that name the dispatch 487ms + ✓ direct agent spawn > completes the step when the tail directory cannot be created 395ms + ✓ tests/authored-flow-operation.test.ts (24 tests) 480ms + ✓ tests/flow-requirements.test.ts (14 tests) 491ms + ✓ flows check prints REQUIRES > names the helper, the harness and the mcp server of an authored flow 326ms + ✓ tests/backlog-picker.test.ts (14 tests) 44ms + ✓ tests/plugin-store-bounds.test.ts (11 tests | 1 skipped) 129ms + ✓ tests/backlog-picker-flow.test.ts (6 tests) 587ms + ✓ tests/hosted-extension-protocol-intrinsics.test.ts (6 tests) 11ms + ✓ tests/preflight-permissions-unenforced.test.ts (17 tests) 760ms + ✓ tests/authored-helpers.test.ts (6 tests) 5409ms + ✓ lowers the named acceptance helpers and Slack to confirmed journal effects 352ms + ✓ runs every available provider through the real kernel and resumes completed effects without a second write 3426ms + ✓ replays after SIGKILL before confirm with the same token and one successful completion 723ms + ✓ replays after SIGKILL before complete with the same token and one successful completion 623ms + ❯ tests/hosted-extension-protocol.test.ts (24 tests | 7 failed | 1 skipped) 40365ms + × hosted extension hostile protocol > rejects an import-time different PR frame with zero adapter calls 12ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects an import-time different delivery frame with zero adapter calls 4ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects an import-time different event frame with zero adapter calls 6ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects two forged calls after the authoritative first outcome settles 10009ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > waits for a pending adapter to reject after a forged child error 10002ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > waits for a pending adapter to resolve after a forged child error 10004ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > returns a typed adapter rejection even when the hostile child hangs 10264ms + → hostile child did not invoke the adapter + ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) 222ms + × stuck-run-triage input validation > refuses an 8-character run-id prefix: Cloud has no prefix lookup 5ms + → expected [Function] to throw error matching /not full Cloud run ids: c649fe14/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses the whole batch when any id is invalid, rather than dropping it 1ms + → expected [Function] to throw error matching /not full Cloud run ids: nope!/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses an empty batch 1ms + → expected [Function] to throw error matching /needs runIds/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses a batch too large for the edge step lease 0ms + → expected [Function] to throw error matching /exceeds the 8 that fit/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > accepts eight ids — the incident batch is inside the bound 4ms + → promise rejected "TypeError: expected an @relayflows/surfac…" instead of resolving + × stuck-run-triage apiUrl > refuses to send the Cloud bearer token to an unapproved origin 1ms + → expected [Function] to throw error matching /refusing to send the Cloud bearer to…/\ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage apiUrl > refuses a non-URL apiUrl 0ms + → expected [Function] to throw error matching /is not a URL/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage apiUrl > allows an approved origin and uses it in the curl 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage apiUrl > defaults to production Cloud 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage apiUrl > never publishes a run record the fetch did not produce 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > names the Worker on every wrangler invocation 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > accepts caller-supplied Workers and rejects option-shaped ones 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > falls back when GNU timeout is absent, as it is on macOS 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > runs the tails concurrently so wall time does not scale with the batch 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > records wrangler's own exit status rather than head's 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage shell text > parses under both sh and bash 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage shell text > collects tails with no GNU timeout on PATH, as on a stock macOS 205ms + → expected an @relayflows/surface flow handle + × stuck-run-triage agents > declares read-only permissions on every agent 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage agents > tells the forensics agents their evidence is untrusted 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage fan-out > refuses a duplicate run id: two tails would share one evidence file 1ms + → expected [Function] to throw error matching /duplicate runIds: c649fe14-0c2e-4e51-…/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage fan-out > refuses a duplicate Worker name for the same reason 1ms + → expected [Function] to throw error matching /duplicate workers: w-one/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage fan-out > bounds ids x workers, not just ids 1ms + → expected [Function] to throw error matching /24 concurrent tails, over the 16/ but got 'expected an @relayflows/surface flow …' + ✓ tests/hosted-extension-routing.test.ts (7 tests) 7ms + ✓ tests/daemon-lifecycle-live.test.ts (9 tests) 12418ms + ✓ flows run against a data dir with no daemon (§6 test 7) > cold start spawns exactly one daemon, the run succeeds, and the daemon outlives the CLI 956ms + ✓ flows run against a data dir with no daemon (§6 test 7) > polls, bounded, for a daemon that holds the lock before it binds 1437ms + ✓ flows run against a data dir with no daemon (§6 test 7) > attaches to a serving daemon that has not published a connection file 1085ms + ✓ flows run against a data dir with no daemon (§6 test 7) > a second run attaches to the daemon the first one started, spawning nothing 2215ms + ✓ flows run against a data dir with no daemon (§6 test 7) > detects a stale connection file left by a hard kill and starts a fresh daemon 1903ms + ✓ concurrent invocations against one empty data dir (§6 test 15) > ends with exactly one daemon owning the socket, and both runs succeed 1282ms + ✓ refusals from a spawn that cannot produce a daemon > names relayflowd_not_found rather than falling through to PATH 1077ms + ✓ refusals from a spawn that cannot produce a daemon > names daemon_start_failed and quotes the daemon log when startup dies 1234ms + ✓ refusals from a spawn that cannot produce a daemon > refuses a daemon speaking another protocol version instead of binding over it 1227ms + ✓ tests/artifact-gates.test.ts (7 tests) 337ms + ✓ tests/webhook.test.ts (9 tests) 851ms + ✓ webhook ingress > checks TS declarations against flows.json without invoking handlers 527ms + ✓ tests/wrapper-exit-drain.test.ts (8 tests) 5385ms + ✓ settles a wrapper that exits leaving stdout held open, at default limits 502ms + ✓ settles a wrapper that exits leaving only stderr held open 1338ms + ✓ reports the wrapper’s nonzero exit code rather than the drain 600ms + ✓ reports a signalled wrapper death while a pipe is held, with its output intact 611ms + ✓ flushes an un-terminated final line exactly once when the drain settles 533ms + ✓ refuses a duplicate execute frame found at finalization, despite a clean exit 525ms + ✓ ignores output a holder writes after the drain has settled 520ms + ✓ lets a lease abort outrank a successful exit still being drained 756ms + ✓ tests/authored-step-failed.test.ts (10 tests) 32ms + ✓ tests/human-live.test.ts (3 tests) 4660ms + ✓ f.human against a real daemon > parks with the question, refuses wrong answers, records one, and resumes to success 2887ms + ✓ f.human against a real daemon > a "no" is a value the body branches on: declined, exit 0, no effect 1133ms + ✓ f.human against a real daemon > refuses to answer a run the daemon does not know 640ms + ✓ tests/budget-preflight.test.ts (25 tests) 10ms + ✓ tests/webhook-live.test.ts (6 tests) 9757ms + ✓ executes and deduplicates 'app_mention' only for its provider and matching payload 1738ms + ✓ executes and deduplicates 'reaction_added' only for its provider and matching payload 1514ms + ✓ executes and deduplicates 'pull_request' only for its provider and matching payload 1483ms + ✓ flows serve-webhook writes JSON before the daemon starts, then journals and archives exactly once 1369ms + ✓ replays a dropped file after SIGKILL before spawn 318ms + ✓ resumes the same journal after SIGKILL after spawn and before acknowledgement 3334ms + ✓ tests/budget-unmetered-live.test.ts (3 tests) 1798ms + ✓ unmetered budget spend through the live kernel > runs an unpriced step under a dollar budget without tripping it, journaling unknown dollars 746ms + ✓ unmetered budget spend through the live kernel > still counts an unpriced step toward a token budget 630ms + ✓ unmetered budget spend through the live kernel > accrues a priced step and stops the run when it crosses the dollar budget 422ms + ✓ tests/provider-trigger-contract.test.ts (7 tests) 523ms + ✓ tests/work-package-consumer.test.ts (13 tests) 145ms + ✓ tests/spec-parity.test.ts (31 tests) 279ms + ✓ tests/helpers-fanout.test.ts (96 tests) 609ms +(node:96239) ExperimentalWarning: SQLite is an experimental feature and might change at any time +(Use `node --trace-warnings ...` to show where the warning was created) + ✓ tests/authored-step-graph-live.test.ts (1 test) 760ms + ✓ the authored step DAG through the live kernel > carries labels and predecessors on every index record and journal step, ids unchanged 760ms + ❯ tests/cli-watch.test.ts (10 tests | 1 failed) 16175ms + ✓ flows check --watch > rechecks syntax errors, clears once, and returns the last refusal on Ctrl-C 1312ms + ✓ flows check --watch > streams JSON lines without ANSI, recovers after atomic saves, and exits zero after repair 1498ms + ✓ flows check --watch > coalesces 20 concurrent saves into at most two rechecks 1518ms + ✓ flows check --watch > watches transitive relative use imports, cycles, and nearest config changes 1741ms + ✓ flows check --watch > refreshes the import graph and notices missing imports being created 1738ms + ✓ flows check --watch > reloads authored TypeScript instead of reusing the first imported definition 1048ms + × flows check --watch > detects a nearer config appearing and falls back after it is deleted 5011ms + → Test timed out in 5000ms. +If this is a long-running test, pass a timeout value as the last argument or configure it globally with "testTimeout". + ✓ flows check --watch > keeps watching after the target is deleted and recreated 1543ms + ✓ flows check --watch > queues changes during a slow check without overlapping checks 763ms + ✓ tests/generate-triggers.test.ts (7 tests) 1082ms + ✓ discovers new adapters, preserves exact event names, and prefers adapter-local mappings 314ms + ✓ tests/webhook-hardening.test.ts (11 tests) 102ms + ✓ tests/human-to.test.ts (8 tests) 9ms + ✓ tests/plugin-loader.test.ts (9 tests) 186ms +stdout | tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI +LIVE_ANALYZER ready: claude -p --model claude-haiku-4-5-20251001 round-trip OK + + ✓ tests/authored-parallel-agents.test.ts (8 tests) 14518ms + ✓ authored steps under local workers with capacity > runs more concurrent f.llm calls than the worker holds side by side, never more than its capacity 1527ms + ✓ authored steps under local workers with capacity > completes more concurrent f.agent calls than the worker holds: the overflow waits for a slot instead of parking 3140ms + ✓ authored steps under local workers with capacity > runs agents in distinct working directories side by side (the kernel carries cwd) 1179ms + ✓ authored steps under local workers with capacity > serializes agents whose cwd is a symlink alias of the same directory 1334ms + ✓ authored steps under local workers with capacity > serializes agents whose cwd is a directory nested inside the other 1387ms + ✓ authored steps under local workers with capacity > never starts queued agents once the body has failed 2256ms + ✓ authored steps under local workers with capacity > never starts a queued agent when the agent holding the only slot fails 2502ms + ✓ authored steps under local workers with capacity > parks the overflow when the body is not told the capacity (the defect this closes) 1191ms + ✓ tests/worker-lease.test.ts (7 tests) 7ms + ✓ tests/yaml-helpers.test.ts (33 tests) 78ms + ✓ tests/pty-sidechannel.test.ts (11 tests) 8078ms + ✓ view attach preserves worker completion and marks only drive 1238ms + ✓ drive attach preserves worker completion and marks only drive 550ms + ✓ passthrough attach preserves worker completion and marks only drive 986ms + ✓ none attach preserves worker completion and marks only drive 1028ms + ✓ none subscriber lets an unattended CLI read EOF 461ms + ✓ view subscriber lets an unattended CLI read EOF 561ms + ✓ passthrough subscriber lets an unattended CLI read EOF 719ms + ✓ incomplete subscriber lets an unattended CLI read EOF 773ms + ✓ rejects drive after EOF without marking human intervention 953ms + ✓ delivers all drive bytes in order across child stdin backpressure 805ms + ✓ tests/authored-agent-permissions.test.ts (26 tests) 3294ms + ✓ authored agent permissions > lowers the full readonly declaration 332ms + ✓ authored agent permissions > preserves partial declarations without defaults: {"fileGlobs":["drafts/**"]} 300ms + ✓ authored agent permissions > preserves partial declarations without defaults: {"accessPreset":"readonly"} 424ms + ✓ authored agent permissions > preserves absence: {} 411ms +stdout | tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI +LIVE_ANALYZER analysis: {"reasoning":"This story directly describes an autonomous AI agent performing software development tasks (opening and reviewing pull requests), which is core to the AI agents and automation domain. The capability demonstrated aligns with advanced agent reasoning and self-directed task execution patterns central to modern automation systems.","relevance_score":9,"story_title":"Show HN: an agent that opens and reviews its own pull requests [wake-nonce-7f3a91c4]"} + + ✓ tests/babysitter-catalog-export.test.ts (14 tests) 795ms + ✓ Babysitter catalog artifact export > CLI refuses an existing output and leaves no file on validation failure 648ms + ✓ tests/redact.test.ts (35 tests) 4ms + ✓ tests/communication.test.ts (10 tests) 20ms + ✓ tests/typed-output.test.ts (14 tests) 209ms + ❯ tests/canonical-software-factory.test.ts (3 tests | 3 failed) 172ms + × canonical software-factory metadata contract > opens the actual catalog flow with the ticket title and exactly one GitHub closing line 119ms + → expected an @relayflows/surface flow handle + × canonical software-factory metadata contract > normalizes whitespace and caps the title at 240 Unicode code points 52ms + → expected an @relayflows/surface flow handle + × canonical software-factory metadata contract > fails closed before push when GitHub identity is missing or the final body duplicates its closing line 0ms + → expected an @relayflows/surface flow handle + ✓ tests/budget-attribution.test.ts (5 tests) 5ms + ✓ tests/deploy.test.ts (11 tests) 5078ms + ✓ flows deploy file buckets > publishes the full signed layout byte-for-byte and redeploys as a noop 870ms + ✓ flows deploy file buckets > answers --json with one object per outcome 703ms + ✓ flows deploy file buckets > reports a refusal as JSON under --json 366ms + ✓ flows deploy file buckets > refuses a missing local bundle before creating the bucket 397ms + ✓ flows deploy file buckets > refuses an unreachable bucket before copying 386ms + ✓ flows deploy file buckets > refuses an unwritable bucket 460ms + ✓ flows deploy file buckets > refuses local tampering of spec.canonical.json 399ms + ✓ flows deploy file buckets > refuses local tampering of identity.json 344ms + ✓ flows deploy file buckets > refuses asset bundles instead of using daemon-relative files 338ms + ✓ flows deploy file buckets > never labels a corrupt existing deployment as a noop 753ms + ✓ tests/json-schema-bound.test.ts (71 tests) 2392ms + ✓ JSON Schema termination bound > walks a deep schema with an explicit stack rather than recursion 1931ms + ✓ tests/mcp-lifecycle.test.ts (4 tests) 10ms + ✓ tests/effect-channel.test.ts (5 tests) 701ms + ✓ tests/model-selection.test.ts (10 tests) 20ms + ✓ tests/relayflowd-path.test.ts (10 tests) 5ms + ✓ tests/agent-artifacts.test.ts (9 tests) 39ms + ✓ tests/agent-transcript-live.test.ts (4 tests) 43748ms + ✓ the transcript digest through the built CLI, a real daemon and the local agent > preserves structured agent failure details and its completed root index 15268ms + ✓ the transcript digest through the built CLI, a real daemon and the local agent > preserves structured llm failure details and its completed root index 13064ms + ✓ the transcript digest through the built CLI, a real daemon and the local agent > journals the digest in trajectory_tail on a successful agent step and writes the file it points at 1032ms + ✓ the transcript digest through the built CLI, a real daemon and the local agent > on a failed agent step, names the failure and the transcript in the terminal diagnostic, redacted 14384ms + ✓ tests/authored-plugin-effect.test.ts (6 tests) 162ms + ✓ tests/agent-artifacts-live.test.ts (5 tests) 47193ms + ✓ agent artifacts and gates through the built CLI, a real daemon and the local agent > journals the files the agent wrote, including under a dot-directory, and every artifact gate passes on that journal 1424ms + ✓ agent artifacts and gates through the built CLI, a real daemon and the local agent > fails the run when the artifact_exists gate names a file the agent did not write 12923ms + ✓ agent artifacts and gates through the built CLI, a real daemon and the local agent > fails the run with the author reason when a predicate gate returns false, journaling the verdict 12930ms + ✓ review follow-ups > applies a predicate gate on a helper step too, and journals its verdict 14443ms + ✓ review follow-ups > records predicate verdicts on the root run so a resume reuses them instead of re-running the closure 5472ms + ✓ tests/f-memory.test.ts (7 tests) 1370ms + ✓ attaches read helpers without journaling read steps 336ms + ✓ tests/local-dev-ux.test.ts (8 tests) 61ms + ✓ tests/worker-slots.test.ts (7 tests) 7ms +stdout | tests/live-kernel.test.ts > surface resume after a real daemon kill > resumes a three-step run with each successful completion exactly once +LIVE_KERNEL kill -9 pid=98068 run=01M36BD299CBXXA2GRZWSKXZ25 while step=two state=Running + + ↓ tests/relay-cli-surface-live.test.ts (3 tests | 3 skipped) + ✓ tests/authored-declined.test.ts (13 tests) 110ms + ✓ tests/resume-failure.test.ts (2 tests) 5ms + ✓ tests/authored-hooks.test.ts (5 tests) 8ms + ✓ tests/worker-cli-result-exit.test.ts (5 tests) 33809ms + ✓ a Claude agent step completes on its result, not only on process exit > settles a hung, successful run within the grace and stops its whole tree 31838ms + ✓ a Claude agent step completes on its result, not only on process exit > maps an error result on a hung run to a failed exit 32206ms + ✓ a Claude agent step completes on its result, not only on process exit > leaves a hang before any result to the existing stops 32045ms + ✓ a Claude agent step completes on its result, not only on process exit > settles a normal exit at once with the real exit code 345ms + ✓ an agent tree does not outlive the process that spawned it > kills the agent group when the run process is terminated by SIGTERM 1247ms + ✓ tests/input-binding.test.ts (12 tests) 201ms + ✓ tests/dependency-validation.test.ts (6 tests) 621ms + ✓ dependency validation > accepts a valid 10,000-step reverse chain through every direct public boundary 363ms + ✓ tests/communication-review.test.ts (5 tests) 1510ms + ✓ rejects missing, incorrect, and another session token before invoking journal operations 1196ms + ✓ tests/deterministic-llm.test.ts (5 tests) 46ms + ✓ tests/scope-preflight.test.ts (6 tests) 274ms + ✓ tests/yaml-helper-effect.test.ts (4 tests) 888ms + ✓ dispatches the compiled helper from the SDK agent worker without a CLI 630ms + ✓ tests/scope-compiler.test.ts (25 tests) 21ms + ✓ tests/live-kernel.test.ts (31 tests) 88941ms + ✓ built flows CLI against live relayflowd > twenty-six-step reuses 25 durable completions after editing the failed final step 4728ms + ✓ built flows CLI against live relayflowd > runs rung (a), parks rung (b), and keeps JSON report-shaped 5455ms + ✓ built flows CLI against live relayflowd > allows a deterministic run to exceed the bounded request timeout 32375ms + ✓ built flows CLI against live relayflowd > follows a live worker dispatch through flows run 2141ms + ✓ built flows CLI against live relayflowd > runs an agent CLI end to end through the SDK worker 1125ms + ✓ built flows CLI against live relayflowd > f.agent lowers to a real agent step and dispatches through a live worker 921ms + ✓ built flows CLI against live relayflowd > f.agent's default flowPath anchors on cwd, not cwd's parent 382ms + ✓ built flows CLI against live relayflowd > can always get a parked run to a late-attaching worker 5577ms + ✓ built flows CLI against live relayflowd > reports a real manual-recovery NeedsHuman state as parked 904ms + ✓ built flows CLI against live relayflowd > hn-monitor analyze-story FAILS verification when the CLI omits required schema fields 425ms + ✓ built flows CLI against live relayflowd > agent step preserves the CliResult wrapper as output when the CLI emits non-JSON text 345ms + ✓ built flows CLI against live relayflowd > AgentWorker exposes wake_context to the CLI via RELAYFLOW_WAKE_CONTEXT env var (real analyzer prerequisite) 678ms + ✓ built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_WAKE_CONTEXT UNSET when the run has no wake_context (undefined-vs-null pin) 552ms + ✓ built flows CLI against live relayflowd > AgentWorker passes a declared model to an identified wrapper as RELAYFLOW_MODEL 658ms + ✓ built flows CLI against live relayflowd > AgentWorker refuses a nonconforming journal-submitted wrapper before exposing RELAYFLOW_MODEL 797ms + ✓ built flows CLI against live relayflowd > AgentWorker executes the raw claude adapter with its real model flag 606ms + ✓ built flows CLI against live relayflowd > AgentWorker executes the raw codex adapter with its real model flag 597ms + ✓ built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_MODEL UNSET when the step declares no model 379ms + ✓ built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI 12263ms + ✓ built flows CLI against live relayflowd > preflights before journaling and names an unreachable socket 6195ms + ✓ built flows CLI against live relayflowd > starts exactly one daemon when two runs race for one empty data dir 1325ms + ✓ surface resume after a real daemon kill > resumes a three-step run with each successful completion exactly once 3836ms + ✓ a relayflow can be scheduled: tick source against live relayflowd > a tick spawns a real run whose step reports the SCHEDULED instant 5707ms + ✓ tests/build-gate.test.ts (3 tests) 2043ms + ✓ flows build gates on flows check green (#318) > refuses a flow with an unresolvable named-agent CLI and leaves no artifacts 1010ms + ✓ flows build gates on flows check green (#318) > --json emits one CheckReport object on stdout on refusal, exits 2, no artifacts 584ms + ✓ flows build gates on flows check green (#318) > builds the bundle on success (regression: gate must not block valid flows) 448ms + ✓ tests/hn-poller.test.ts (6 tests) 9ms + ✓ tests/bin.test.ts (7 tests) 3643ms + ✓ built flows binary > refuses through a symlink to the built artifact 1013ms + ✓ built flows binary > refuses through a symlinked directory component 790ms + ✓ built flows binary > classifies a signal-terminated auth probe as probe_failed 488ms + ✓ built flows binary > classifies an unavailable PATH resolver as probe_failed 406ms + ✓ built flows binary > does not describe a present non-executable CLI as missing 458ms + ✓ built flows binary > runs one auth probe for three steps sharing a flow CLI 487ms + ✓ tests/communication-worker.test.ts (15 tests) 1533ms + ✓ tests/authored-step-failed-exit.test.ts (3 tests) 6ms + ✓ tests/yaml-local-agent-live.test.ts (7 tests) 12669ms + ✓ YAML --local-agent through the built CLI and real daemon > runs with the checked step CLI and model and journals done 1162ms + ✓ YAML --local-agent through the built CLI and real daemon > runs with the checked named CLI and model and journals done 1356ms + ✓ YAML --local-agent through the built CLI and real daemon > runs with the checked flow CLI and model and journals done 5863ms + ✓ YAML --local-agent through the built CLI and real daemon > runs with the checked project CLI and model and journals done 1412ms + ✓ YAML --local-agent through the built CLI and real daemon > still parks without --local-agent 782ms + ✓ YAML --local-agent through the built CLI and real daemon > reports the agent process failure 960ms + ✓ YAML --local-agent through the built CLI and real daemon > preserves declared workspace surfaces that the local worker cannot pin 1133ms + ✓ tests/dir-watcher-poller.test.ts (6 tests) 7ms + ✓ tests/model-pricing.test.ts (10 tests) 6ms + ✓ tests/direct-run-failure.test.ts (8 tests) 24ms + ✓ tests/plugin-add.test.ts (7 tests) 1989ms + ✓ installs a real offline npm fixture and includes declarations 479ms + ✓ typechecks the augmented verb and rejects unknown namespaces 1482ms + ✓ tests/provider-trigger-executor.test.ts (4 tests) 280ms + ✓ tests/hello-deterministic.test.ts (5 tests) 27ms + ✓ tests/yaml-helper-live.test.ts (1 test) 1327ms + ✓ runs compiled YAML helpers through the built CLI and kernel effect journal 1327ms + ✓ tests/wrapper-artifacts-cwd.test.ts (2 tests) 816ms + ✓ wrapper artifacts follow the requested cwd > spawns the wrapper in `cwd` and measures artifacts there, not in the worker process directory 413ms + ✓ wrapper artifacts follow the requested cwd > never reports the kernel data dir, even when the daemon writes there while the agent runs 402ms + ✓ tests/cli-adapter.test.ts (4 tests) 3ms + ✓ tests/transcript-exclusion-timeout.test.ts (1 test) 423ms + ✓ a transcript close that outruns its deadline > still keeps the transcript out of the agent's artifacts 422ms + ✓ tests/work-package-validator.test.ts (7 tests) 21ms + ❯ tests/communication-mixed-resume.test.ts (1 test | 1 failed) 175ms + × resumes mixed ordinary and linked agents through the real daemon without stealing peer capacity 174ms + → {"ok":false,"command":"resume","resolutions":[],"diagnostics":[{"severity":"failure","kind":"protocol_error","message":"relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`"}],"runId":"01M36BDG520KC244R6W7NNWDHX","socketPath":"/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-bcaf86b64a3c.sock"}: expected { ok: false, command: 'resume', …(4) } to match object { status: 'completed', …(2) } +(6 matching properties omitted from actual) + ✓ tests/transcript-tail-close.test.ts (2 tests) 2663ms + ✓ a stalled transcript-tail close > does not hold the spawn open past its bounded window 1471ms + ✓ a stalled tail close beside a transcript that finished > still journals the transcript pointer 1191ms + ✓ tests/run-from-digest.test.ts (6 tests) 6453ms + ✓ flows run digest input > submits the sealed canonical spec through the normal journal path without checkout 738ms + ✓ flows run digest input > uses a verified cache hit even after the bucket is removed 527ms + ✓ flows run digest input > resolves deploy.bucket from flows.json and honors explicit override 1805ms + ✓ flows run digest input > refuses an unconfigured bucket 976ms + ✓ flows run digest input > refuses tampered spec.canonical.json before creating run data 1503ms + ✓ flows run digest input > refuses tampered identity.json before creating run data 903ms + ✓ tests/cli-answer.test.ts (15 tests) 9ms + ✓ tests/agent-relay-hardening.test.ts (12 tests) 11ms + ✓ tests/authored-use-loader.test.ts (5 tests) 1381ms + ✓ authored use graph loader > loads a diamond in dependency order with one node per canonical path 667ms + ✓ tests/authored-declined-live.test.ts (1 test) 2010ms + ✓ runs an input guard and resumes its completed declined root without repeated effects 2010ms + ↓ tests/real-cli-adapters.test.ts (3 tests | 3 skipped) + ✓ tests/fs-descriptor.test.ts (1 test) 3ms + ✓ tests/bundle-preflight.test.ts (4 tests) 851ms + ✓ bundle execution preflight > ignores surrounding cache configuration on a verified cache hit 479ms + ✓ bundle execution preflight > uses the built alias for a nameless flow even in a digest-only cache directory 348ms + ✓ tests/parse-json-output.test.ts (7 tests) 2ms + ✓ tests/journal-client-completion.test.ts (4 tests) 98ms + ✓ tests/communication-preflight.test.ts (13 tests) 716ms + ✓ managed CLI preflight > does not loosen ordinary wrapper preflight or cache a managed result for an ordinary step 314ms + ✓ tests/communication-environment-preflight.test.ts (6 tests) 2ms + ✓ tests/memoization.test.ts (57 tests) 650ms + ✓ refuses invalid reuse invocation "run" 598ms + ✓ tests/budget-authored-live.test.ts (2 tests) 217ms + ✓ tests/slack-writeback.test.ts (1 test) 255ms + ✓ tests/adapters/claude.test.ts (7 tests) 2ms + ✓ tests/authored-surface-authority.test.ts (2 tests) 17ms + ✓ tests/adapters/codex.test.ts (7 tests) 4ms + ✓ tests/communication-history.test.ts (1 test) 2ms + ✓ tests/worker-cli-cwd.test.ts (2 tests) 279ms + ✓ tests/adapters/registry.test.ts (4 tests) 3ms + ✓ tests/slack-block-kit.test.ts (5 tests) 81ms + ✓ tests/promise-ancestry.test.ts (2 tests) 230ms + ✓ tests/authored-declined-report.test.ts (6 tests) 6ms + ✓ tests/classify-outcome.test.ts (2 tests) 2191ms + ✓ classifyOutcome > gives up and reports when a running run never becomes classifiable 2037ms + ✓ tests/communication-refusal.test.ts (1 test) 10ms + ✓ tests/catalog-plugins.test.ts (2 tests) 2ms + ✓ tests/check-command-cwd.test.ts (1 test) 12ms + ✓ tests/communication-lazy.test.ts (1 test) 2ms + ✓ tests/agent-cwd-validation.test.ts (2 tests) 319ms + ✓ declarative agent cwd > is refused by `flows check` on a YAML flow before anything runs 315ms + ↓ tests/run-digest-live.test.ts (1 test | 1 skipped) + ✓ tests/canonical-tree.test.ts (1 test) 2ms + ✓ tests/placement.test.ts (54 tests) 11ms + ✓ tests/communication-tools.test.ts (1 test) 81ms + ✓ tests/authored-admission.test.ts (2 tests) 3ms + ✓ tests/worker-cli-abort.test.ts (2 tests) 3196ms + ✓ stops claude and its process group when lease ownership is lost 1467ms + ✓ stops wrapper.mjs and its process group when lease ownership is lost 1728ms + ✓ tests/memory.test.ts (18 tests) 7ms + ✓ tests/worker-platform.test.ts (1 test) 3ms + ✓ tests/cli-progress-wait.test.ts (2 tests) 1360ms + ✓ run starts the wait clock on its first observed lease 698ms + ✓ resume starts the wait clock on its first observed lease 662ms + ✓ tests/bundle-transport.test.ts (20 tests) 2326ms + ✓ digest references > accepts and deploys the build output for hello 400ms + ✓ digest references > accepts and deploys the build output for Hello 307ms + ✓ digest references > accepts and deploys the build output for hello.world 337ms + ✓ digest references > accepts and deploys the build output for hello_world 627ms + ✓ digest references > accepts and deploys the build output for 123 379ms + ✓ tests/run-digest.test.ts (4 tests) 1434ms + ✓ digest run configuration refusals > reports config_invalid before fetching or starting a run for {invalid json 575ms + ✓ digest run configuration refusals > reports config_invalid before fetching or starting a run for {"deploy":{}} 349ms + ✓ tests/step-lease.test.ts (36 tests) 66605ms + ✓ f.run leases against the live kernel > enforces 10000 ms for 'sleep 5; printf ok' 5125ms + ✓ f.run leases against the live kernel > enforces 40000 ms for 'sleep 31; printf ok' 31151ms + ✓ f.run leases against the live kernel > enforces 30000 ms for 'sleep 31; printf ok' 30126ms + ✓ tests/local-agent-live.test.ts (5 tests) 74097ms + ✓ built CLI local agent against a real daemon > dispatches through the wrapper and keeps --json stdout report-shaped 2171ms + ✓ built CLI local agent against a real daemon > runs beyond the initial 30-second lease without a second invocation 37175ms + ✓ built CLI local agent against a real daemon > renders actual agent completion in text output 5424ms + ✓ built CLI local agent against a real daemon > returns a failed run when the agent process fails 14857ms + ✓ built CLI local agent against a real daemon > refuses a workspace it cannot pin before invoking the agent 14470ms + +⎯⎯⎯⎯⎯⎯ Failed Suites 3 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/authored-node-runtime.test.ts [ tests/authored-node-runtime.test.ts ] +AssertionError: expected '1.3.14' to be '1.4.0' // Object.is equality + +Expected: "1.4.0" +Received: "1.3.14" + + ❯ tests/authored-node-runtime.test.ts:18:77 + 16| + 17| beforeAll(() => { + 18| expect(spawnSync(bun, ['--version'], { encoding: 'utf8' }).stdout.tr… + | ^ + 19| expect(existsSync(daemon), 'build the current kernel or set RELAYFLO… + 20| stage = mkdtempSync(join(tmpdir(), 'authored-standalone-build-')); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/64]⎯ + + FAIL tests/babysitter-native-extension.test.ts [ tests/babysitter-native-extension.test.ts ] +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ baseAt src/hosted-extension-runtime.ts:234:20 + ❯ Module.loadHostedExtensionRuntime src/hosted-extension-runtime.ts:90:16 + ❯ composed tests/babysitter-native-extension.test.ts:64:25 + ❯ tests/babysitter-native-extension.test.ts:100:37 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[2/64]⎯ + + FAIL tests/mcp.test.ts > authored MCP effects against the real kernel +Error: journal client: connect failed: connect ENOENT /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-649148efd2a9.sock + ❯ Socket.onError src/journal-client.ts:102:16 + 100| socket.removeAllListeners(); + 101| this.failAll(err); + 102| reject(new Error(`journal client: connect failed: ${err.messag… + | ^ + 103| }; + 104| socket.once('error', onError); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[3/64]⎯ + +⎯⎯⎯⎯⎯⎯ Failed Tests 61 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > opens the actual catalog flow with the ticket title and exactly one GitHub closing line +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ runCanonical tests/canonical-software-factory.test.ts:69:22 + 67| }; + 68| + 69| const definition = getFlowDefinition(softwareFactory); + | ^ + 70| return definition.body(context as never, { issue, approver: 'khaliq'… + 71| root, + ❯ tests/canonical-software-factory.test.ts:82:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[4/64]⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > normalizes whitespace and caps the title at 240 Unicode code points +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ runCanonical tests/canonical-software-factory.test.ts:69:22 + 67| }; + 68| + 69| const definition = getFlowDefinition(softwareFactory); + | ^ + 70| return definition.body(context as never, { issue, approver: 'khaliq'… + 71| root, + ❯ tests/canonical-software-factory.test.ts:99:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[5/64]⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > fails closed before push when GitHub identity is missing or the final body duplicates its closing line +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ tests/canonical-software-factory.test.ts:108:24 + 106| + 107| it('fails closed before push when GitHub identity is missing or the … + 108| const definition = getFlowDefinition(softwareFactory); + | ^ + 109| const commands: string[] = []; + 110| let completionReason = ''; + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[6/64]⎯ + + FAIL tests/cli-watch.test.ts > flows check --watch > detects a nearer config appearing and falls back after it is deleted +Error: Test timed out in 5000ms. +If this is a long-running test, pass a timeout value as the last argument or configure it globally with "testTimeout". +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[7/64]⎯ + + FAIL tests/communication-mixed-resume.test.ts > resumes mixed ordinary and linked agents through the real daemon without stealing peer capacity +AssertionError: {"ok":false,"command":"resume","resolutions":[],"diagnostics":[{"severity":"failure","kind":"protocol_error","message":"relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`"}],"runId":"01M36BDG520KC244R6W7NNWDHX","socketPath":"/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-bcaf86b64a3c.sock"}: expected { ok: false, command: 'resume', …(4) } to match object { status: 'completed', …(2) } +(6 matching properties omitted from actual) + +- Expected ++ Received + + Object { +- "completedSteps": 4, +- "completionReason": "success", +- "status": "completed", ++ "command": "resume", ++ "diagnostics": Array [ ++ Object { ++ "kind": "protocol_error", ++ "message": "relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`", ++ "severity": "failure", ++ }, ++ ], ++ "ok": false, ++ "resolutions": Array [], ++ "runId": "01M36BDG520KC244R6W7NNWDHX", ++ "socketPath": "/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-bcaf86b64a3c.sock", + } + + ❯ tests/communication-mixed-resume.test.ts:48:60 + 46| expect(parked.status).toBe('parked'); + 47| const resumed = await resumeFlow(parked.run_id, dataDir, { localAg… + 48| expect(resumed.report, JSON.stringify(resumed.report)).toMatchObje… + | ^ + 49| expect(state.started).toEqual(new Set(['a', 'b'])); + 50| const entries = (await client.journalRead(parked.run_id, 1)).entri… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[8/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > shadows inherited thenables on completed snapshot authority +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:33:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[9/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > locks the inherited then slot before authored code can schedule a replacement +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:58:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[10/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps the buffered base bytes when the live source changes +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:65:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[11/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps source digests sensitive when ambient Array.map is poisoned +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:78:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[12/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > defines source entries without consulting inherited numeric setters +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:118:75 + 116| }, + 117| }); + 118| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 119| } finally { + 120| if (previous === undefined) delete (Array.prototype as unknown a… + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:118:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[13/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > uses the module-captured platform during source traversal +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:142:75 + 140| }, + 141| }); + 142| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 143| } finally { + 144| Object.defineProperty(process, 'platform', descriptor); + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:142:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[14/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > shadows native directory arrays before promise resolution can substitute them +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:168:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[15/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > excludes project node_modules from the admitted generation +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:184:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[16/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > refuses an oversized source file before buffering its contents +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "snapshot entry or byte limit", + } + + ❯ tests/hosted-base-snapshot.test.ts:200:5 + 198| writeFileSync(oversized, ''); + 199| truncateSync(oversized, 64 * 1024 * 1024 + 1); + 200| await expect(createHostedBaseSnapshot(flowPath)).rejects.toMatchOb… + | ^ + 201| code: 'plugin_source_invalid', + 202| message: expect.stringContaining('snapshot entry or byte limit'), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[17/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > captures conversion and allocation across declarations and source snapshots +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "reviewed Software Factory base source", + } + + ❯ tests/hosted-base-snapshot.test.ts:247:7 + 245| return allocUnsafe(size); + 246| }) as typeof Buffer.allocUnsafe; + 247| await expect(loadHostedExtensionRuntime(flowPath)).rejects.toMat… + | ^ + 248| code: 'plugin_source_invalid', + 249| message: expect.stringContaining('reviewed Software Factory ba… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[18/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps declaration and reviewed-base reads bound after builtin export synchronization +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "reviewed Software Factory base source", + } + + ❯ tests/hosted-base-snapshot.test.ts:321:7 + 319| } + 320| syncBuiltinESMExports(); + 321| await expect(loadHostedExtensionRuntime(flowPath)).rejects.toMat… + | ^ + 322| code: 'plugin_source_invalid', + 323| message: expect.stringContaining('reviewed Software Factory ba… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[19/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > reads through captured descriptors and charges admitted descriptor sizes +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:380:75 + 378| }, + 379| }); + 380| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 381| } finally { + 382| fileHandlePrototype.read = fileHandleRead; + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:380:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[20/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > ignores inherited snapshot test hooks when production omits them +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:403:75 + 401| }); + 402| } + 403| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 404| } finally { + 405| for (let index = 0; index < names.length; index += 1) { + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:403:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[21/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > bounds a file that grows after its admitted size was checked +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "changed while reading \"raced.bin\"", + } + + ❯ tests/hosted-base-snapshot.test.ts:418:5 + 416| const raced = join(project, 'raced.bin'); + 417| writeFileSync(raced, 'small'); + 418| await expect( + | ^ + 419| hostedBaseSourceDigest([{ root: project, prefix: '' }], { + 420| afterStat: async path => { + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[22/64]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > stops streaming project entries at the shared count limit +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "snapshot entry or byte limit", + } + + ❯ tests/hosted-base-snapshot.test.ts:467:5 + 465| writeFileSync(join(project, `empty-${index}.txt`), ''); + 466| } + 467| await expect(createHostedBaseSnapshot(flowPath)).rejects.toMatchOb… + | ^ + 468| code: 'plugin_source_invalid', + 469| message: expect.stringContaining('snapshot entry or byte limit'), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[23/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > resolves locked artifacts without importing extension top-level code +AssertionError: expected [ { …(6) } ] to deeply equal [ { …(6) } ] + +- Expected ++ Received + + Array [ + Object { + "digest": "ada877550e2bc84fb8050d89d40580eeae0fa9425b078bee5e6b02568a561925", +- "directory": "/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/hosted-extension-test-555Up2/.flows/plugins/babysitter@sha256:ada877550e2bc84fb8050d89d40580eeae0fa9425b078bee5e6b02568a561925", ++ "directory": "/private/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/hosted-extension-test-555Up2/.flows/plugins/babysitter@sha256:ada877550e2bc84fb8050d89d40580eeae0fa9425b078bee5e6b02568a561925", + "manifestSha256": "5681622ac4930a1b9c7eda1736b52111e7ae8711c0a29eed0db00bb639fc7375", + "name": "babysitter", + "ref": "github:AgentWorkforce/flows@aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa#extensions/babysitter", + "version": "0.2.0", + }, + ] + + ❯ tests/hosted-extension-isolation.test.ts:159:70 + 157| const flowPath = join(root, 'software-factory.flow.ts'); + 158| writeFileSync(flowPath, 'export default {};'); + 159| expect((await loadHostedExtensionArtifacts(flowPath)).artifacts).t… + | ^ + 160| expect(() => readFileSync(marker)).toThrow(); + 161| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[24/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > executes the exact capability-only handler for a queued receipt + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > executes the exact capability-only handler for a duplicate receipt +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:221:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[25/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > streams verified bytes when the live store is replaced and no writable staging path exists +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:431:7 + 429| return { receiptId: 'receipt-1', status: 'queued' }; + 430| } }, + 431| })).resolves.toEqual({ completionReason: 'success', capabilityCall… + | ^ + 432| expect(calls).toBe(1); + 433| expect(readFileSync(join(installed.directory, 'babysitter.flow.ts'… + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:424:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[26/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > mounts pinned private Surface bytes when the live package changes before launch +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:456:7 + 454| return { receiptId: 'receipt-1', status: 'queued' }; + 455| } }, + 456| })).resolves.toEqual({ completionReason: 'success', capabilityCall… + | ^ + 457| expect(calls).toBe(1); + 458| expect(readFileSync(join(surfaceRoot, 'dist/flow.js'), 'utf8')).to… + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:443:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[27/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > refuses oversized Surface files through the bounded descriptor reader +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_unsupported", +- "message": StringContaining "cannot read pinned Surface runtime flow.js", + } + + ❯ tests/hosted-extension-isolation.test.ts:489:5 + 487| }); + 488| let calls = 0; + 489| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 490| artifact: await artifact(), + 491| manifest: validateFlowExtensionManifest(manifest()), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[28/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > shields verified Surface files before async settlement +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:532:9 + 530| surfaceRoot, + 531| babysitterTurn: { queue: async () => ({ receiptId: 'receipt-1'… + 532| })).resolves.toEqual({ completionReason: 'success', capabilityCa… + | ^ + 533| } finally { + 534| if (previous === undefined) delete (Array.prototype as { then?: … + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:525:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[29/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > preserves a typed host refusal while disclosing only a fixed marker to the child +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to be Error: private Cloud policy detail { code: '…' } // Object.is equality + +- Expected ++ Received + +- [Error: private Cloud policy detail] ++ [Error: hosted extension isolation requires Linux] + + ❯ tests/hosted-extension-isolation.test.ts:560:5 + 558| provider: 'github', eventType: 'pull_request.labeled', deliveryI… + 559| }); + 560| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 561| artifact: await artifact(source), manifest: validateFlowExtensio… + 562| babysitterTurn: { queue: async () => { throw refusal; } }, + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[30/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > denies ambient credentials, host files, writes, network, subprocesses, and undeclared context verbs +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:624:13 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[31/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > enforces OS address-space and data bounds on native Buffer allocation +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_unsupported", +- "message": StringMatching /(?:Failed to allocate memory|Array buffer allocation failed)/u, + } + + ❯ tests/hosted-extension-isolation.test.ts:646:5 + 644| }); + 645| let calls = 0; + 646| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 647| artifact: installed, + 648| manifest: validateFlowExtensionManifest(manifest()), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[32/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > blocks extra handler fields and authority-bearing receipt fields at the parent port +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + +- Expected ++ Received + +- Object { +- "code": "plugin_event_unroutable", ++ PluginError { ++ "code": "plugin_unsupported", + } + + ❯ tests/hosted-extension-isolation.test.ts:675:5 + 673| }); + 674| let calls = 0; + 675| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 676| artifact: await artifact(source), manifest: validateFlowExtensio… + 677| babysitterTurn: { queue: async () => { calls += 1; return { rece… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[33/64]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > constructs adapter authority with the captured freeze intrinsic +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:745:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[34/64]⎯ + + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different PR frame with zero adapter calls + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different delivery frame with zero adapter calls + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different event frame with zero adapter calls +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + +- Expected ++ Received + +- Object { +- "code": "plugin_event_unroutable", ++ PluginError { ++ "code": "plugin_unsupported", + } + + ❯ tests/hosted-extension-protocol.test.ts:430:5 + 428| ])('rejects an import-time %s frame with zero adapter calls', async … + 429| let calls = 0; + 430| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 431| artifact: await artifact(hostileImport([frame, { type: 'error', … + 432| manifest: validateFlowExtensionManifest(manifest()), dispatch: d… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[35/64]⎯ + + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects two forged calls after the authoritative first outcome settles + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > waits for a pending adapter to reject after a forged child error + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > waits for a pending adapter to resolve after a forged child error + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > returns a typed adapter rejection even when the hostile child hangs +Error: hostile child did not invoke the adapter + ❯ Timeout._onTimeout tests/hosted-extension-protocol.test.ts:118:45 + 116| async function waitForInvocation(invoked: Promise): Promise((resolve, reject) => { + 118| const timeout = setTimeout(() => reject(new Error('hostile child d… + | ^ + 119| void invoked.then(() => { clearTimeout(timeout); resolve(); }, rej… + 120| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[36/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses an 8-character run-id prefix: Cloud has no prefix lookup +AssertionError: expected [Function] to throw error matching /not full Cloud run ids: c649fe14/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/not full Cloud run ids: c649fe14/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[37/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses the whole batch when any id is invalid, rather than dropping it +AssertionError: expected [Function] to throw error matching /not full Cloud run ids: nope!/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/not full Cloud run ids: nope!/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[38/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses an empty batch +AssertionError: expected [Function] to throw error matching /needs runIds/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/needs runIds/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[39/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses a batch too large for the edge step lease +AssertionError: expected [Function] to throw error matching /exceeds the 8 that fit/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/exceeds the 8 that fit/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[40/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > accepts eight ids — the incident batch is inside the bound +AssertionError: promise rejected "TypeError: expected an @relayflows/surfac…" instead of resolving + ❯ tests/stuck-run-triage.test.ts:62:40 + 60| it('accepts eight ids — the incident batch is inside the bound', asy… + 61| const ids = Array.from({ length: 8 }, (_, i) => `${ID_A.slice(0, -… + 62| await expect(drive({ runIds: ids })).resolves.toBeDefined(); + | ^ + 63| }); + 64| }); + +Caused by: TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + ❯ tests/stuck-run-triage.test.ts:62:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[41/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > refuses to send the Cloud bearer token to an unapproved origin +AssertionError: expected [Function] to throw error matching /refusing to send the Cloud bearer to…/\ but got 'expected an @relayflows/surface flow …' + +- Expected: +/refusing to send the Cloud bearer token to https:\/\/evil\.example/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[42/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > refuses a non-URL apiUrl +AssertionError: expected [Function] to throw error matching /is not a URL/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/is not a URL/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[43/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > allows an approved origin and uses it in the curl +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:77:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[44/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > defaults to production Cloud +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:82:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[45/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > never publishes a run record the fetch did not produce +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:89:34 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[46/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > names the Worker on every wrangler invocation +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:98:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[47/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > accepts caller-supplied Workers and rejects option-shaped ones +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:107:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[48/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > falls back when GNU timeout is absent, as it is on macOS +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:115:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[49/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > runs the tails concurrently so wall time does not scale with the batch +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:123:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[50/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > records wrangler's own exit status rather than head's +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:129:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[51/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage shell text > parses under both sh and bash +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:137:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[52/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage shell text > collects tails with no GNU timeout on PATH, as on a stock macOS +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:157:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[53/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage agents > declares read-only permissions on every agent +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:176:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[54/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage agents > tells the forensics agents their evidence is untrusted +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:182:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[55/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > refuses a duplicate run id: two tails would share one evidence file +AssertionError: expected [Function] to throw error matching /duplicate runIds: c649fe14-0c2e-4e51-…/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/duplicate runIds: c649fe14-0c2e-4e51-9a6a-4f0d1b0f77aa/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[56/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > refuses a duplicate Worker name for the same reason +AssertionError: expected [Function] to throw error matching /duplicate workers: w-one/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/duplicate workers: w-one/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[57/64]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > bounds ids x workers, not just ids +AssertionError: expected [Function] to throw error matching /24 concurrent tails, over the 16/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/24 concurrent tails, over the 16/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[58/64]⎯ + +⎯⎯⎯⎯⎯⎯ Unhandled Errors ⎯⎯⎯⎯⎯⎯ + +Vitest caught 2 unhandled errors during the test run. +This might cause false positive tests. Resolve unhandled errors to make sure your tests are not affected. + +⎯⎯⎯⎯ Unhandled Rejection ⎯⎯⎯⎯⎯ +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-protocol.test.ts:454:17 + ❯ node_modules/@vitest/runner/dist/index.js:533:5 + ❯ runTest node_modules/@vitest/runner/dist/index.js:1056:11 + ❯ runSuite node_modules/@vitest/runner/dist/index.js:1205:15 + ❯ runSuite node_modules/@vitest/runner/dist/index.js:1205:15 + ❯ runFiles node_modules/@vitest/runner/dist/index.js:1262:5 + ❯ startTests node_modules/@vitest/runner/dist/index.js:1271:3 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +This error originated in "tests/hosted-extension-protocol.test.ts" test file. It doesn't mean the error was thrown inside the file itself, but while it was running. +The latest test that might've caused the error is "rejects two forged calls after the authoritative first outcome settles". It might mean one of the following: +- The error was thrown, while Vitest was running this test. +- If the error occurred after the test had been completed, this was the last documented test before it was thrown. + +⎯⎯⎯⎯⎯ Uncaught Exception ⎯⎯⎯⎯⎯ +Error: spawn /Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd ENOENT + ❯ Process.ChildProcess._handle.onexit node:internal/child_process:285:19 + ❯ onErrorNT node:internal/child_process:483:16 + ❯ processTicksAndRejections node:internal/process/task_queues:89:21 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { errno: -2, code: 'ENOENT', syscall: 'spawn /Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd', path: '/Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd', spawnargs: [ '--data-dir', '/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/flows-mcp-daemon-wnknsm', 'serve' ] } +This error originated in "tests/mcp.test.ts" test file. It doesn't mean the error was thrown inside the file itself, but while it was running. +The latest test that might've caused the error is "authored MCP effects against the real kernel". It might mean one of the following: +- The error was thrown, while Vitest was running this test. +- If the error occurred after the test had been completed, this was the last documented test before it was thrown. +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ + + Test Files 10 failed | 169 passed | 3 skipped (182) + Tests 61 failed | 2759 passed | 72 skipped (2892) + Errors 2 errors + Start at 22:20:28 + Duration 152.83s (transform 2.97s, setup 0ms, collect 51.96s, tests 706.78s, environment 24ms, prepare 7.70s) + +exit=1 diff --git a/evidence/llm-fence-predicate-resume/mutation-1-fence.txt b/evidence/llm-fence-predicate-resume/mutation-1-fence.txt new file mode 100644 index 000000000..c66633d08 --- /dev/null +++ b/evidence/llm-fence-predicate-resume/mutation-1-fence.txt @@ -0,0 +1,58 @@ +$ cp src/llm-worker.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts +exit=0 +$ sed -i.bak 's/JSON.parse(unfenced(result.stdout_tail))/JSON.parse(result.stdout_tail)/' src/llm-worker.ts && rm src/llm-worker.ts.bak +exit=0 +$ git diff --stat -- src/llm-worker.ts + packages/sdk/src/llm-worker.ts | 2 +- + 1 file changed, 1 insertion(+), 1 deletion(-) +exit=0 +$ npx vitest run tests/worker-transcript.test.ts -t 'fenced reply' + + RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk + + ❯ tests/worker-transcript.test.ts (8 tests | 1 failed | 7 skipped) 205ms + × the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value 205ms + → expected 'verification_failed' to be 'success' // Object.is equality + +⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/worker-transcript.test.ts > the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value +AssertionError: expected 'verification_failed' to be 'success' // Object.is equality + +Expected: "success" +Received: "verification_failed" + + ❯ tests/worker-transcript.test.ts:215:27 + 213| it('accepts a fenced reply whose content matches the schema, and jou… + 214| const completion = await completeWith('```json\n{"x": 4}\n```'); + 215| expect(completion[4]).toBe('success'); + | ^ + 216| expect((completion[5] as { output: unknown }).output).toEqual({ x:… + 217| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/1]⎯ + + Test Files 1 failed (1) + Tests 1 failed | 7 skipped (8) + Start at 22:18:46 + Duration 1.36s (transform 256ms, setup 0ms, collect 901ms, tests 205ms, environment 0ms, prepare 49ms) + +exit=1 +$ cp /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts src/llm-worker.ts +exit=0 +$ cmp src/llm-worker.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts +exit=0 +$ git diff --stat -- src/llm-worker.ts +exit=0 +$ npx vitest run tests/worker-transcript.test.ts -t 'fenced reply' + + RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk + + ✓ tests/worker-transcript.test.ts (8 tests | 7 skipped) 195ms + + Test Files 1 passed (1) + Tests 1 passed | 7 skipped (8) + Start at 22:18:48 + Duration 901ms (transform 206ms, setup 0ms, collect 461ms, tests 195ms, environment 0ms, prepare 43ms) + +exit=0 diff --git a/evidence/llm-fence-predicate-resume/mutation-2-predicate.txt b/evidence/llm-fence-predicate-resume/mutation-2-predicate.txt new file mode 100644 index 000000000..e0161f0df --- /dev/null +++ b/evidence/llm-fence-predicate-resume/mutation-2-predicate.txt @@ -0,0 +1,58 @@ +$ cp src/authored-flow-executor.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts +exit=0 +$ sed -i.bak 's/JSON.stringify(canonical)/JSON.stringify(record)/' src/authored-flow-executor.ts && rm src/authored-flow-executor.ts.bak +exit=0 +$ git diff --stat -- src/authored-flow-executor.ts + packages/sdk/src/authored-flow-executor.ts | 2 +- + 1 file changed, 1 insertion(+), 1 deletion(-) +exit=0 +$ npx vitest run tests/authored-agent-artifacts.test.ts -t 'sorted keys' + + RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk + + ❯ tests/authored-agent-artifacts.test.ts (5 tests | 1 failed | 4 skipped) 42ms + × predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run 41ms + → expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality + +⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/authored-agent-artifacts.test.ts > predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run +AssertionError: expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality + +Expected: "printf '%s' '{"gate":"predicate","step":"run-1","verdict":"pass","because":"plan covers every question"}'" +Received: "printf '%s' '{"because":"plan covers every question","gate":"predicate","step":"run-1","verdict":"pass"}'" + + ❯ tests/authored-agent-artifacts.test.ts:300:25 + 298| expect(stream).toHaveLength(1); + 299| expect(commands).toHaveLength(2); + 300| expect(commands[1]).toBe(commands[0]); + | ^ + 301| expect(commands[0]).toContain('{"gate":"predicate","step":"run-1",… + 302| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/1]⎯ + + Test Files 1 failed (1) + Tests 1 failed | 4 skipped (5) + Start at 22:18:50 + Duration 1.56s (transform 573ms, setup 0ms, collect 1.26s, tests 42ms, environment 0ms, prepare 66ms) + +exit=1 +$ cp /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts src/authored-flow-executor.ts +exit=0 +$ cmp src/authored-flow-executor.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts +exit=0 +$ git diff --stat -- src/authored-flow-executor.ts +exit=0 +$ npx vitest run tests/authored-agent-artifacts.test.ts -t 'sorted keys' + + RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk + + ✓ tests/authored-agent-artifacts.test.ts (5 tests | 4 skipped) 33ms + + Test Files 1 passed (1) + Tests 1 passed | 4 skipped (5) + Start at 22:18:52 + Duration 982ms (transform 371ms, setup 0ms, collect 736ms, tests 33ms, environment 0ms, prepare 57ms) + +exit=0 diff --git a/evidence/llm-fence-predicate-resume/thirteen-files-at-main.txt b/evidence/llm-fence-predicate-resume/thirteen-files-at-main.txt new file mode 100644 index 000000000..23c162ce7 --- /dev/null +++ b/evidence/llm-fence-predicate-resume/thirteen-files-at-main.txt @@ -0,0 +1,1209 @@ +$ cp src/llm-worker.ts src/authored-flow-executor.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/ +exit=0 +$ git show main:packages/sdk/src/llm-worker.ts > src/llm-worker.ts +exit=0 +$ git show main:packages/sdk/src/authored-flow-executor.ts > src/authored-flow-executor.ts +exit=0 +$ git diff --stat main -- src/llm-worker.ts src/authored-flow-executor.ts +exit=0 +$ npx vitest run tests/authored-node-runtime.test.ts tests/babysitter-native-extension.test.ts tests/bundle.test.ts tests/canonical-software-factory.test.ts tests/cli-watch.test.ts tests/communication-mixed-resume.test.ts tests/hosted-base-snapshot.test.ts tests/hosted-extension-isolation.test.ts tests/hosted-extension-protocol.test.ts tests/mcp.test.ts tests/stop-process-group.test.ts tests/stuck-run-triage.test.ts tests/worker-cli.test.ts + + RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk + + ❯ tests/hosted-extension-isolation.test.ts (22 tests | 12 failed | 3 skipped) 237ms + × hosted extension capability isolation > resolves locked artifacts without importing extension top-level code 25ms + → expected [ { …(6) } ] to deeply equal [ { …(6) } ] + × hosted extension capability isolation > executes the exact capability-only handler for a queued receipt 8ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > executes the exact capability-only handler for a duplicate receipt 8ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > streams verified bytes when the live store is replaced and no writable staging path exists 72ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > mounts pinned private Surface bytes when the live package changes before launch 18ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > refuses oversized Surface files through the bounded descriptor reader 9ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + × hosted extension capability isolation > shields verified Surface files before async settlement 7ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > preserves a typed host refusal while disclosing only a fixed marker to the child 11ms + → expected Error: hosted extension isolation require… { code: '…' } to be Error: private Cloud policy detail { code: '…' } // Object.is equality + × hosted extension capability isolation > denies ambient credentials, host files, writes, network, subprocesses, and undeclared context verbs 8ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > enforces OS address-space and data bounds on native Buffer allocation 11ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + × hosted extension capability isolation > blocks extra handler fields and authority-bearing receipt fields at the parent port 4ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension capability isolation > constructs adapter authority with the captured freeze intrinsic 3ms + → hosted extension isolation requires Linux + ❯ tests/hosted-base-snapshot.test.ts (18 tests | 15 failed | 1 skipped) 1345ms + × hosted base private snapshot > shadows inherited thenables on completed snapshot authority 14ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > locks the inherited then slot before authored code can schedule a replacement 4ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > keeps the buffered base bytes when the live source changes 3ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > keeps source digests sensitive when ambient Array.map is poisoned 1ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > defines source entries without consulting inherited numeric setters 8ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > uses the module-captured platform during source traversal 2ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > shadows native directory arrays before promise resolution can substitute them 2ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > excludes project node_modules from the admitted generation 2ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > refuses an oversized source file before buffering its contents 4ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > captures conversion and allocation across declarations and source snapshots 4ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > keeps declaration and reviewed-base reads bound after builtin export synchronization 4ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > reads through captured descriptors and charges admitted descriptor sizes 3ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > ignores inherited snapshot test hooks when production omits them 2ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > bounds a file that grows after its admitted size was checked 1ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > stops streaming project entries at the shared count limit 1288ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + ❯ tests/babysitter-native-extension.test.ts (41 tests | 41 skipped) 740ms + ❯ tests/authored-node-runtime.test.ts (14 tests | 14 skipped) 63ms + ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) 168ms + × stuck-run-triage input validation > refuses an 8-character run-id prefix: Cloud has no prefix lookup 5ms + → expected [Function] to throw error matching /not full Cloud run ids: c649fe14/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses the whole batch when any id is invalid, rather than dropping it 1ms + → expected [Function] to throw error matching /not full Cloud run ids: nope!/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses an empty batch 0ms + → expected [Function] to throw error matching /needs runIds/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses a batch too large for the edge step lease 0ms + → expected [Function] to throw error matching /exceeds the 8 that fit/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > accepts eight ids — the incident batch is inside the bound 3ms + → promise rejected "TypeError: expected an @relayflows/surfac…" instead of resolving + × stuck-run-triage apiUrl > refuses to send the Cloud bearer token to an unapproved origin 0ms + → expected [Function] to throw error matching /refusing to send the Cloud bearer to…/\ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage apiUrl > refuses a non-URL apiUrl 0ms + → expected [Function] to throw error matching /is not a URL/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage apiUrl > allows an approved origin and uses it in the curl 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage apiUrl > defaults to production Cloud 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage apiUrl > never publishes a run record the fetch did not produce 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > names the Worker on every wrangler invocation 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > accepts caller-supplied Workers and rejects option-shaped ones 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > falls back when GNU timeout is absent, as it is on macOS 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > runs the tails concurrently so wall time does not scale with the batch 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > records wrangler's own exit status rather than head's 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage shell text > parses under both sh and bash 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage shell text > collects tails with no GNU timeout on PATH, as on a stock macOS 154ms + → expected an @relayflows/surface flow handle + × stuck-run-triage agents > declares read-only permissions on every agent 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage agents > tells the forensics agents their evidence is untrusted 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage fan-out > refuses a duplicate run id: two tails would share one evidence file 0ms + → expected [Function] to throw error matching /duplicate runIds: c649fe14-0c2e-4e51-…/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage fan-out > refuses a duplicate Worker name for the same reason 0ms + → expected [Function] to throw error matching /duplicate workers: w-one/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage fan-out > bounds ids x workers, not just ids 0ms + → expected [Function] to throw error matching /24 concurrent tails, over the 16/ but got 'expected an @relayflows/surface flow …' + ❯ tests/canonical-software-factory.test.ts (3 tests | 3 failed) 66ms + × canonical software-factory metadata contract > opens the actual catalog flow with the ticket title and exactly one GitHub closing line 41ms + → expected an @relayflows/surface flow handle + × canonical software-factory metadata contract > normalizes whitespace and caps the title at 240 Unicode code points 24ms + → expected an @relayflows/surface flow handle + × canonical software-factory metadata contract > fails closed before push when GitHub identity is missing or the final body duplicates its closing line 1ms + → expected an @relayflows/surface flow handle + ❯ tests/communication-mixed-resume.test.ts (1 test | 1 failed) 90ms + × resumes mixed ordinary and linked agents through the real daemon without stealing peer capacity 90ms + → {"ok":false,"command":"resume","resolutions":[],"diagnostics":[{"severity":"failure","kind":"protocol_error","message":"relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`"}],"runId":"01M36B7SMVPW00H67FE03NC6XS","socketPath":"/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-c074cf92d10f.sock"}: expected { ok: false, command: 'resume', …(4) } to match object { status: 'completed', …(2) } +(6 matching properties omitted from actual) + ✓ tests/bundle.test.ts (26 tests) 7669ms + ✓ immutable bundles > returns exit 2 naming a byte-flipped payload and refuses to reuse corruption 364ms + ✓ immutable bundles > verifies with --verify in any position and answers --json with one object 690ms + ✓ immutable bundles > refuses --out with --verify rather than ignoring the destination 323ms + ✓ immutable bundles > builds and verifies the canonical YAML fixture through the compiled CLI 757ms + ✓ immutable bundles > emits the ephemeral warning on CLI stderr and uses the default output directory 495ms + ✓ immutable bundles > builds a standalone TS fixture twice with identical executable hashes 2896ms + ❯ tests/mcp.test.ts (30 tests | 4 skipped) 9774ms + ✓ MCP preflight and transports > flows check refuses an undeclared server with exit 2 and no daemon 848ms + ✓ MCP preflight and transports > flows check reports a refusing server and leaves no PID 773ms + ✓ MCP preflight and transports > kills a SIGTERM-resistant silent child after a parent-owned handshake deadline 1331ms + ✓ MCP preflight and transports > reaps a SIGTERM-resistant descendant with inherit stdio before cleanup finishes 1142ms + ✓ MCP preflight and transports > reaps a SIGTERM-resistant descendant with ignore stdio before cleanup finishes 2074ms + ✓ MCP preflight and transports > reports malformed connection configuration as config_invalid 429ms + ✓ tests/cli-watch.test.ts (10 tests) 12241ms + ✓ flows check --watch > rechecks syntax errors, clears once, and returns the last refusal on Ctrl-C 1168ms + ✓ flows check --watch > streams JSON lines without ANSI, recovers after atomic saves, and exits zero after repair 1351ms + ✓ flows check --watch > coalesces 20 concurrent saves into at most two rechecks 1476ms + ✓ flows check --watch > watches transitive relative use imports, cycles, and nearest config changes 1827ms + ✓ flows check --watch > refreshes the import graph and notices missing imports being created 1818ms + ✓ flows check --watch > reloads authored TypeScript instead of reusing the first imported definition 1034ms + ✓ flows check --watch > detects a nearer config appearing and falls back after it is deleted 1317ms + ✓ flows check --watch > keeps watching after the target is deleted and recreated 1481ms + ✓ flows check --watch > queues changes during a slow check without overlapping checks 767ms + ✓ tests/stop-process-group.test.ts (9 tests) 15421ms + ✓ every stop reaches the process group, not just the direct child > exits the run after an execution-timeout stop 2203ms + ✓ every stop reaches the process group, not just the direct child > exits the run after a protocol terminate stop 668ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after a protocol terminate stop 1837ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after an execution-timeout stop 2389ms + ✓ every stop reaches the process group, not just the direct child > holds the loop open long enough for the escalation to run 1107ms + ✓ a wrapper that exits with no execution deadline still drains > reports the wrapper result and reaps a grandchild holding its pipes 811ms + ✓ a wrapper that exits with no execution deadline still drains > reaps a SIGTERM-deaf grandchild holding its pipes 1742ms + ✓ a wrapper that exits with no execution deadline still drains > settles on its own deadline when an escaped holder withholds close 4381ms + ✓ tests/worker-cli.test.ts (18 tests) 25934ms + ✓ registered CLI model defaults > passes the same priced Claude default to the real provider invocation 835ms + ✓ step discovery environment > names the run, step, attempt and an absolute data dir for a direct agent spawn 591ms + ✓ step discovery environment > exports none of the four without a data dir, even when the worker inherited them 503ms + ✓ wrapper discovery environment > sets the four names from the dispatch and still refuses ambient values and other secrets 549ms + ✓ wrapper discovery environment > exports none of the four to a wrapper without a data dir, even when the worker inherited them 378ms + ✓ custom wrapper execution identity > passes an explicit safe environment at identification and execution 392ms + ✓ custom wrapper execution identity > refuses a wrapper symlink retarget before delivering private values 369ms + ✓ custom wrapper execution identity > bounds wrapper execution after acknowledgement 405ms + ✓ custom wrapper execution identity > bounds captured wrapper output 435ms + ✓ custom wrapper execution identity > refuses a duplicate execute protocol frame 425ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a conforming wrapper leaks a stdio pipe to a background helper 1928ms + ✓ custom wrapper execution bounds are reader-owned > resolves when the leaked helper inherits stderr only 1956ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a wrapper leaks a stdio pipe and exits before identifying 3432ms + ✓ custom wrapper execution bounds are reader-owned > journals a completionReason at the default bound when a wrapper leaks a stdio pipe 11479ms + ✓ custom wrapper execution bounds are reader-owned > accepts an execute token and an over-8KiB payload flushed in one write 351ms + ✓ custom wrapper execution bounds are reader-owned > accepts the same over-8KiB payload whether or not it coalesces with the execute token 1194ms + ✓ custom wrapper execution bounds are reader-owned > still bounds an un-terminated handshake buffer and names the bound 357ms + ✓ delivers the journaled memory pack to the real wrapper and excludes its charge from completion usage 355ms + ❯ tests/hosted-extension-protocol.test.ts (24 tests | 7 failed | 1 skipped) 40154ms + × hosted extension hostile protocol > rejects an import-time different PR frame with zero adapter calls 14ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects an import-time different delivery frame with zero adapter calls 10ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects an import-time different event frame with zero adapter calls 17ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects two forged calls after the authoritative first outcome settles 10008ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > waits for a pending adapter to reject after a forged child error 10004ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > waits for a pending adapter to resolve after a forged child error 10005ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > returns a typed adapter rejection even when the hostile child hangs 10009ms + → hostile child did not invoke the adapter + +⎯⎯⎯⎯⎯⎯ Failed Suites 3 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/authored-node-runtime.test.ts [ tests/authored-node-runtime.test.ts ] +AssertionError: expected '1.3.14' to be '1.4.0' // Object.is equality + +Expected: "1.4.0" +Received: "1.3.14" + + ❯ tests/authored-node-runtime.test.ts:18:77 + 16| + 17| beforeAll(() => { + 18| expect(spawnSync(bun, ['--version'], { encoding: 'utf8' }).stdout.tr… + | ^ + 19| expect(existsSync(daemon), 'build the current kernel or set RELAYFLO… + 20| stage = mkdtempSync(join(tmpdir(), 'authored-standalone-build-')); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/63]⎯ + + FAIL tests/babysitter-native-extension.test.ts [ tests/babysitter-native-extension.test.ts ] +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ baseAt src/hosted-extension-runtime.ts:234:20 + ❯ Module.loadHostedExtensionRuntime src/hosted-extension-runtime.ts:90:16 + ❯ composed tests/babysitter-native-extension.test.ts:64:25 + ❯ tests/babysitter-native-extension.test.ts:100:37 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[2/63]⎯ + + FAIL tests/mcp.test.ts > authored MCP effects against the real kernel +Error: journal client: connect failed: connect ENOENT /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-34682962b0a5.sock + ❯ Socket.onError src/journal-client.ts:102:16 + 100| socket.removeAllListeners(); + 101| this.failAll(err); + 102| reject(new Error(`journal client: connect failed: ${err.messag… + | ^ + 103| }; + 104| socket.once('error', onError); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[3/63]⎯ + +⎯⎯⎯⎯⎯⎯ Failed Tests 60 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > opens the actual catalog flow with the ticket title and exactly one GitHub closing line +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ runCanonical tests/canonical-software-factory.test.ts:69:22 + 67| }; + 68| + 69| const definition = getFlowDefinition(softwareFactory); + | ^ + 70| return definition.body(context as never, { issue, approver: 'khaliq'… + 71| root, + ❯ tests/canonical-software-factory.test.ts:82:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[4/63]⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > normalizes whitespace and caps the title at 240 Unicode code points +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ runCanonical tests/canonical-software-factory.test.ts:69:22 + 67| }; + 68| + 69| const definition = getFlowDefinition(softwareFactory); + | ^ + 70| return definition.body(context as never, { issue, approver: 'khaliq'… + 71| root, + ❯ tests/canonical-software-factory.test.ts:99:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[5/63]⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > fails closed before push when GitHub identity is missing or the final body duplicates its closing line +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ tests/canonical-software-factory.test.ts:108:24 + 106| + 107| it('fails closed before push when GitHub identity is missing or the … + 108| const definition = getFlowDefinition(softwareFactory); + | ^ + 109| const commands: string[] = []; + 110| let completionReason = ''; + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[6/63]⎯ + + FAIL tests/communication-mixed-resume.test.ts > resumes mixed ordinary and linked agents through the real daemon without stealing peer capacity +AssertionError: {"ok":false,"command":"resume","resolutions":[],"diagnostics":[{"severity":"failure","kind":"protocol_error","message":"relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`"}],"runId":"01M36B7SMVPW00H67FE03NC6XS","socketPath":"/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-c074cf92d10f.sock"}: expected { ok: false, command: 'resume', …(4) } to match object { status: 'completed', …(2) } +(6 matching properties omitted from actual) + +- Expected ++ Received + + Object { +- "completedSteps": 4, +- "completionReason": "success", +- "status": "completed", ++ "command": "resume", ++ "diagnostics": Array [ ++ Object { ++ "kind": "protocol_error", ++ "message": "relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`", ++ "severity": "failure", ++ }, ++ ], ++ "ok": false, ++ "resolutions": Array [], ++ "runId": "01M36B7SMVPW00H67FE03NC6XS", ++ "socketPath": "/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-c074cf92d10f.sock", + } + + ❯ tests/communication-mixed-resume.test.ts:48:60 + 46| expect(parked.status).toBe('parked'); + 47| const resumed = await resumeFlow(parked.run_id, dataDir, { localAg… + 48| expect(resumed.report, JSON.stringify(resumed.report)).toMatchObje… + | ^ + 49| expect(state.started).toEqual(new Set(['a', 'b'])); + 50| const entries = (await client.journalRead(parked.run_id, 1)).entri… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[7/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > shadows inherited thenables on completed snapshot authority +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:33:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[8/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > locks the inherited then slot before authored code can schedule a replacement +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:58:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[9/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps the buffered base bytes when the live source changes +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:65:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[10/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps source digests sensitive when ambient Array.map is poisoned +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:78:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[11/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > defines source entries without consulting inherited numeric setters +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:118:75 + 116| }, + 117| }); + 118| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 119| } finally { + 120| if (previous === undefined) delete (Array.prototype as unknown a… + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:118:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[12/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > uses the module-captured platform during source traversal +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:142:75 + 140| }, + 141| }); + 142| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 143| } finally { + 144| Object.defineProperty(process, 'platform', descriptor); + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:142:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[13/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > shadows native directory arrays before promise resolution can substitute them +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:168:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[14/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > excludes project node_modules from the admitted generation +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:184:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[15/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > refuses an oversized source file before buffering its contents +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "snapshot entry or byte limit", + } + + ❯ tests/hosted-base-snapshot.test.ts:200:5 + 198| writeFileSync(oversized, ''); + 199| truncateSync(oversized, 64 * 1024 * 1024 + 1); + 200| await expect(createHostedBaseSnapshot(flowPath)).rejects.toMatchOb… + | ^ + 201| code: 'plugin_source_invalid', + 202| message: expect.stringContaining('snapshot entry or byte limit'), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[16/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > captures conversion and allocation across declarations and source snapshots +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "reviewed Software Factory base source", + } + + ❯ tests/hosted-base-snapshot.test.ts:247:7 + 245| return allocUnsafe(size); + 246| }) as typeof Buffer.allocUnsafe; + 247| await expect(loadHostedExtensionRuntime(flowPath)).rejects.toMat… + | ^ + 248| code: 'plugin_source_invalid', + 249| message: expect.stringContaining('reviewed Software Factory ba… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[17/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps declaration and reviewed-base reads bound after builtin export synchronization +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "reviewed Software Factory base source", + } + + ❯ tests/hosted-base-snapshot.test.ts:321:7 + 319| } + 320| syncBuiltinESMExports(); + 321| await expect(loadHostedExtensionRuntime(flowPath)).rejects.toMat… + | ^ + 322| code: 'plugin_source_invalid', + 323| message: expect.stringContaining('reviewed Software Factory ba… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[18/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > reads through captured descriptors and charges admitted descriptor sizes +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:380:75 + 378| }, + 379| }); + 380| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 381| } finally { + 382| fileHandlePrototype.read = fileHandleRead; + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:380:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[19/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > ignores inherited snapshot test hooks when production omits them +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:403:75 + 401| }); + 402| } + 403| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 404| } finally { + 405| for (let index = 0; index < names.length; index += 1) { + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:403:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[20/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > bounds a file that grows after its admitted size was checked +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "changed while reading \"raced.bin\"", + } + + ❯ tests/hosted-base-snapshot.test.ts:418:5 + 416| const raced = join(project, 'raced.bin'); + 417| writeFileSync(raced, 'small'); + 418| await expect( + | ^ + 419| hostedBaseSourceDigest([{ root: project, prefix: '' }], { + 420| afterStat: async path => { + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[21/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > stops streaming project entries at the shared count limit +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "snapshot entry or byte limit", + } + + ❯ tests/hosted-base-snapshot.test.ts:467:5 + 465| writeFileSync(join(project, `empty-${index}.txt`), ''); + 466| } + 467| await expect(createHostedBaseSnapshot(flowPath)).rejects.toMatchOb… + | ^ + 468| code: 'plugin_source_invalid', + 469| message: expect.stringContaining('snapshot entry or byte limit'), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[22/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > resolves locked artifacts without importing extension top-level code +AssertionError: expected [ { …(6) } ] to deeply equal [ { …(6) } ] + +- Expected ++ Received + + Array [ + Object { + "digest": "ff644144d69775d9ccc63954eea7bedff16571a45c8ac01f2f4c8d23a5ecdc9a", +- "directory": "/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/hosted-extension-test-LRLpYP/.flows/plugins/babysitter@sha256:ff644144d69775d9ccc63954eea7bedff16571a45c8ac01f2f4c8d23a5ecdc9a", ++ "directory": "/private/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/hosted-extension-test-LRLpYP/.flows/plugins/babysitter@sha256:ff644144d69775d9ccc63954eea7bedff16571a45c8ac01f2f4c8d23a5ecdc9a", + "manifestSha256": "5681622ac4930a1b9c7eda1736b52111e7ae8711c0a29eed0db00bb639fc7375", + "name": "babysitter", + "ref": "github:AgentWorkforce/flows@aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa#extensions/babysitter", + "version": "0.2.0", + }, + ] + + ❯ tests/hosted-extension-isolation.test.ts:159:70 + 157| const flowPath = join(root, 'software-factory.flow.ts'); + 158| writeFileSync(flowPath, 'export default {};'); + 159| expect((await loadHostedExtensionArtifacts(flowPath)).artifacts).t… + | ^ + 160| expect(() => readFileSync(marker)).toThrow(); + 161| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[23/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > executes the exact capability-only handler for a queued receipt + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > executes the exact capability-only handler for a duplicate receipt +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:221:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[24/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > streams verified bytes when the live store is replaced and no writable staging path exists +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:431:7 + 429| return { receiptId: 'receipt-1', status: 'queued' }; + 430| } }, + 431| })).resolves.toEqual({ completionReason: 'success', capabilityCall… + | ^ + 432| expect(calls).toBe(1); + 433| expect(readFileSync(join(installed.directory, 'babysitter.flow.ts'… + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:424:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[25/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > mounts pinned private Surface bytes when the live package changes before launch +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:456:7 + 454| return { receiptId: 'receipt-1', status: 'queued' }; + 455| } }, + 456| })).resolves.toEqual({ completionReason: 'success', capabilityCall… + | ^ + 457| expect(calls).toBe(1); + 458| expect(readFileSync(join(surfaceRoot, 'dist/flow.js'), 'utf8')).to… + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:443:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[26/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > refuses oversized Surface files through the bounded descriptor reader +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_unsupported", +- "message": StringContaining "cannot read pinned Surface runtime flow.js", + } + + ❯ tests/hosted-extension-isolation.test.ts:489:5 + 487| }); + 488| let calls = 0; + 489| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 490| artifact: await artifact(), + 491| manifest: validateFlowExtensionManifest(manifest()), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[27/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > shields verified Surface files before async settlement +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:532:9 + 530| surfaceRoot, + 531| babysitterTurn: { queue: async () => ({ receiptId: 'receipt-1'… + 532| })).resolves.toEqual({ completionReason: 'success', capabilityCa… + | ^ + 533| } finally { + 534| if (previous === undefined) delete (Array.prototype as { then?: … + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:525:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[28/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > preserves a typed host refusal while disclosing only a fixed marker to the child +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to be Error: private Cloud policy detail { code: '…' } // Object.is equality + +- Expected ++ Received + +- [Error: private Cloud policy detail] ++ [Error: hosted extension isolation requires Linux] + + ❯ tests/hosted-extension-isolation.test.ts:560:5 + 558| provider: 'github', eventType: 'pull_request.labeled', deliveryI… + 559| }); + 560| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 561| artifact: await artifact(source), manifest: validateFlowExtensio… + 562| babysitterTurn: { queue: async () => { throw refusal; } }, + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[29/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > denies ambient credentials, host files, writes, network, subprocesses, and undeclared context verbs +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:624:13 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[30/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > enforces OS address-space and data bounds on native Buffer allocation +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_unsupported", +- "message": StringMatching /(?:Failed to allocate memory|Array buffer allocation failed)/u, + } + + ❯ tests/hosted-extension-isolation.test.ts:646:5 + 644| }); + 645| let calls = 0; + 646| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 647| artifact: installed, + 648| manifest: validateFlowExtensionManifest(manifest()), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[31/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > blocks extra handler fields and authority-bearing receipt fields at the parent port +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + +- Expected ++ Received + +- Object { +- "code": "plugin_event_unroutable", ++ PluginError { ++ "code": "plugin_unsupported", + } + + ❯ tests/hosted-extension-isolation.test.ts:675:5 + 673| }); + 674| let calls = 0; + 675| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 676| artifact: await artifact(source), manifest: validateFlowExtensio… + 677| babysitterTurn: { queue: async () => { calls += 1; return { rece… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[32/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > constructs adapter authority with the captured freeze intrinsic +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:745:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[33/63]⎯ + + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different PR frame with zero adapter calls + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different delivery frame with zero adapter calls + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different event frame with zero adapter calls +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + +- Expected ++ Received + +- Object { +- "code": "plugin_event_unroutable", ++ PluginError { ++ "code": "plugin_unsupported", + } + + ❯ tests/hosted-extension-protocol.test.ts:430:5 + 428| ])('rejects an import-time %s frame with zero adapter calls', async … + 429| let calls = 0; + 430| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 431| artifact: await artifact(hostileImport([frame, { type: 'error', … + 432| manifest: validateFlowExtensionManifest(manifest()), dispatch: d… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[34/63]⎯ + + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects two forged calls after the authoritative first outcome settles + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > waits for a pending adapter to reject after a forged child error + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > waits for a pending adapter to resolve after a forged child error + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > returns a typed adapter rejection even when the hostile child hangs +Error: hostile child did not invoke the adapter + ❯ Timeout._onTimeout tests/hosted-extension-protocol.test.ts:118:45 + 116| async function waitForInvocation(invoked: Promise): Promise((resolve, reject) => { + 118| const timeout = setTimeout(() => reject(new Error('hostile child d… + | ^ + 119| void invoked.then(() => { clearTimeout(timeout); resolve(); }, rej… + 120| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[35/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses an 8-character run-id prefix: Cloud has no prefix lookup +AssertionError: expected [Function] to throw error matching /not full Cloud run ids: c649fe14/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/not full Cloud run ids: c649fe14/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[36/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses the whole batch when any id is invalid, rather than dropping it +AssertionError: expected [Function] to throw error matching /not full Cloud run ids: nope!/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/not full Cloud run ids: nope!/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[37/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses an empty batch +AssertionError: expected [Function] to throw error matching /needs runIds/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/needs runIds/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[38/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses a batch too large for the edge step lease +AssertionError: expected [Function] to throw error matching /exceeds the 8 that fit/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/exceeds the 8 that fit/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[39/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > accepts eight ids — the incident batch is inside the bound +AssertionError: promise rejected "TypeError: expected an @relayflows/surfac…" instead of resolving + ❯ tests/stuck-run-triage.test.ts:62:40 + 60| it('accepts eight ids — the incident batch is inside the bound', asy… + 61| const ids = Array.from({ length: 8 }, (_, i) => `${ID_A.slice(0, -… + 62| await expect(drive({ runIds: ids })).resolves.toBeDefined(); + | ^ + 63| }); + 64| }); + +Caused by: TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + ❯ tests/stuck-run-triage.test.ts:62:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[40/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > refuses to send the Cloud bearer token to an unapproved origin +AssertionError: expected [Function] to throw error matching /refusing to send the Cloud bearer to…/\ but got 'expected an @relayflows/surface flow …' + +- Expected: +/refusing to send the Cloud bearer token to https:\/\/evil\.example/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[41/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > refuses a non-URL apiUrl +AssertionError: expected [Function] to throw error matching /is not a URL/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/is not a URL/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[42/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > allows an approved origin and uses it in the curl +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:77:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[43/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > defaults to production Cloud +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:82:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[44/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > never publishes a run record the fetch did not produce +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:89:34 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[45/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > names the Worker on every wrangler invocation +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:98:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[46/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > accepts caller-supplied Workers and rejects option-shaped ones +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:107:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[47/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > falls back when GNU timeout is absent, as it is on macOS +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:115:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[48/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > runs the tails concurrently so wall time does not scale with the batch +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:123:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[49/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > records wrangler's own exit status rather than head's +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:129:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[50/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage shell text > parses under both sh and bash +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:137:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[51/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage shell text > collects tails with no GNU timeout on PATH, as on a stock macOS +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:157:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[52/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage agents > declares read-only permissions on every agent +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:176:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[53/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage agents > tells the forensics agents their evidence is untrusted +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:182:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[54/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > refuses a duplicate run id: two tails would share one evidence file +AssertionError: expected [Function] to throw error matching /duplicate runIds: c649fe14-0c2e-4e51-…/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/duplicate runIds: c649fe14-0c2e-4e51-9a6a-4f0d1b0f77aa/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[55/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > refuses a duplicate Worker name for the same reason +AssertionError: expected [Function] to throw error matching /duplicate workers: w-one/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/duplicate workers: w-one/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[56/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > bounds ids x workers, not just ids +AssertionError: expected [Function] to throw error matching /24 concurrent tails, over the 16/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/24 concurrent tails, over the 16/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[57/63]⎯ + +⎯⎯⎯⎯⎯⎯ Unhandled Errors ⎯⎯⎯⎯⎯⎯ + +Vitest caught 2 unhandled errors during the test run. +This might cause false positive tests. Resolve unhandled errors to make sure your tests are not affected. + +⎯⎯⎯⎯ Unhandled Rejection ⎯⎯⎯⎯⎯ +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-protocol.test.ts:454:17 + ❯ node_modules/@vitest/runner/dist/index.js:533:5 + ❯ runTest node_modules/@vitest/runner/dist/index.js:1056:11 + ❯ runSuite node_modules/@vitest/runner/dist/index.js:1205:15 + ❯ runSuite node_modules/@vitest/runner/dist/index.js:1205:15 + ❯ runFiles node_modules/@vitest/runner/dist/index.js:1262:5 + ❯ startTests node_modules/@vitest/runner/dist/index.js:1271:3 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +This error originated in "tests/hosted-extension-protocol.test.ts" test file. It doesn't mean the error was thrown inside the file itself, but while it was running. +The latest test that might've caused the error is "rejects two forged calls after the authoritative first outcome settles". It might mean one of the following: +- The error was thrown, while Vitest was running this test. +- If the error occurred after the test had been completed, this was the last documented test before it was thrown. + +⎯⎯⎯⎯⎯ Uncaught Exception ⎯⎯⎯⎯⎯ +Error: spawn /Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd ENOENT + ❯ Process.ChildProcess._handle.onexit node:internal/child_process:285:19 + ❯ onErrorNT node:internal/child_process:483:16 + ❯ processTicksAndRejections node:internal/process/task_queues:89:21 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { errno: -2, code: 'ENOENT', syscall: 'spawn /Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd', path: '/Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd', spawnargs: [ '--data-dir', '/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/flows-mcp-daemon-GLuXGX', 'serve' ] } +This error originated in "tests/mcp.test.ts" test file. It doesn't mean the error was thrown inside the file itself, but while it was running. +The latest test that might've caused the error is "authored MCP effects against the real kernel". It might mean one of the following: +- The error was thrown, while Vitest was running this test. +- If the error occurred after the test had been completed, this was the last documented test before it was thrown. +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ + + Test Files 9 failed | 4 passed (13) + Tests 60 failed | 114 passed | 64 skipped (238) + Errors 2 errors + Start at 22:18:53 + Duration 41.74s (transform 1.54s, setup 0ms, collect 7.80s, tests 113.90s, environment 16ms, prepare 649ms) + +exit=1 +$ cp /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts src/ +exit=0 +$ cmp src/llm-worker.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts && cmp src/authored-flow-executor.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts +exit=0 diff --git a/evidence/llm-fence-predicate-resume/thirteen-files-on-branch.txt b/evidence/llm-fence-predicate-resume/thirteen-files-on-branch.txt new file mode 100644 index 000000000..3c1f33c79 --- /dev/null +++ b/evidence/llm-fence-predicate-resume/thirteen-files-on-branch.txt @@ -0,0 +1,1193 @@ +$ npx vitest run tests/authored-node-runtime.test.ts tests/babysitter-native-extension.test.ts tests/bundle.test.ts tests/canonical-software-factory.test.ts tests/cli-watch.test.ts tests/communication-mixed-resume.test.ts tests/hosted-base-snapshot.test.ts tests/hosted-extension-isolation.test.ts tests/hosted-extension-protocol.test.ts tests/mcp.test.ts tests/stop-process-group.test.ts tests/stuck-run-triage.test.ts tests/worker-cli.test.ts + + RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk + + ❯ tests/hosted-base-snapshot.test.ts (18 tests | 15 failed | 1 skipped) 1008ms + × hosted base private snapshot > shadows inherited thenables on completed snapshot authority 11ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > locks the inherited then slot before authored code can schedule a replacement 3ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > keeps the buffered base bytes when the live source changes 1ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > keeps source digests sensitive when ambient Array.map is poisoned 1ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > defines source entries without consulting inherited numeric setters 3ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > uses the module-captured platform during source traversal 1ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > shadows native directory arrays before promise resolution can substitute them 1ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > excludes project node_modules from the admitted generation 2ms + → Hosted base source snapshotting requires Linux. + × hosted base private snapshot > refuses an oversized source file before buffering its contents 4ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > captures conversion and allocation across declarations and source snapshots 5ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > keeps declaration and reviewed-base reads bound after builtin export synchronization 3ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > reads through captured descriptors and charges admitted descriptor sizes 4ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > ignores inherited snapshot test hooks when production omits them 1ms + → promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + × hosted base private snapshot > bounds a file that grows after its admitted size was checked 1ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + × hosted base private snapshot > stops streaming project entries at the shared count limit 961ms + → expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + ❯ tests/hosted-extension-isolation.test.ts (22 tests | 12 failed | 3 skipped) 174ms + × hosted extension capability isolation > resolves locked artifacts without importing extension top-level code 16ms + → expected [ { …(6) } ] to deeply equal [ { …(6) } ] + × hosted extension capability isolation > executes the exact capability-only handler for a queued receipt 27ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > executes the exact capability-only handler for a duplicate receipt 9ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > streams verified bytes when the live store is replaced and no writable staging path exists 17ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > mounts pinned private Surface bytes when the live package changes before launch 15ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > refuses oversized Surface files through the bounded descriptor reader 6ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + × hosted extension capability isolation > shields verified Surface files before async settlement 5ms + → promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + × hosted extension capability isolation > preserves a typed host refusal while disclosing only a fixed marker to the child 4ms + → expected Error: hosted extension isolation require… { code: '…' } to be Error: private Cloud policy detail { code: '…' } // Object.is equality + × hosted extension capability isolation > denies ambient credentials, host files, writes, network, subprocesses, and undeclared context verbs 7ms + → hosted extension isolation requires Linux + × hosted extension capability isolation > enforces OS address-space and data bounds on native Buffer allocation 3ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + × hosted extension capability isolation > blocks extra handler fields and authority-bearing receipt fields at the parent port 3ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension capability isolation > constructs adapter authority with the captured freeze intrinsic 3ms + → hosted extension isolation requires Linux + ❯ tests/babysitter-native-extension.test.ts (41 tests | 41 skipped) 681ms + ❯ tests/authored-node-runtime.test.ts (14 tests | 14 skipped) 52ms + ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) 141ms + × stuck-run-triage input validation > refuses an 8-character run-id prefix: Cloud has no prefix lookup 3ms + → expected [Function] to throw error matching /not full Cloud run ids: c649fe14/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses the whole batch when any id is invalid, rather than dropping it 1ms + → expected [Function] to throw error matching /not full Cloud run ids: nope!/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses an empty batch 0ms + → expected [Function] to throw error matching /needs runIds/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > refuses a batch too large for the edge step lease 0ms + → expected [Function] to throw error matching /exceeds the 8 that fit/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage input validation > accepts eight ids — the incident batch is inside the bound 3ms + → promise rejected "TypeError: expected an @relayflows/surfac…" instead of resolving + × stuck-run-triage apiUrl > refuses to send the Cloud bearer token to an unapproved origin 0ms + → expected [Function] to throw error matching /refusing to send the Cloud bearer to…/\ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage apiUrl > refuses a non-URL apiUrl 0ms + → expected [Function] to throw error matching /is not a URL/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage apiUrl > allows an approved origin and uses it in the curl 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage apiUrl > defaults to production Cloud 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage apiUrl > never publishes a run record the fetch did not produce 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > names the Worker on every wrangler invocation 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > accepts caller-supplied Workers and rejects option-shaped ones 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > falls back when GNU timeout is absent, as it is on macOS 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > runs the tails concurrently so wall time does not scale with the batch 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage edge collection > records wrangler's own exit status rather than head's 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage shell text > parses under both sh and bash 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage shell text > collects tails with no GNU timeout on PATH, as on a stock macOS 129ms + → expected an @relayflows/surface flow handle + × stuck-run-triage agents > declares read-only permissions on every agent 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage agents > tells the forensics agents their evidence is untrusted 0ms + → expected an @relayflows/surface flow handle + × stuck-run-triage fan-out > refuses a duplicate run id: two tails would share one evidence file 0ms + → expected [Function] to throw error matching /duplicate runIds: c649fe14-0c2e-4e51-…/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage fan-out > refuses a duplicate Worker name for the same reason 0ms + → expected [Function] to throw error matching /duplicate workers: w-one/ but got 'expected an @relayflows/surface flow …' + × stuck-run-triage fan-out > bounds ids x workers, not just ids 0ms + → expected [Function] to throw error matching /24 concurrent tails, over the 16/ but got 'expected an @relayflows/surface flow …' + ❯ tests/canonical-software-factory.test.ts (3 tests | 3 failed) 57ms + × canonical software-factory metadata contract > opens the actual catalog flow with the ticket title and exactly one GitHub closing line 35ms + → expected an @relayflows/surface flow handle + × canonical software-factory metadata contract > normalizes whitespace and caps the title at 240 Unicode code points 21ms + → expected an @relayflows/surface flow handle + × canonical software-factory metadata contract > fails closed before push when GitHub identity is missing or the final body duplicates its closing line 0ms + → expected an @relayflows/surface flow handle + ❯ tests/communication-mixed-resume.test.ts (1 test | 1 failed) 86ms + × resumes mixed ordinary and linked agents through the real daemon without stealing peer capacity 86ms + → {"ok":false,"command":"resume","resolutions":[],"diagnostics":[{"severity":"failure","kind":"protocol_error","message":"relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`"}],"runId":"01M36B92A5PSZYQ4AK0G5F51XC","socketPath":"/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-1d0c651808a5.sock"}: expected { ok: false, command: 'resume', …(4) } to match object { status: 'completed', …(2) } +(6 matching properties omitted from actual) + ✓ tests/bundle.test.ts (26 tests) 6547ms + ✓ immutable bundles > verifies with --verify in any position and answers --json with one object 556ms + ✓ immutable bundles > builds and verifies the canonical YAML fixture through the compiled CLI 845ms + ✓ immutable bundles > emits the ephemeral warning on CLI stderr and uses the default output directory 491ms + ✓ immutable bundles > builds a standalone TS fixture twice with identical executable hashes 2133ms + ❯ tests/mcp.test.ts (30 tests | 4 skipped) 9628ms + ✓ MCP preflight and transports > flows check refuses an undeclared server with exit 2 and no daemon 847ms + ✓ MCP preflight and transports > flows check reports a refusing server and leaves no PID 645ms + ✓ MCP preflight and transports > kills a SIGTERM-resistant silent child after a parent-owned handshake deadline 1332ms + ✓ MCP preflight and transports > reaps a SIGTERM-resistant descendant with inherit stdio before cleanup finishes 1141ms + ✓ MCP preflight and transports > reaps a SIGTERM-resistant descendant with ignore stdio before cleanup finishes 2074ms + ✓ MCP preflight and transports > reports malformed connection configuration as config_invalid 422ms + ✓ tests/cli-watch.test.ts (10 tests) 12012ms + ✓ flows check --watch > rechecks syntax errors, clears once, and returns the last refusal on Ctrl-C 1066ms + ✓ flows check --watch > streams JSON lines without ANSI, recovers after atomic saves, and exits zero after repair 1413ms + ✓ flows check --watch > coalesces 20 concurrent saves into at most two rechecks 1456ms + ✓ flows check --watch > watches transitive relative use imports, cycles, and nearest config changes 1732ms + ✓ flows check --watch > refreshes the import graph and notices missing imports being created 1831ms + ✓ flows check --watch > reloads authored TypeScript instead of reusing the first imported definition 1200ms + ✓ flows check --watch > detects a nearer config appearing and falls back after it is deleted 1248ms + ✓ flows check --watch > keeps watching after the target is deleted and recreated 1292ms + ✓ flows check --watch > queues changes during a slow check without overlapping checks 772ms + ✓ tests/stop-process-group.test.ts (9 tests) 15379ms + ✓ every stop reaches the process group, not just the direct child > exits the run after an execution-timeout stop 1225ms + ✓ every stop reaches the process group, not just the direct child > exits the run after a protocol terminate stop 1946ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after a protocol terminate stop 1790ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after an execution-timeout stop 2122ms + ✓ every stop reaches the process group, not just the direct child > holds the loop open long enough for the escalation to run 1120ms + ✓ every stop reaches the process group, not just the direct child > terminate() forces a group that outlives SIGTERM 303ms + ✓ a wrapper that exits with no execution deadline still drains > reports the wrapper result and reaps a grandchild holding its pipes 777ms + ✓ a wrapper that exits with no execution deadline still drains > reaps a SIGTERM-deaf grandchild holding its pipes 1727ms + ✓ a wrapper that exits with no execution deadline still drains > settles on its own deadline when an escaped holder withholds close 4369ms + ✓ tests/worker-cli.test.ts (18 tests) 24622ms + ✓ registered CLI model defaults > passes the same priced Claude default to the real provider invocation 796ms + ✓ step discovery environment > names the run, step, attempt and an absolute data dir for a direct agent spawn 636ms + ✓ step discovery environment > exports none of the four without a data dir, even when the worker inherited them 427ms + ✓ wrapper discovery environment > sets the four names from the dispatch and still refuses ambient values and other secrets 387ms + ✓ wrapper discovery environment > exports none of the four to a wrapper without a data dir, even when the worker inherited them 338ms + ✓ custom wrapper execution identity > passes an explicit safe environment at identification and execution 396ms + ✓ custom wrapper execution identity > refuses a wrapper symlink retarget before delivering private values 392ms + ✓ custom wrapper execution identity > bounds wrapper execution after acknowledgement 316ms + ✓ custom wrapper execution identity > bounds captured wrapper output 304ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a conforming wrapper leaks a stdio pipe to a background helper 1879ms + ✓ custom wrapper execution bounds are reader-owned > resolves when the leaked helper inherits stderr only 1881ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a wrapper leaks a stdio pipe and exits before identifying 3430ms + ✓ custom wrapper execution bounds are reader-owned > journals a completionReason at the default bound when a wrapper leaks a stdio pipe 11390ms + ✓ custom wrapper execution bounds are reader-owned > accepts an execute token and an over-8KiB payload flushed in one write 379ms + ✓ custom wrapper execution bounds are reader-owned > accepts the same over-8KiB payload whether or not it coalesces with the execute token 832ms + ❯ tests/hosted-extension-protocol.test.ts (24 tests | 7 failed | 1 skipped) 40125ms + × hosted extension hostile protocol > rejects an import-time different PR frame with zero adapter calls 10ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects an import-time different delivery frame with zero adapter calls 4ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects an import-time different event frame with zero adapter calls 4ms + → expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + × hosted extension hostile protocol > rejects two forged calls after the authoritative first outcome settles 10006ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > waits for a pending adapter to reject after a forged child error 10003ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > waits for a pending adapter to resolve after a forged child error 10005ms + → hostile child did not invoke the adapter + × hosted extension hostile protocol > returns a typed adapter rejection even when the hostile child hangs 10009ms + → hostile child did not invoke the adapter + +⎯⎯⎯⎯⎯⎯ Failed Suites 3 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/authored-node-runtime.test.ts [ tests/authored-node-runtime.test.ts ] +AssertionError: expected '1.3.14' to be '1.4.0' // Object.is equality + +Expected: "1.4.0" +Received: "1.3.14" + + ❯ tests/authored-node-runtime.test.ts:18:77 + 16| + 17| beforeAll(() => { + 18| expect(spawnSync(bun, ['--version'], { encoding: 'utf8' }).stdout.tr… + | ^ + 19| expect(existsSync(daemon), 'build the current kernel or set RELAYFLO… + 20| stage = mkdtempSync(join(tmpdir(), 'authored-standalone-build-')); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/63]⎯ + + FAIL tests/babysitter-native-extension.test.ts [ tests/babysitter-native-extension.test.ts ] +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ baseAt src/hosted-extension-runtime.ts:234:20 + ❯ Module.loadHostedExtensionRuntime src/hosted-extension-runtime.ts:90:16 + ❯ composed tests/babysitter-native-extension.test.ts:64:25 + ❯ tests/babysitter-native-extension.test.ts:100:37 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[2/63]⎯ + + FAIL tests/mcp.test.ts > authored MCP effects against the real kernel +Error: journal client: connect failed: connect ENOENT /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-59cc00f3b2b2.sock + ❯ Socket.onError src/journal-client.ts:102:16 + 100| socket.removeAllListeners(); + 101| this.failAll(err); + 102| reject(new Error(`journal client: connect failed: ${err.messag… + | ^ + 103| }; + 104| socket.once('error', onError); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[3/63]⎯ + +⎯⎯⎯⎯⎯⎯ Failed Tests 60 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > opens the actual catalog flow with the ticket title and exactly one GitHub closing line +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ runCanonical tests/canonical-software-factory.test.ts:69:22 + 67| }; + 68| + 69| const definition = getFlowDefinition(softwareFactory); + | ^ + 70| return definition.body(context as never, { issue, approver: 'khaliq'… + 71| root, + ❯ tests/canonical-software-factory.test.ts:82:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[4/63]⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > normalizes whitespace and caps the title at 240 Unicode code points +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ runCanonical tests/canonical-software-factory.test.ts:69:22 + 67| }; + 68| + 69| const definition = getFlowDefinition(softwareFactory); + | ^ + 70| return definition.body(context as never, { issue, approver: 'khaliq'… + 71| root, + ❯ tests/canonical-software-factory.test.ts:99:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[5/63]⎯ + + FAIL tests/canonical-software-factory.test.ts > canonical software-factory metadata contract > fails closed before push when GitHub identity is missing or the final body duplicates its closing line +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ tests/canonical-software-factory.test.ts:108:24 + 106| + 107| it('fails closed before push when GitHub identity is missing or the … + 108| const definition = getFlowDefinition(softwareFactory); + | ^ + 109| const commands: string[] = []; + 110| let completionReason = ''; + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[6/63]⎯ + + FAIL tests/communication-mixed-resume.test.ts > resumes mixed ordinary and linked agents through the real daemon without stealing peer capacity +AssertionError: {"ok":false,"command":"resume","resolutions":[],"diagnostics":[{"severity":"failure","kind":"protocol_error","message":"relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`"}],"runId":"01M36B92A5PSZYQ4AK0G5F51XC","socketPath":"/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-1d0c651808a5.sock"}: expected { ok: false, command: 'resume', …(4) } to match object { status: 'completed', …(2) } +(6 matching properties omitted from actual) + +- Expected ++ Received + + Object { +- "completedSteps": 4, +- "completionReason": "success", +- "status": "completed", ++ "command": "resume", ++ "diagnostics": Array [ ++ Object { ++ "kind": "protocol_error", ++ "message": "relayflowd could not complete the resume request: bad_request: unknown field `required_streams`, expected one of `worker_id`, `step_types`, `capacity`, `pins`", ++ "severity": "failure", ++ }, ++ ], ++ "ok": false, ++ "resolutions": Array [], ++ "runId": "01M36B92A5PSZYQ4AK0G5F51XC", ++ "socketPath": "/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/relayflowd-1d0c651808a5.sock", + } + + ❯ tests/communication-mixed-resume.test.ts:48:60 + 46| expect(parked.status).toBe('parked'); + 47| const resumed = await resumeFlow(parked.run_id, dataDir, { localAg… + 48| expect(resumed.report, JSON.stringify(resumed.report)).toMatchObje… + | ^ + 49| expect(state.started).toEqual(new Set(['a', 'b'])); + 50| const entries = (await client.journalRead(parked.run_id, 1)).entri… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[7/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > shadows inherited thenables on completed snapshot authority +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:33:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[8/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > locks the inherited then slot before authored code can schedule a replacement +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:58:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[9/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps the buffered base bytes when the live source changes +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:65:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[10/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps source digests sensitive when ambient Array.map is poisoned +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:78:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[11/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > defines source entries without consulting inherited numeric setters +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:118:75 + 116| }, + 117| }); + 118| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 119| } finally { + 120| if (previous === undefined) delete (Array.prototype as unknown a… + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:118:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[12/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > uses the module-captured platform during source traversal +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:142:75 + 140| }, + 141| }); + 142| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 143| } finally { + 144| Object.defineProperty(process, 'platform', descriptor); + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:142:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[13/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > shadows native directory arrays before promise resolution can substitute them +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:168:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[14/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > excludes project node_modules from the admitted generation +Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + 348| + 349| function invalid(message: string): PluginError { + 350| return new PluginError('plugin_source_invalid', message); + | ^ + 351| } + 352| + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.createHostedBaseSnapshot src/hosted-base-snapshot.ts:126:23 + ❯ tests/hosted-base-snapshot.test.ts:184:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[15/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > refuses an oversized source file before buffering its contents +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "snapshot entry or byte limit", + } + + ❯ tests/hosted-base-snapshot.test.ts:200:5 + 198| writeFileSync(oversized, ''); + 199| truncateSync(oversized, 64 * 1024 * 1024 + 1); + 200| await expect(createHostedBaseSnapshot(flowPath)).rejects.toMatchOb… + | ^ + 201| code: 'plugin_source_invalid', + 202| message: expect.stringContaining('snapshot entry or byte limit'), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[16/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > captures conversion and allocation across declarations and source snapshots +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "reviewed Software Factory base source", + } + + ❯ tests/hosted-base-snapshot.test.ts:247:7 + 245| return allocUnsafe(size); + 246| }) as typeof Buffer.allocUnsafe; + 247| await expect(loadHostedExtensionRuntime(flowPath)).rejects.toMat… + | ^ + 248| code: 'plugin_source_invalid', + 249| message: expect.stringContaining('reviewed Software Factory ba… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[17/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > keeps declaration and reviewed-base reads bound after builtin export synchronization +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "reviewed Software Factory base source", + } + + ❯ tests/hosted-base-snapshot.test.ts:321:7 + 319| } + 320| syncBuiltinESMExports(); + 321| await expect(loadHostedExtensionRuntime(flowPath)).rejects.toMat… + | ^ + 322| code: 'plugin_source_invalid', + 323| message: expect.stringContaining('reviewed Software Factory ba… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[18/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > reads through captured descriptors and charges admitted descriptor sizes +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:380:75 + 378| }, + 379| }); + 380| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 381| } finally { + 382| fileHandlePrototype.read = fileHandleRead; + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:380:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[19/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > ignores inherited snapshot test hooks when production omits them +AssertionError: promise rejected "Error: Hosted base source snapshotting re… { code: '…' }" instead of resolving + ❯ tests/hosted-base-snapshot.test.ts:403:75 + 401| }); + 402| } + 403| await expect(hostedBaseSourceDigest([{ root: project, prefix: ''… + | ^ + 404| } finally { + 405| for (let index = 0; index < names.length; index += 1) { + +Caused by: Error: Hosted base source snapshotting requires Linux. + ❯ invalid src/hosted-base-snapshot.ts:350:10 + ❯ readTree src/hosted-base-snapshot.ts:235:11 + ❯ readAuthorityFiles src/hosted-base-snapshot.ts:181:31 + ❯ Module.hostedBaseSourceDigest src/hosted-base-snapshot.ts:166:29 + ❯ tests/hosted-base-snapshot.test.ts:403:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_source_invalid' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[20/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > bounds a file that grows after its admitted size was checked +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "changed while reading \"raced.bin\"", + } + + ❯ tests/hosted-base-snapshot.test.ts:418:5 + 416| const raced = join(project, 'raced.bin'); + 417| writeFileSync(raced, 'small'); + 418| await expect( + | ^ + 419| hostedBaseSourceDigest([{ root: project, prefix: '' }], { + 420| afterStat: async path => { + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[21/63]⎯ + + FAIL tests/hosted-base-snapshot.test.ts > hosted base private snapshot > stops streaming project entries at the shared count limit +AssertionError: expected Error: Hosted base source snapshotting re… { code: '…' } to match object { code: 'plugin_source_invalid', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_source_invalid", +- "message": StringContaining "snapshot entry or byte limit", + } + + ❯ tests/hosted-base-snapshot.test.ts:467:5 + 465| writeFileSync(join(project, `empty-${index}.txt`), ''); + 466| } + 467| await expect(createHostedBaseSnapshot(flowPath)).rejects.toMatchOb… + | ^ + 468| code: 'plugin_source_invalid', + 469| message: expect.stringContaining('snapshot entry or byte limit'), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[22/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > resolves locked artifacts without importing extension top-level code +AssertionError: expected [ { …(6) } ] to deeply equal [ { …(6) } ] + +- Expected ++ Received + + Array [ + Object { + "digest": "b816d7a5647f967d5253cd481a86f5c6894f0952e4223db114185ef3193f9107", +- "directory": "/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/hosted-extension-test-9oxlGe/.flows/plugins/babysitter@sha256:b816d7a5647f967d5253cd481a86f5c6894f0952e4223db114185ef3193f9107", ++ "directory": "/private/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/hosted-extension-test-9oxlGe/.flows/plugins/babysitter@sha256:b816d7a5647f967d5253cd481a86f5c6894f0952e4223db114185ef3193f9107", + "manifestSha256": "5681622ac4930a1b9c7eda1736b52111e7ae8711c0a29eed0db00bb639fc7375", + "name": "babysitter", + "ref": "github:AgentWorkforce/flows@aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa#extensions/babysitter", + "version": "0.2.0", + }, + ] + + ❯ tests/hosted-extension-isolation.test.ts:159:70 + 157| const flowPath = join(root, 'software-factory.flow.ts'); + 158| writeFileSync(flowPath, 'export default {};'); + 159| expect((await loadHostedExtensionArtifacts(flowPath)).artifacts).t… + | ^ + 160| expect(() => readFileSync(marker)).toThrow(); + 161| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[23/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > executes the exact capability-only handler for a queued receipt + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > executes the exact capability-only handler for a duplicate receipt +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:221:26 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[24/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > streams verified bytes when the live store is replaced and no writable staging path exists +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:431:7 + 429| return { receiptId: 'receipt-1', status: 'queued' }; + 430| } }, + 431| })).resolves.toEqual({ completionReason: 'success', capabilityCall… + | ^ + 432| expect(calls).toBe(1); + 433| expect(readFileSync(join(installed.directory, 'babysitter.flow.ts'… + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:424:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[25/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > mounts pinned private Surface bytes when the live package changes before launch +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:456:7 + 454| return { receiptId: 'receipt-1', status: 'queued' }; + 455| } }, + 456| })).resolves.toEqual({ completionReason: 'success', capabilityCall… + | ^ + 457| expect(calls).toBe(1); + 458| expect(readFileSync(join(surfaceRoot, 'dist/flow.js'), 'utf8')).to… + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:443:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[26/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > refuses oversized Surface files through the bounded descriptor reader +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_unsupported", +- "message": StringContaining "cannot read pinned Surface runtime flow.js", + } + + ❯ tests/hosted-extension-isolation.test.ts:489:5 + 487| }); + 488| let calls = 0; + 489| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 490| artifact: await artifact(), + 491| manifest: validateFlowExtensionManifest(manifest()), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[27/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > shields verified Surface files before async settlement +AssertionError: promise rejected "Error: hosted extension isolation require… { code: '…' }" instead of resolving + ❯ tests/hosted-extension-isolation.test.ts:532:9 + 530| surfaceRoot, + 531| babysitterTurn: { queue: async () => ({ receiptId: 'receipt-1'… + 532| })).resolves.toEqual({ completionReason: 'success', capabilityCa… + | ^ + 533| } finally { + 534| if (previous === undefined) delete (Array.prototype as { then?: … + +Caused by: Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:525:20 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[28/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > preserves a typed host refusal while disclosing only a fixed marker to the child +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to be Error: private Cloud policy detail { code: '…' } // Object.is equality + +- Expected ++ Received + +- [Error: private Cloud policy detail] ++ [Error: hosted extension isolation requires Linux] + + ❯ tests/hosted-extension-isolation.test.ts:560:5 + 558| provider: 'github', eventType: 'pull_request.labeled', deliveryI… + 559| }); + 560| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 561| artifact: await artifact(source), manifest: validateFlowExtensio… + 562| babysitterTurn: { queue: async () => { throw refusal; } }, + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[29/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > denies ambient credentials, host files, writes, network, subprocesses, and undeclared context verbs +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:624:13 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[30/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > enforces OS address-space and data bounds on native Buffer allocation +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_unsupported', …(1) } + +- Expected ++ Received + +- Object { ++ PluginError { + "code": "plugin_unsupported", +- "message": StringMatching /(?:Failed to allocate memory|Array buffer allocation failed)/u, + } + + ❯ tests/hosted-extension-isolation.test.ts:646:5 + 644| }); + 645| let calls = 0; + 646| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 647| artifact: installed, + 648| manifest: validateFlowExtensionManifest(manifest()), + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[31/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > blocks extra handler fields and authority-bearing receipt fields at the parent port +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + +- Expected ++ Received + +- Object { +- "code": "plugin_event_unroutable", ++ PluginError { ++ "code": "plugin_unsupported", + } + + ❯ tests/hosted-extension-isolation.test.ts:675:5 + 673| }); + 674| let calls = 0; + 675| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 676| artifact: await artifact(source), manifest: validateFlowExtensio… + 677| babysitterTurn: { queue: async () => { calls += 1; return { rece… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[32/63]⎯ + + FAIL tests/hosted-extension-isolation.test.ts > hosted extension capability isolation > constructs adapter authority with the captured freeze intrinsic +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-isolation.test.ts:745:22 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[33/63]⎯ + + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different PR frame with zero adapter calls + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different delivery frame with zero adapter calls + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects an import-time different event frame with zero adapter calls +AssertionError: expected Error: hosted extension isolation require… { code: '…' } to match object { code: 'plugin_event_unroutable' } + +- Expected ++ Received + +- Object { +- "code": "plugin_event_unroutable", ++ PluginError { ++ "code": "plugin_unsupported", + } + + ❯ tests/hosted-extension-protocol.test.ts:430:5 + 428| ])('rejects an import-time %s frame with zero adapter calls', async … + 429| let calls = 0; + 430| await expect(runVerifiedNativeExtensionSandbox({ + | ^ + 431| artifact: await artifact(hostileImport([frame, { type: 'error', … + 432| manifest: validateFlowExtensionManifest(manifest()), dispatch: d… + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[34/63]⎯ + + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > rejects two forged calls after the authoritative first outcome settles + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > waits for a pending adapter to reject after a forged child error + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > waits for a pending adapter to resolve after a forged child error + FAIL tests/hosted-extension-protocol.test.ts > hosted extension hostile protocol > returns a typed adapter rejection even when the hostile child hangs +Error: hostile child did not invoke the adapter + ❯ Timeout._onTimeout tests/hosted-extension-protocol.test.ts:118:45 + 116| async function waitForInvocation(invoked: Promise): Promise((resolve, reject) => { + 118| const timeout = setTimeout(() => reject(new Error('hostile child d… + | ^ + 119| void invoked.then(() => { clearTimeout(timeout); resolve(); }, rej… + 120| }); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[35/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses an 8-character run-id prefix: Cloud has no prefix lookup +AssertionError: expected [Function] to throw error matching /not full Cloud run ids: c649fe14/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/not full Cloud run ids: c649fe14/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[36/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses the whole batch when any id is invalid, rather than dropping it +AssertionError: expected [Function] to throw error matching /not full Cloud run ids: nope!/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/not full Cloud run ids: nope!/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[37/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses an empty batch +AssertionError: expected [Function] to throw error matching /needs runIds/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/needs runIds/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[38/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > refuses a batch too large for the edge step lease +AssertionError: expected [Function] to throw error matching /exceeds the 8 that fit/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/exceeds the 8 that fit/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[39/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage input validation > accepts eight ids — the incident batch is inside the bound +AssertionError: promise rejected "TypeError: expected an @relayflows/surfac…" instead of resolving + ❯ tests/stuck-run-triage.test.ts:62:40 + 60| it('accepts eight ids — the incident batch is inside the bound', asy… + 61| const ids = Array.from({ length: 8 }, (_, i) => `${ID_A.slice(0, -… + 62| await expect(drive({ runIds: ids })).resolves.toBeDefined(); + | ^ + 63| }); + 64| }); + +Caused by: TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + ❯ tests/stuck-run-triage.test.ts:62:18 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[40/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > refuses to send the Cloud bearer token to an unapproved origin +AssertionError: expected [Function] to throw error matching /refusing to send the Cloud bearer to…/\ but got 'expected an @relayflows/surface flow …' + +- Expected: +/refusing to send the Cloud bearer token to https:\/\/evil\.example/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[41/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > refuses a non-URL apiUrl +AssertionError: expected [Function] to throw error matching /is not a URL/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/is not a URL/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[42/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > allows an approved origin and uses it in the curl +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:77:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[43/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > defaults to production Cloud +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:82:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[44/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage apiUrl > never publishes a run record the fetch did not produce +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:89:34 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[45/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > names the Worker on every wrangler invocation +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:98:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[46/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > accepts caller-supplied Workers and rejects option-shaped ones +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:107:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[47/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > falls back when GNU timeout is absent, as it is on macOS +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:115:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[48/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > runs the tails concurrently so wall time does not scale with the batch +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:123:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[49/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage edge collection > records wrangler's own exit status rather than head's +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:129:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[50/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage shell text > parses under both sh and bash +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:137:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[51/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage shell text > collects tails with no GNU timeout on PATH, as on a stock macOS +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:157:33 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[52/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage agents > declares read-only permissions on every agent +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:176:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[53/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage agents > tells the forensics agents their evidence is untrusted +TypeError: expected an @relayflows/surface flow handle + ❯ Module.getFlowDefinition node_modules/@relayflows/surface/src/flow.ts:149:11 + ❯ drive tests/stuck-run-triage.test.ts:34:9 + 32| done: () => {}, + 33| }; + 34| await getFlowDefinition(triage).body(f as never… + | ^ + 35| return rec; + 36| } + ❯ tests/stuck-run-triage.test.ts:182:23 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[54/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > refuses a duplicate run id: two tails would share one evidence file +AssertionError: expected [Function] to throw error matching /duplicate runIds: c649fe14-0c2e-4e51-…/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/duplicate runIds: c649fe14-0c2e-4e51-9a6a-4f0d1b0f77aa/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[55/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > refuses a duplicate Worker name for the same reason +AssertionError: expected [Function] to throw error matching /duplicate workers: w-one/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/duplicate workers: w-one/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[56/63]⎯ + + FAIL tests/stuck-run-triage.test.ts > stuck-run-triage fan-out > bounds ids x workers, not just ids +AssertionError: expected [Function] to throw error matching /24 concurrent tails, over the 16/ but got 'expected an @relayflows/surface flow …' + +- Expected: +/24 concurrent tails, over the 16/ + ++ Received: +"expected an @relayflows/surface flow handle" + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[57/63]⎯ + +⎯⎯⎯⎯⎯⎯ Unhandled Errors ⎯⎯⎯⎯⎯⎯ + +Vitest caught 2 unhandled errors during the test run. +This might cause false positive tests. Resolve unhandled errors to make sure your tests are not affected. + +⎯⎯⎯⎯ Unhandled Rejection ⎯⎯⎯⎯⎯ +Error: hosted extension isolation requires Linux + ❯ unsupported src/hosted-extension-sandbox.ts:488:9 + 486| + 487| function unsupported(message: string): never { + 488| throw new PluginError('plugin_unsupported', message); + | ^ + 489| } + 490| + ❯ Module.runHostedExtensionSandbox src/hosted-extension-sandbox.ts:155:44 + ❯ Module.runVerifiedNativeExtensionSandbox src/hosted-extension-isolation.ts:173:16 + ❯ tests/hosted-extension-protocol.test.ts:454:17 + ❯ node_modules/@vitest/runner/dist/index.js:533:5 + ❯ runTest node_modules/@vitest/runner/dist/index.js:1056:11 + ❯ runSuite node_modules/@vitest/runner/dist/index.js:1205:15 + ❯ runSuite node_modules/@vitest/runner/dist/index.js:1205:15 + ❯ runFiles node_modules/@vitest/runner/dist/index.js:1262:5 + ❯ startTests node_modules/@vitest/runner/dist/index.js:1271:3 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { code: 'plugin_unsupported' } +This error originated in "tests/hosted-extension-protocol.test.ts" test file. It doesn't mean the error was thrown inside the file itself, but while it was running. +The latest test that might've caused the error is "rejects two forged calls after the authoritative first outcome settles". It might mean one of the following: +- The error was thrown, while Vitest was running this test. +- If the error occurred after the test had been completed, this was the last documented test before it was thrown. + +⎯⎯⎯⎯⎯ Uncaught Exception ⎯⎯⎯⎯⎯ +Error: spawn /Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd ENOENT + ❯ Process.ChildProcess._handle.onexit node:internal/child_process:285:19 + ❯ onErrorNT node:internal/child_process:483:16 + ❯ processTicksAndRejections node:internal/process/task_queues:89:21 + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ +Serialized Error: { errno: -2, code: 'ENOENT', syscall: 'spawn /Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd', path: '/Users/khaliqgant/Projects/AgentWorkforce/flows/kernel/target/release/relayflowd', spawnargs: [ '--data-dir', '/var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/flows-mcp-daemon-tMydRy', 'serve' ] } +This error originated in "tests/mcp.test.ts" test file. It doesn't mean the error was thrown inside the file itself, but while it was running. +The latest test that might've caused the error is "authored MCP effects against the real kernel". It might mean one of the following: +- The error was thrown, while Vitest was running this test. +- If the error occurred after the test had been completed, this was the last documented test before it was thrown. +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯ + + Test Files 9 failed | 4 passed (13) + Tests 60 failed | 114 passed | 64 skipped (238) + Errors 2 errors + Start at 22:19:35 + Duration 41.46s (transform 1.57s, setup 0ms, collect 6.50s, tests 110.51s, environment 1ms, prepare 578ms) + +exit=1