Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion docs/SURFACE.md
Original file line number Diff line number Diff line change
Expand Up @@ -399,7 +399,11 @@ authentication probes, and exact `flows.json` model allow-list as agent steps;
a declared `model` must be in that project's `models` array. A template call
such as ``await f.llm`Summarize ${text}` `` returns text. The structured overload
parses JSON and checks `output` before submitting a successful completion;
the kernel independently checks the schema before accepting the output.
the kernel independently checks the schema before accepting the output. A
reply that is exactly one markdown code fence around a value (three or more
backticks or tildes, closed by a run of the same character at least as long)
is judged by the value inside it; prose around the JSON, or two fenced values,
is still invalid.
Invalid JSON or a schema mismatch completes with `verification_failed` and
prevents downstream work. Retry and lease handling use the existing kernel
policies; this overload introduces no separate retry contract.
Expand Down
1,655 changes: 1,655 additions & 0 deletions evidence/llm-fence-predicate-resume/full-suite-on-branch.txt

Large diffs are not rendered by default.

58 changes: 58 additions & 0 deletions evidence/llm-fence-predicate-resume/mutation-1-fence.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
$ cp src/llm-worker.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts
exit=0
$ sed -i.bak 's/JSON.parse(unfenced(result.stdout_tail))/JSON.parse(result.stdout_tail)/' src/llm-worker.ts && rm src/llm-worker.ts.bak
exit=0
$ git diff --stat -- src/llm-worker.ts
packages/sdk/src/llm-worker.ts | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
exit=0
$ npx vitest run tests/worker-transcript.test.ts -t 'fenced reply'

RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk

❯ tests/worker-transcript.test.ts (8 tests | 1 failed | 7 skipped) 205ms
× the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value 205ms
→ expected 'verification_failed' to be 'success' // Object.is equality

⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯

FAIL tests/worker-transcript.test.ts > the llm worker judges the value inside one markdown fence > accepts a fenced reply whose content matches the schema, and journals the parsed value
AssertionError: expected 'verification_failed' to be 'success' // Object.is equality

Expected: "success"
Received: "verification_failed"

❯ tests/worker-transcript.test.ts:215:27
213| it('accepts a fenced reply whose content matches the schema, and jou…
214| const completion = await completeWith('```json\n{"x": 4}\n```');
215| expect(completion[4]).toBe('success');
| ^
216| expect((completion[5] as { output: unknown }).output).toEqual({ x:…
217| });

⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/1]⎯

Test Files 1 failed (1)
Tests 1 failed | 7 skipped (8)
Start at 22:18:46
Duration 1.36s (transform 256ms, setup 0ms, collect 901ms, tests 205ms, environment 0ms, prepare 49ms)

exit=1
$ cp /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts src/llm-worker.ts
exit=0
$ cmp src/llm-worker.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/llm-worker.ts
exit=0
$ git diff --stat -- src/llm-worker.ts
exit=0
$ npx vitest run tests/worker-transcript.test.ts -t 'fenced reply'

RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk

✓ tests/worker-transcript.test.ts (8 tests | 7 skipped) 195ms

Test Files 1 passed (1)
Tests 1 passed | 7 skipped (8)
Start at 22:18:48
Duration 901ms (transform 206ms, setup 0ms, collect 461ms, tests 195ms, environment 0ms, prepare 43ms)

exit=0
58 changes: 58 additions & 0 deletions evidence/llm-fence-predicate-resume/mutation-2-predicate.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
$ cp src/authored-flow-executor.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts
exit=0
$ sed -i.bak 's/JSON.stringify(canonical)/JSON.stringify(record)/' src/authored-flow-executor.ts && rm src/authored-flow-executor.ts.bak
exit=0
$ git diff --stat -- src/authored-flow-executor.ts
packages/sdk/src/authored-flow-executor.ts | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
exit=0
$ npx vitest run tests/authored-agent-artifacts.test.ts -t 'sorted keys'

RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk

❯ tests/authored-agent-artifacts.test.ts (5 tests | 1 failed | 4 skipped) 42ms
× predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run 41ms
→ expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality

⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯

FAIL tests/authored-agent-artifacts.test.ts > predicate verdicts recorded on the root run > a resumed verdict read back with sorted keys lowers the same gate command as the first run
AssertionError: expected 'printf \'%s\' \'{"because":"plan cove…' to be 'printf \'%s\' \'{"gate":"predicate","…' // Object.is equality

Expected: "printf '%s' '{"gate":"predicate","step":"run-1","verdict":"pass","because":"plan covers every question"}'"
Received: "printf '%s' '{"because":"plan covers every question","gate":"predicate","step":"run-1","verdict":"pass"}'"

❯ tests/authored-agent-artifacts.test.ts:300:25
298| expect(stream).toHaveLength(1);
299| expect(commands).toHaveLength(2);
300| expect(commands[1]).toBe(commands[0]);
| ^
301| expect(commands[0]).toContain('{"gate":"predicate","step":"run-1",…
302| });

⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/1]⎯

Test Files 1 failed (1)
Tests 1 failed | 4 skipped (5)
Start at 22:18:50
Duration 1.56s (transform 573ms, setup 0ms, collect 1.26s, tests 42ms, environment 0ms, prepare 66ms)

exit=1
$ cp /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts src/authored-flow-executor.ts
exit=0
$ cmp src/authored-flow-executor.ts /var/folders/6d/0x5fkt8d01gfmmjdzkxqzwnh0000gn/T/tmp.oIUGoRyujb/authored-flow-executor.ts
exit=0
$ git diff --stat -- src/authored-flow-executor.ts
exit=0
$ npx vitest run tests/authored-agent-artifacts.test.ts -t 'sorted keys'

RUN v2.1.9 /Users/khaliqgant/Projects/AgentWorkforce/flows/packages/sdk

✓ tests/authored-agent-artifacts.test.ts (5 tests | 4 skipped) 33ms

Test Files 1 passed (1)
Tests 1 passed | 4 skipped (5)
Start at 22:18:52
Duration 982ms (transform 371ms, setup 0ms, collect 736ms, tests 33ms, environment 0ms, prepare 57ms)

exit=0
Loading
Loading