Skip to content

test(server): record Grok replay fixtures from a live grok agent - #13537

Merged
juliusmarminge merged 3 commits into
t3code/codex-turn-mappingfrom
v2/grok-recorder
Sep 25, 2026
Merged

juliusmarminge merged 3 commits into
t3code/codex-turn-mappingfrom
v2/grok-recorder

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Every committed Grok replay transcript was synthetic (live-grok-shape-probe, protocol-semantic-fixture, compiled-provider-log), and there was no way to record a real one. Several of those frames don't match what Grok 1.0.41 actually sends, so the replay suite was checking our assumptions about Grok, not Grok itself.

What changed

Recorder: apps/server/scripts/record-grok-acp-replay-fixture.ts (vp run record:grok-replay -- --scenario simple,multi_turn)

  • It drives the real adapter. It takes the fixture's own buildInput(), materializes it into exactly the commands and steps the replay test dispatches, and runs them through the real orchestrator plus makeGrokAdapterV2 against a live grok agent stdio process. Only the runtime's protocolLogging is swapped, for a tee of raw ACP lines. Outbound frames are therefore byte-for-byte what T3 sends, inbound frames are what Grok answered, and the live run uses the same TestClock and seeded ids as replay, so dispatch order matches too.
  • It checks the live run against the fixture's assertOutput and prints any failure. This is how the assertion mismatches below surfaced.
  • Normalization:
    • Session ids become fixed UUIDs and the workspace becomes <workspace>, including Grok's URL-encoded session directory.
    • HOME becomes /home/grok-replay and the recording user becomes grok-replay.
    • T3-owned prompt text, MCP servers and initialize params become <any>.
    • Personal (scope: "user") skills, host/agent ids, account settings and announcement broadcasts are dropped, along with responses to Grok-internal request ids ("skills-reload") that T3's protocol already discards.
    • The header records generatedBy: "live-grok-recorder", grokVersion, and what was normalized.

Replay plumbing

  • The Grok replay harness now wraps its runtime the way makeGrokAcpRuntime does: Ctrl+C cancelMeta plus makeXAiPromptCompletionRuntime. Replay therefore checks the _meta.promptId T3 really sends on session/prompt and exercises the turn_completed / _x.ai/session/prompt_complete settlement race. The four synthetic transcripts that were not re-recorded gain that _meta field and nothing else.
  • acp-replay-agent reads the transcript from a file. A recorded transcript exceeds Linux's 128 KiB cap on a single environment variable (spawn E2BIG). The agent also materializes <workspace> in inbound frames (tool inputs, fs/* requests), matching how expectations already work.
  • The ACP registry variants used to replay the Grok files. They now keep the previous synthetic frames in their own registry_transcript.ndjson, because real Grok frames (x.ai extensions, _meta.promptId) describe Grok, not a generic registry agent.

Re-recorded from Grok 1.0.41: simple, multi_turn, queued_turn, todo_list, message_steering, turn_interrupt. All existing output.ts assertions pass unchanged against the live frames.

Not re-recorded, and why

  • tool_call_read_only: the prompt names /tmp/claude-replay-tool_call_read_only/*, but the fixture never creates those files (input.ts has no workspaceFiles, and the recorded workspace is a fresh temp dir). Live, Grok's fs/read_text_file gets "Could not read text file", it falls back to a shell ls that the read-only policy rejects, and the turn ends cancelled. The assertion is right and the fixture input is incomplete. Fix: seed the two files at a workspace-relative path and point the prompt there (this touches the Claude and Cursor variants that share the prompt). Follow-up.
  • tool_call_read_only_on_request: live, Grok uses its write tool (kind: "edit"), so the approval is a file_change. The shared assertToolCallReadOnlyOnRequestOutput requires command_execution and a command request, which only holds when the agent picks the shell. The adapter is correct; the assertion is Codex-shaped. Follow-up: a Grok-specific output assertion, then record.
  • plan_questions, grok_subagent_lineage: not attempted in this PR. plan_questions needs the x.ai/ask_user_question round trip, and grok_subagent_lineage came from a real session log, not a straightforward scenario.

Real Grok frames the synthetic transcripts didn't have

  • initialize answers protocolVersion: 1 to our v2 request, as expected, and carries the model list and available commands in _meta.
  • A turn settles through _x.ai/session_notification {sessionUpdate: "turn_completed", prompt_id} and _x.ai/session/prompt_complete {promptId} before the session/prompt response. Both echo T3's _meta.promptId (t3-xai-prompt-N).
  • Reads go through client-mediated fs/read_text_file after pending_interaction / interaction_resolved notifications. Tool calls start as title: "read_file" with _meta["x.ai/tool"] and are later renamed by tool_call_update.
  • session/cancel → turn_completed {stop_reason: "cancelled"} → prompt_complete {cancelTrigger: "ctrl_c", cancellationCategory: "MidTurnAbort"} → session/prompt {stopReason: "cancelled"}.
  • For turn_interrupt, the user Stop hard-kills the process group (restartRuntimeAfterInterrupt), so the recording ends after session/prompt with no session/cancel on the wire (runtime_exit: cancelled).

Verification

  • cd apps/server && vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts src/provider/acp/GrokAcpSupport.test.ts: 104 passed, including every */grok and */acpRegistry replay (19).
  • Live recordings: every re-recorded scenario was recorded against Grok 1.0.41 on Linux, and the live orchestration passed its fixture assertions before the transcript was written. simple was recorded again after the last recorder change.
  • cd apps/server && vp exec tsc --noEmit -p .: 0 error TS / warning TS.
  • vp run knip:check: clean. vp lint and vp fmt on the touched files: clean.
  • Not run: repo-wide test, typecheck or lint.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

juliusmarminge and others added 2 commits September 24, 2026 18:55
Adds scripts/record-grok-acp-replay-fixture.ts. It runs a fixture's own
scenario through the real orchestrator and GrokAdapterV2 against a live
`grok agent stdio`, tees the raw ACP lines from the runtime's protocol
logger, and writes them as a replay transcript.

Grok replay now wraps its runtime the way production does (Ctrl+C cancel
metadata, the x.ai prompt-completion race), so replay checks the
`_meta.promptId` T3 actually sends on session/prompt. The synthetic
transcripts that are not re-recorded gain that field. The replay agent
reads large transcripts from a file, since recorded ones exceed the
128 KiB limit on one environment variable, and materializes <workspace>
in inbound frames as well as expectations.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Re-records simple, multi_turn, queued_turn, todo_list, message_steering
and turn_interrupt from Grok 1.0.41 with the new recorder. The ACP
registry variants replayed the Grok files; they keep the previous
synthetic frames in their own registry_transcript.ndjson, since real
Grok frames describe Grok, not a generic registry agent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting).

This review would cost an estimated $38.53, which exceeds your per-review limit of $15.00.

The top 3 files driving up this estimate:

File Diff Size Estimate
apps/server/src/orchestration-v2/testkit/fixtures/todo_list/grok_transcript.ndjson 291.35KB $14.57
apps/server/src/orchestration-v2/testkit/fixtures/multi_turn/grok_transcript.ndjson 103.79KB $5.19
apps/server/src/orchestration-v2/testkit/fixtures/queued_turn/grok_transcript.ndjson 100.43KB $5.02

Tip

To get this pull request reviewed, you can:

  1. Comment @macroscope-app on this PR to request a manual review (monthly spend limits still apply).
  2. Exclude the file(s) above from review by adding a pattern to your .macroscope/ignore.md — note that creating this file replaces Macroscope's built-in default ignores rather than extending them.
  3. Raise your cost limit in your workspace billing settings.

Turn off this reminder going forward

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Sep 25, 2026
Comment on lines +33 to +34
import { layer as idAllocatorLayer, IdAllocatorV2 } from "../src/orchestration-v2/IdAllocator.ts";
import { makeLayerEffect as makeProviderAdapterRegistryLayerEffect } from "../src/orchestration-v2/ProviderAdapterRegistry.ts";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These service-module imports rename layer and makeLayerEffect, hiding the modules' public service boundaries. Could you import IdAllocator and ProviderAdapterRegistry as namespaces and use IdAllocator.layer and ProviderAdapterRegistry.makeLayerEffect at the call sites?

Posted via Macroscope — Effect Service Conventions

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.7 KiB — 29.3 KiB ✅
Claude Live turn messages — 1 — 8 ✅

Baseline: unavailable · PR result: 42c26da · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — The PR preserves existing Grok production cancellation behavior and mainly updates replay infrastructure and fixtures, but it also introduces a substantial live-agent recording workflow with protocol parsing, normalization, and integration logic. That scope is broader than a simple test adjustment and merits human review.

Not approved because:

  • Per-review cost limit exceeded (workspace setting). Approvability relies on correctness review in order to determine eligibility

Review your spending limits in Billing settings, or comment @macroscope-app review this PR to bypass the limit and review now. You can add or adjust custom eligibility rules. Learn more.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@juliusmarminge
juliusmarminge merged commit 8d09ecb into t3code/codex-turn-mapping Sep 25, 2026
24 checks passed
@juliusmarminge
juliusmarminge deleted the v2/grok-recorder branch September 25, 2026 05:26
juliusmarminge added a commit that referenced this pull request Sep 25, 2026
)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant