test(server): record Grok replay fixtures from a live grok agent - #13537
Conversation
Adds scripts/record-grok-acp-replay-fixture.ts. It runs a fixture's own scenario through the real orchestrator and GrokAdapterV2 against a live `grok agent stdio`, tees the raw ACP lines from the runtime's protocol logger, and writes them as a replay transcript. Grok replay now wraps its runtime the way production does (Ctrl+C cancel metadata, the x.ai prompt-completion race), so replay checks the `_meta.promptId` T3 actually sends on session/prompt. The synthetic transcripts that are not re-recorded gain that field. The replay agent reads large transcripts from a file, since recorded ones exceed the 128 KiB limit on one environment variable, and materializes <workspace> in inbound frames as well as expectations. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Re-records simple, multi_turn, queued_turn, todo_list, message_steering and turn_interrupt from Grok 1.0.41 with the new recorder. The ACP registry variants replayed the Grok files; they keep the previous synthetic frames in their own registry_transcript.ndjson, since real Grok frames describe Grok, not a generic registry agent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting). This review would cost an estimated $38.53, which exceeds your per-review limit of $15.00. The top 3 files driving up this estimate:
Tip To get this pull request reviewed, you can:
|
| import { layer as idAllocatorLayer, IdAllocatorV2 } from "../src/orchestration-v2/IdAllocator.ts"; | ||
| import { makeLayerEffect as makeProviderAdapterRegistryLayerEffect } from "../src/orchestration-v2/ProviderAdapterRegistry.ts"; |
There was a problem hiding this comment.
These service-module imports rename layer and makeLayerEffect, hiding the modules' public service boundaries. Could you import IdAllocator and ProviderAdapterRegistry as namespaces and use IdAllocator.layer and ProviderAdapterRegistry.makeLayerEffect at the call sites?
Posted via Macroscope — Effect Service Conventions
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — The PR preserves existing Grok production cancellation behavior and mainly updates replay infrastructure and fixtures, but it also introduces a substantial live-agent recording workflow with protocol parsing, normalization, and integration logic. That scope is broader than a simple test adjustment and merits human review. Not approved because:
Review your spending limits in Billing settings, or comment |
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Every committed Grok replay transcript was synthetic (
live-grok-shape-probe,protocol-semantic-fixture,compiled-provider-log), and there was no way to record a real one. Several of those frames don't match what Grok 1.0.41 actually sends, so the replay suite was checking our assumptions about Grok, not Grok itself.What changed
Recorder:
apps/server/scripts/record-grok-acp-replay-fixture.ts(vp run record:grok-replay -- --scenario simple,multi_turn)buildInput(), materializes it into exactly the commands and steps the replay test dispatches, and runs them through the real orchestrator plusmakeGrokAdapterV2against a livegrok agent stdioprocess. Only the runtime'sprotocolLoggingis swapped, for a tee of raw ACP lines. Outbound frames are therefore byte-for-byte what T3 sends, inbound frames are what Grok answered, and the live run uses the same TestClock and seeded ids as replay, so dispatch order matches too.assertOutputand prints any failure. This is how the assertion mismatches below surfaced.<workspace>, including Grok's URL-encoded session directory./home/grok-replayand the recording user becomesgrok-replay.initializeparams become<any>.scope: "user") skills, host/agent ids, account settings and announcement broadcasts are dropped, along with responses to Grok-internal request ids ("skills-reload") that T3's protocol already discards.generatedBy: "live-grok-recorder",grokVersion, and what was normalized.Replay plumbing
makeGrokAcpRuntimedoes: Ctrl+CcancelMetaplusmakeXAiPromptCompletionRuntime. Replay therefore checks the_meta.promptIdT3 really sends onsession/promptand exercises theturn_completed/_x.ai/session/prompt_completesettlement race. The four synthetic transcripts that were not re-recorded gain that_metafield and nothing else.acp-replay-agentreads the transcript from a file. A recorded transcript exceeds Linux's 128 KiB cap on a single environment variable (spawn E2BIG). The agent also materializes<workspace>in inbound frames (tool inputs,fs/*requests), matching how expectations already work.registry_transcript.ndjson, because real Grok frames (x.ai extensions,_meta.promptId) describe Grok, not a generic registry agent.Re-recorded from Grok 1.0.41:
simple,multi_turn,queued_turn,todo_list,message_steering,turn_interrupt. All existingoutput.tsassertions pass unchanged against the live frames.Not re-recorded, and why
tool_call_read_only: the prompt names/tmp/claude-replay-tool_call_read_only/*, but the fixture never creates those files (input.tshas noworkspaceFiles, and the recorded workspace is a fresh temp dir). Live, Grok'sfs/read_text_filegets "Could not read text file", it falls back to a shelllsthat the read-only policy rejects, and the turn endscancelled. The assertion is right and the fixture input is incomplete. Fix: seed the two files at a workspace-relative path and point the prompt there (this touches the Claude and Cursor variants that share the prompt). Follow-up.tool_call_read_only_on_request: live, Grok uses itswritetool (kind: "edit"), so the approval is afile_change. The sharedassertToolCallReadOnlyOnRequestOutputrequirescommand_executionand acommandrequest, which only holds when the agent picks the shell. The adapter is correct; the assertion is Codex-shaped. Follow-up: a Grok-specific output assertion, then record.plan_questions,grok_subagent_lineage: not attempted in this PR.plan_questionsneeds thex.ai/ask_user_questionround trip, andgrok_subagent_lineagecame from a real session log, not a straightforward scenario.Real Grok frames the synthetic transcripts didn't have
initializeanswersprotocolVersion: 1to our v2 request, as expected, and carries the model list and available commands in_meta._x.ai/session_notification {sessionUpdate: "turn_completed", prompt_id}and_x.ai/session/prompt_complete {promptId}before thesession/promptresponse. Both echo T3's_meta.promptId(t3-xai-prompt-N).fs/read_text_fileafterpending_interaction/interaction_resolvednotifications. Tool calls start astitle: "read_file"with_meta["x.ai/tool"]and are later renamed bytool_call_update.session/cancel→turn_completed {stop_reason: "cancelled"}→prompt_complete {cancelTrigger: "ctrl_c", cancellationCategory: "MidTurnAbort"}→session/prompt {stopReason: "cancelled"}.turn_interrupt, the user Stop hard-kills the process group (restartRuntimeAfterInterrupt), so the recording ends aftersession/promptwith nosession/cancelon the wire (runtime_exit: cancelled).Verification
cd apps/server && vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts src/provider/acp/GrokAcpSupport.test.ts: 104 passed, including every*/grokand*/acpRegistryreplay (19).simplewas recorded again after the last recorder change.cd apps/server && vp exec tsc --noEmit -p .: 0error TS/warning TS.vp run knip:check: clean.vp lintandvp fmton the touched files: clean.Model: Claude Opus 5.5 (Claude Code)
🤖 Generated with Claude Code