Skip to content

refactor(server): record Cursor replay fixtures with Effect - #13556

Merged
juliusmarminge merged 2 commits into
t3code/codex-turn-mappingfrom
v2/effect-cursor-recorder
Sep 25, 2026
Merged

juliusmarminge merged 2 commits into
t3code/codex-turn-mappingfrom
v2/effect-cursor-recorder

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Maintainer question: "Why is the Cursor testkit async-based and not Effect?" The Cursor replay recorder was an async function. It called @cursor/sdk directly, used hand-rolled Promise signals and a Promise.race timeout, and rebuilt every protocol frame by hand. The adapter talks to the SDK through CursorAgentSdkRunner, so recording and production could drift.

What changed

The recorder now uses the runner. CursorAgentSdk.ts builds the live runner with makeCursorAgentSdkRunner(protocolLoggerFor). The live layer passes the native event-log writer, as before. The recorder passes a logger that appends each frame (agent.open, agent.opened, run.start, run.started, interaction.update, run.completed, run.cancel, agent.close, and resume) to the transcript. The SDK calls, the pending-update buffering before run.started, and the treatment of an AbortError from run.cancel as success now come from the same code the adapter runs. The runner change is a pure extraction: git diff -w on that file is +27/-14.

The recorder is an Effect (CursorAdapterV2.testkit.ts, now Effect.fn):

  • Deferred replaces the Promise signals: first update for the run-start interrupt, first tool-call-started for the mid-tool interrupt.
  • Effect.timeoutOrElse replaces the 30 s Promise.race.
  • In a mid-tool interrupt recording, the recording logger buffers every frame after the first tool-call-started and appends them right after run.cancel, so the cancel always directly follows its trigger. It appends instead of waiting, because updates that arrive before send returns are flushed inside send, and waiting there would never return. When tool-call-started arrives after send returns, the SDK callback also pauses until the cancel is sent, as the async recorder did (see the crash note below).
  • Effect.acquireUseRelease opens and closes each agent, so agent.close runs (and is recorded) on every exit, including the restart before a resumed prompt.
  • Invalid inputs fail with a tagged CursorReplayRecordingError instead of throwing.
  • The 10 ms cancel deferral is now Effect.sleep("10 millis"), with its comment unchanged. It works around the unhandled AbortError inside @cursor/sdk when cancelling synchronously from the tool-call-started callback.

The script is an Effect CLI command (record-cursor-agent-sdk-replay-fixture.ts), following migrate-dev-db.ts and t3-sqlite-state.ts:

  • It runs Command.run with NodeServices.layer and NodeRuntime.runMain.
  • --scenario is a Flag.Literals over the recording names, falling back to T3_CURSOR_REPLAY_SCENARIO. --out is optional.
  • CURSOR_API_KEY is read as Config.Redacted. T3_CURSOR_REPLAY_MODEL and T3_CURSOR_REPLAY_CWD are read through Config.
  • The recording workspace is a scoped checkpointWorkspace, so it is removed on every exit. The runFileSystem promise bridge is gone.

Docs: docs/user/cursor.md and docs/orchestration-v2/testing-strategy.md drop the -- in pnpm --filter t3 record:cursor-replay -- --scenario …. The new argument parsing requires it: pnpm 11 forwards the -- literally (checked: argv becomes ["--","--scenario","simple"]), and the Effect CLI treats everything after -- as trailing operands. With the --, the command fails with Missing required flag: --scenario. Without it, pnpm passes the flags straight through.

Line delta: +545/-576 across 5 files. Most of that is reindentation from the runner extraction; ignoring whitespace it is +316/-347.

Verification

  • Committed Cursor transcripts are untouched (no fixture files in the diff).

  • vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts -t cursor: 10 passed.

  • vp test run on OrchestratorReplayRecovery.integration.test.ts, OrchestratorReplayFixtures.contract.test.ts, CursorAgentSdk.test.ts, CursorAdapterV2.test.ts: 18 passed. Before rebasing onto the current base, CursorAdapterV2.testkit.test.ts and cursorReplayRecordingWorkspace.test.ts also passed (31 total); test(server): remove replay harness self-tests #13531 has since removed both files.

  • vp exec tsc --noEmit -p . in apps/server: no error TS or warning TS.

  • vp run knip:check: clean at the first push. On the current base it reports one unused export, THREAD_DETAILS_PANEL_SPLIT_BUTTON_SURFACE_CLASS in apps/web, added by 99d76f2 on the base branch; this PR touches no web files. vp lint on the three touched TS files: clean.

  • Live, with composer-2.5, recorded to a scratch dir outside the repo and not committed. The same scenarios were recorded with the old recorder (at the pre-change commit) and the new one, and compared in two ways:

    1. Frame sequence, collapsing runs of streaming thinking/token/text-delta updates whose count depends on model output: identical for all four scenarios.
    2. Every non-update entry (header, agent.open options, run.start message and options, run.cancel, run.completed status and keys, runtime_exit, agent.close, labels), serialized byte-for-byte with ids, durationMs, result, and usage masked: identical.
    Scenario Path exercised Old / new entries Match
    simple single run 17 / 18 yes
    turn_interrupt_mid_tool cancel after tool-call-started (10 ms deferral, gate) 42 / 34 yes: tool-call-started, run.cancel, run.completed (cancelled), agent.close
    message_steering cancel after first update 23 / 25 yes: thinking-delta, run.cancel, run.completed (cancelled), then run 2
    provider_thread_resume close, resume, second prompt 61 / 69 yes, including agent.close:before-prompt-2, agent.resume:before-prompt-2, agent.resumed:before-prompt-2

    The entry-count differences are only in how many streaming deltas the model produced. tool_call_read_only was also recorded with the new recorder. Its non-update entries match the committed fixture, including the prompt rewritten back to the fixture path, and nothing from the recording host's paths leaks into the transcript. As a stronger check, the four new recordings were copied over the committed fixtures and OrchestratorReplayFixtures -t cursor passed 10/10. The fixtures were then restored.

  • Review follow-up (updates flushed inside send): a throwaway test with a mocked SDK whose send delivers tool-call-started then text-delta before returning failed on the first commit (text-delta was recorded before run.cancel:1) and passes with the fix (order tool-call-started, run.cancel:1, text-delta). The test was not committed, since the maintainer prefers no tests for the testkit. After the fix, turn_interrupt_mid_tool was re-recorded live: the transcript ends tool-call-started, run.cancel:1, run.completed:1 (cancelled), agent.close.

  • Crash note: turn_interrupt_mid_tool recording sometimes dies with the SDK's unhandled AbortError whatever recorder runs it. It failed 1 of 3 runs with the pre-PR async recorder, 1 of 3 with this PR's first commit, and 2 of 6 with the fix. A variant that did not pause the SDK callback failed all 3 runs, which is why the pause stays. Re-run on failure; the 10 ms deferral reduces the crash rate but does not remove it.

  • Diff grepped for crsr_: none.

Not run: repo-wide checks, and live re-recording of the other seven Cursor scenarios (they use the same single-run path as simple).

Recording note for anyone re-running this on a Linux host with pnpm: apps/server/node_modules/@cursor/ only links sdk, not the platform package @cursor/sdk-linux-x64. The SDK looks for its cursorsandbox helper next to the script, so any sandboxed scenario (turn_interrupt_mid_tool, the read-only ones) fails with "sandboxing is not supported in this environment", under both the old and the new recorder. I linked the platform package locally to record; nothing about that is in this PR.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XL 500-999 changed lines (additions + deletions). labels Sep 25, 2026
Comment thread apps/server/src/orchestration-v2/Adapters/CursorAdapterV2.testkit.ts Outdated
export const recordCursorAgentSdkReplayTranscript = Effect.fn(
"recordCursorAgentSdkReplayTranscript",
)(function* (input: CursorReplayRecordingInput) {
const invalid = (reason: string) =>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

invalid only forwards its arguments to new CursorReplayRecordingError, which adds an unnecessary error-construction helper. Consider constructing the tagged error at each validation or timeout boundary instead. This requires edits at multiple call sites, so there is no single-hunk suggestion.

Posted via Macroscope — Effect Service Conventions

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: d5c9e25 · Source CI: failure

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This PR substantially restructures the shared production Cursor SDK runner while also rewriting replay cancellation and buffering logic. An unresolved Medium finding reports a concrete transcript-ordering risk in the new mid-tool interruption path; the remaining comment is stylistic.

Adjust the Minimum Blocking Severity for this repo — including turning it Off — in Settings. You can add or adjust custom eligibility rules. Learn more.

juliusmarminge and others added 2 commits September 24, 2026 21:30
The Cursor recorder was an async function that called @cursor/sdk
directly, with hand-rolled Promise signals and a Promise.race timeout,
while the adapter reaches the SDK through CursorAgentSdkRunner. The two
could drift.

The live runner is now built by makeCursorAgentSdkRunner, which takes the
protocol logger per opened agent. The live layer passes the native event
log writer; the recorder passes a logger that appends each frame to the
transcript. Recording therefore drives the exact code path the adapter
uses, and agent.open/run.start/run.cancel/agent.close frames come from
the runner instead of being rebuilt by hand.

The recorder is an Effect.fn: Deferred for the interrupt triggers,
Effect.timeoutOrElse for the 30 s waits, a Latch to hold updates between
tool-call-started and run.cancel, acquireUseRelease so each agent is
closed on every exit, a tagged CursorReplayRecordingError for invalid
input, and Effect.sleep for the 10 ms cancel deferral.

The script is an Effect CLI command run with NodeRuntime.runMain. Its
scenario flag falls back to T3_CURSOR_REPLAY_SCENARIO, and the recording
workspace is scoped. The docs drop the `--` before `--scenario`: pnpm 11
forwards it literally and the Effect CLI treats everything after it as
operands.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The recorder held the transcript with a latch that it could only close
after send returned, because updates that arrive earlier are flushed
inside send and waiting there never returns. So when tool-call-started
came in that early batch, the updates after it were recorded before
run.cancel.

The transcript is now ordered by buffering instead of waiting: after the
first tool-call-started, frames are held and appended right after
run.cancel is recorded. The SDK callback still pauses until the cancel is
sent when send has already returned, as the async recorder did; without
that pause all three live probes crashed with the SDK's AbortError.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@juliusmarminge
juliusmarminge force-pushed the v2/effect-cursor-recorder branch from d874291 to d5c9e25 Compare September 25, 2026 04:33
@juliusmarminge
juliusmarminge merged commit 29adfb3 into t3code/codex-turn-mapping Sep 25, 2026
23 of 24 checks passed
@juliusmarminge
juliusmarminge deleted the v2/effect-cursor-recorder branch September 25, 2026 05:09
juliusmarminge added a commit that referenced this pull request Sep 25, 2026
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL 500-999 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant