Skip to content

test(server): re-record Cursor replay fixtures and cover skills live - #13493

Merged
juliusmarminge merged 5 commits into
t3code/codex-turn-mappingfrom
v2/cursor-rerecord
Sep 24, 2026
Merged

juliusmarminge merged 5 commits into
t3code/codex-turn-mappingfrom
v2/cursor-rerecord

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Builds on #13499 (merged), which makes Cursor V2 pass local.settingSources. The transcripts here carry that field in agent.open.

The Cursor replay transcripts were recorded in June against @cursor/sdk 1.0.19; the server now runs 1.0.31. They predate the thinking-delta stream, run usage, and today's tool-call shapes, so replays were checking behavior Cursor no longer shows. A skill unit test also used a hand-built runner where a live fixture can cover the same path.

What changed

Recorder (first commit). The recorder opened every agent with full access, but five fixtures replay with a read-only or workspace-write policy. Replay matches the agent.open frame exactly, so any re-recording of those fixtures could not replay. The recorder now takes each scenario's runtime-policy override. It also:

  • sanitizes workspace paths in object keys (grep results are keyed by path);
  • maps the recording workspace's parent to /tmp;
  • waits one timer tick before a mid-tool cancel. Cancelling inside the SDK's tool-call-started callback leaves an unhandled AbortError in @cursor/sdk that kills the recorder process (see below).

Re-recorded fixtures (second commit). Nine of the ten Cursor transcripts were re-recorded live with composer-2.5: simple, multi_turn, message_steering, provider_thread_resume, queued_turn, todo_list, subagent, tool_call_read_only, and turn_interrupt_mid_tool. proposed_plan was left as is. In an empty workspace, the plan-mode agent searches outside it (/tmp, the recording host's cache and worktree) and quotes those files in its plan, so the transcript would publish host paths. Three output assertions pinned the June wording and item order and now match today's run:

  • subagent: a reasoning segment now follows the subagents.
  • todo_list: the progress lines are reworded and there are more reasoning segments.
  • tool_call_read_only: the progress line is compared after trimming, because trailing newlines vary between runs. Grok's transcript shares this assertion.

None of these differences was an adapter bug.

Skill fixture (third commit). The new skill_invocation fixture seeds .cursor/skills/review/SKILL.md and sends $review README.md. Replay matches run.start exactly, so the adapter's rewrite to /review is now proven through the full orchestrator. It was recorded with #13499's settingSources in place, so the SDK loads the workspace skill natively: the model reads SKILL.md straight from the /review invocation, with no workspace search, and answers from it. This replaces the unit test "sends discovered skills as native slash invocations". Fixture inputs gained workspaceFiles, which the recorder and the replay workspace both commit.

The task-lifecycle unit test stays, because a live run can't be made to end without a task completion. Its frames now follow the recorded shape: a partial-tool-call before tool-call-started, and subagentType: {kind: "unspecified"} with agentId and mode.

Not done

The proposed tool_call_ls_lints fixture, which would have replaced the ls/readLints projection unit test, can't be produced live with this SDK:

  • The local SDK's diagnostics provider is a stub that always returns zero diagnostics. Diagnostics come from the IDE, so readLints can never return errors locally.
  • No model advertises an ls tool: composer-2.5, gpt-5.4, and claude-sonnet-4-6 all list directories through Shell or Glob. No Cursor transcript in the repo has ever recorded ls.

The unit test stays.

Verification

  • vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts src/orchestration-v2/testkit/OrchestratorReplayRecovery.integration.test.ts src/orchestration-v2/testkit/OrchestratorReplayFixtures.contract.test.ts src/orchestration-v2/Adapters/CursorAdapterV2.test.ts src/orchestration-v2/Adapters/CursorAdapterV2.testkit.test.ts scripts/cursorReplayRecordingWorkspace.test.ts: 102 passed, which includes all 75 replay fixtures across providers.
  • Negative check: seeding the skill at a non-skill path makes skill_invocation fail with a run.start frame mismatch.
  • Hermeticity: Cursor replays now run with an empty HOME. Before, with HOME pointing at a directory holding a user-level review skill, skill_invocation still passed when the workspace skill was moved to a non-skill path. It now fails with a run.start mismatch, and the seeded fixture still passes under that HOME. The Cursor replay and recovery suites pass (12), and so do the adapter and contract tests (22).
  • vp exec tsc --noEmit -p . in apps/server: no errors. vp lint on touched files: one warning that already exists on the base branch.
  • Transcripts checked for API keys, auth headers, and host paths: none.
  • After rebasing onto fix(server): Cursor V2 threads load project skills and rules #13499: all recorded transcripts carry fix(server): Cursor V2 threads load project skills and rules #13499's settingSources field. The first nine gained only that field, inserted byte-identically to fix(server): Cursor V2 threads load project skills and rules #13499's edit. skill_invocation was re-recorded live with the fix. The Cursor replay and recovery suites pass (12); the adapter, testkit, contract, and recorder-workspace tests pass (26). The hermeticity and negative checks were re-run and hold. Server tsc shows no TS errors or warnings. vp run knip:check passes (the earlier unused-export failure was fixed on the base by chore(web): drop the right panel sheet class left unused by the V2 rebase #13502).
  • Not run: repo-wide checks.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Sep 24, 2026
@macroscopeapp

macroscopeapp Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting).

This review would cost an estimated $19.76, which exceeds your per-review limit of $15.00.

The top 3 files driving up this estimate:

File Diff Size Estimate
apps/server/src/orchestration-v2/testkit/fixtures/subagent/cursor_transcript.ndjson 197.24KB $9.86
apps/server/src/orchestration-v2/testkit/fixtures/todo_list/cursor_transcript.ndjson 69.35KB $3.47
apps/server/src/orchestration-v2/testkit/fixtures/provider_thread_resume/cursor_transcript.ndjson 25.00KB $1.25

Tip

To get this pull request reviewed, you can:

  1. Comment @macroscope-app on this PR to request a manual review (monthly spend limits still apply).
  2. Exclude the file(s) above from review by adding a pattern to your .macroscope/ignore.md — note that creating this file replaces Macroscope's built-in default ignores rather than extending them.
  3. Raise your cost limit in your workspace billing settings.

Turn off this reminder going forward

@macroscopeapp

macroscopeapp Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Not approved

Macroscope's review found this PR not approvable — This is primarily a test-only Cursor fixture refresh with no product-path or default-behavior change. Human review is warranted for the unresolved temporary HOME directory lifecycle issue in the replay layer, which may prevent scoped cleanup.

Not approved because:

  • Per-review cost limit exceeded (workspace setting). Approvability relies on correctness review in order to determine eligibility

Review your spending limits in Billing settings, or comment @macroscope-app review this PR to bypass the limit and review now. You can add or adjust custom eligibility rules. Learn more.

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: cd3ec2f · Source CI: failure

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

).pipe(Layer.provide(NodeServices.layer));
// Skill discovery also scans user roots under HOME; an empty HOME keeps
// replays from picking up the host's own skills.
const hostEnvironmentLayer = Layer.effect(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This layer acquires a scoped temporary directory, so it should use the scoped layer constructor to keep the directory alive for the layer's lifetime and run its finalizer when the layer scope closes.

Suggested change
const hostEnvironmentLayer = Layer.effect(
const hostEnvironmentLayer = Layer.scoped(

Posted via Macroscope — Effect Service Conventions

@juliusmarminge
juliusmarminge force-pushed the t3code/codex-turn-mapping branch from 9fc8676 to ae065ff Compare September 24, 2026 21:34
@juliusmarminge
juliusmarminge changed the base branch from t3code/codex-turn-mapping to v2/cursor-setting-sources September 24, 2026 21:45
Base automatically changed from v2/cursor-setting-sources to t3code/codex-turn-mapping September 24, 2026 21:47
juliusmarminge and others added 5 commits September 24, 2026 14:48
The Cursor recorder always opened agents with full access, while five
fixtures replay with a read-only or workspace-write policy. Replay matches
the agent.open frame exactly, so any re-recording of those fixtures could
not replay. The recorder now takes each scenario's runtime policy override.

It also rewrites workspace paths in object keys (grep results are keyed by
path), maps the recording workspace's parent to /tmp, and waits one timer
tick before cancelling a mid-tool run: cancelling inside the SDK's
tool-call-started callback leaves an unhandled AbortError in @cursor/sdk
that kills the recorder.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…0.31

The Cursor transcripts dated from June (SDK 1.0.19) and predated the
thinking-delta stream, run usage, and the current tool shapes. Nine of the
ten are re-recorded live with composer-2.5; proposed_plan is left as is
because a plan-mode agent in an empty workspace explores the recording
host's filesystem and the transcript would publish those paths.

Three output assertions pinned the June wording and item order. They now
match today's recording: an extra reasoning segment after subagents, the
reworded todo_list progress lines, and a trimmed progress line in
tool_call_read_only (Grok's transcript shares that assertion).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…xture

The unit test for `$skill` rewriting fed a hand-built runner and asserted
on the captured message. The new skill_invocation fixture seeds
.cursor/skills/review/SKILL.md into the workspace and sends `$review`.
Replay matches the run.start frame exactly, so the rewrite to `/review` is
proven through the full orchestrator. (The adapter sets no settingSources,
so the SDK does not load the skill natively; in the recording the model
finds SKILL.md with Glob/Read and answers from it.) Fixture inputs can now
declare workspaceFiles, which the recorder and the replay workspace both
commit.

The task lifecycle test keeps its hand-built frames because a live run
cannot end without a task completion on demand, but they now follow the
recorded shape: a partial-tool-call before tool-call-started, and
subagentType unspecified with agentId and mode.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cursor skill discovery scans user roots under HOME as well as the
workspace, and the replay harness inherited the real process environment.
On a machine with a user-level `review` skill, skill_invocation passed even
without its seeded workspace skill. The Cursor replay registry now gets a
fixed HostProcessEnvironment whose HOME is an empty temp directory.

Also narrows the recorder's mid-tool cancel comment to what was verified.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…loaded

The first recording predated #13499, so the SDK loaded no project skills
and the model found SKILL.md by searching the workspace. Recorded again
with settingSources in agent.open (and an empty HOME), the SDK loads the
workspace skill natively and the model reads it straight from the `/review`
invocation before answering.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@juliusmarminge
juliusmarminge merged commit 393d7b3 into t3code/codex-turn-mapping Sep 24, 2026
23 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the v2/cursor-rerecord branch September 24, 2026 22:00
juliusmarminge added a commit that referenced this pull request Sep 24, 2026
…13493)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
juliusmarminge added a commit that referenced this pull request Sep 25, 2026
…13493)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant