test(server): re-record Cursor replay fixtures and cover skills live - #13493
Conversation
3375075 to
c1bf00f
Compare
|
Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting). This review would cost an estimated $19.76, which exceeds your per-review limit of $15.00. The top 3 files driving up this estimate:
Tip To get this pull request reviewed, you can:
|
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — This is primarily a test-only Cursor fixture refresh with no product-path or default-behavior change. Human review is warranted for the unresolved temporary HOME directory lifecycle issue in the replay layer, which may prevent scoped cleanup. Not approved because:
Review your spending limits in Billing settings, or comment |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
| ).pipe(Layer.provide(NodeServices.layer)); | ||
| // Skill discovery also scans user roots under HOME; an empty HOME keeps | ||
| // replays from picking up the host's own skills. | ||
| const hostEnvironmentLayer = Layer.effect( |
There was a problem hiding this comment.
This layer acquires a scoped temporary directory, so it should use the scoped layer constructor to keep the directory alive for the layer's lifetime and run its finalizer when the layer scope closes.
| const hostEnvironmentLayer = Layer.effect( | |
| const hostEnvironmentLayer = Layer.scoped( |
Posted via Macroscope — Effect Service Conventions
9fc8676 to
ae065ff
Compare
5841eaa to
a171d01
Compare
The Cursor recorder always opened agents with full access, while five fixtures replay with a read-only or workspace-write policy. Replay matches the agent.open frame exactly, so any re-recording of those fixtures could not replay. The recorder now takes each scenario's runtime policy override. It also rewrites workspace paths in object keys (grep results are keyed by path), maps the recording workspace's parent to /tmp, and waits one timer tick before cancelling a mid-tool run: cancelling inside the SDK's tool-call-started callback leaves an unhandled AbortError in @cursor/sdk that kills the recorder. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…0.31 The Cursor transcripts dated from June (SDK 1.0.19) and predated the thinking-delta stream, run usage, and the current tool shapes. Nine of the ten are re-recorded live with composer-2.5; proposed_plan is left as is because a plan-mode agent in an empty workspace explores the recording host's filesystem and the transcript would publish those paths. Three output assertions pinned the June wording and item order. They now match today's recording: an extra reasoning segment after subagents, the reworded todo_list progress lines, and a trimmed progress line in tool_call_read_only (Grok's transcript shares that assertion). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…xture The unit test for `$skill` rewriting fed a hand-built runner and asserted on the captured message. The new skill_invocation fixture seeds .cursor/skills/review/SKILL.md into the workspace and sends `$review`. Replay matches the run.start frame exactly, so the rewrite to `/review` is proven through the full orchestrator. (The adapter sets no settingSources, so the SDK does not load the skill natively; in the recording the model finds SKILL.md with Glob/Read and answers from it.) Fixture inputs can now declare workspaceFiles, which the recorder and the replay workspace both commit. The task lifecycle test keeps its hand-built frames because a live run cannot end without a task completion on demand, but they now follow the recorded shape: a partial-tool-call before tool-call-started, and subagentType unspecified with agentId and mode. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cursor skill discovery scans user roots under HOME as well as the workspace, and the replay harness inherited the real process environment. On a machine with a user-level `review` skill, skill_invocation passed even without its seeded workspace skill. The Cursor replay registry now gets a fixed HostProcessEnvironment whose HOME is an empty temp directory. Also narrows the recorder's mid-tool cancel comment to what was verified. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…loaded The first recording predated #13499, so the SDK loaded no project skills and the model found SKILL.md by searching the workspace. Recorded again with settingSources in agent.open (and an empty HOME), the SDK loads the workspace skill natively and the model reads it straight from the `/review` invocation before answering. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
a171d01 to
cd3ec2f
Compare
393d7b3
into
t3code/codex-turn-mapping
…13493) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…13493) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Builds on #13499 (merged), which makes Cursor V2 pass
local.settingSources. The transcripts here carry that field inagent.open.The Cursor replay transcripts were recorded in June against
@cursor/sdk1.0.19; the server now runs 1.0.31. They predate thethinking-deltastream, run usage, and today's tool-call shapes, so replays were checking behavior Cursor no longer shows. A skill unit test also used a hand-built runner where a live fixture can cover the same path.What changed
Recorder (first commit). The recorder opened every agent with full access, but five fixtures replay with a read-only or workspace-write policy. Replay matches the
agent.openframe exactly, so any re-recording of those fixtures could not replay. The recorder now takes each scenario's runtime-policy override. It also:/tmp;tool-call-startedcallback leaves an unhandledAbortErrorin@cursor/sdkthat kills the recorder process (see below).Re-recorded fixtures (second commit). Nine of the ten Cursor transcripts were re-recorded live with
composer-2.5:simple,multi_turn,message_steering,provider_thread_resume,queued_turn,todo_list,subagent,tool_call_read_only, andturn_interrupt_mid_tool.proposed_planwas left as is. In an empty workspace, the plan-mode agent searches outside it (/tmp, the recording host's cache and worktree) and quotes those files in its plan, so the transcript would publish host paths. Three output assertions pinned the June wording and item order and now match today's run:subagent: a reasoning segment now follows the subagents.todo_list: the progress lines are reworded and there are more reasoning segments.tool_call_read_only: the progress line is compared after trimming, because trailing newlines vary between runs. Grok's transcript shares this assertion.None of these differences was an adapter bug.
Skill fixture (third commit). The new
skill_invocationfixture seeds.cursor/skills/review/SKILL.mdand sends$review README.md. Replay matchesrun.startexactly, so the adapter's rewrite to/reviewis now proven through the full orchestrator. It was recorded with #13499'ssettingSourcesin place, so the SDK loads the workspace skill natively: the model readsSKILL.mdstraight from the/reviewinvocation, with no workspace search, and answers from it. This replaces the unit test "sends discovered skills as native slash invocations". Fixture inputs gainedworkspaceFiles, which the recorder and the replay workspace both commit.The task-lifecycle unit test stays, because a live run can't be made to end without a task completion. Its frames now follow the recorded shape: a
partial-tool-callbeforetool-call-started, andsubagentType: {kind: "unspecified"}withagentIdandmode.Not done
The proposed
tool_call_ls_lintsfixture, which would have replaced thels/readLintsprojection unit test, can't be produced live with this SDK:readLintscan never return errors locally.lstool:composer-2.5,gpt-5.4, andclaude-sonnet-4-6all list directories throughShellorGlob. No Cursor transcript in the repo has ever recordedls.The unit test stays.
Verification
vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts src/orchestration-v2/testkit/OrchestratorReplayRecovery.integration.test.ts src/orchestration-v2/testkit/OrchestratorReplayFixtures.contract.test.ts src/orchestration-v2/Adapters/CursorAdapterV2.test.ts src/orchestration-v2/Adapters/CursorAdapterV2.testkit.test.ts scripts/cursorReplayRecordingWorkspace.test.ts: 102 passed, which includes all 75 replay fixtures across providers.skill_invocationfail with arun.startframe mismatch.HOME. Before, withHOMEpointing at a directory holding a user-levelreviewskill,skill_invocationstill passed when the workspace skill was moved to a non-skill path. It now fails with arun.startmismatch, and the seeded fixture still passes under thatHOME. The Cursor replay and recovery suites pass (12), and so do the adapter and contract tests (22).vp exec tsc --noEmit -p .inapps/server: no errors.vp linton touched files: one warning that already exists on the base branch.settingSourcesfield. The first nine gained only that field, inserted byte-identically to fix(server): Cursor V2 threads load project skills and rules #13499's edit.skill_invocationwas re-recorded live with the fix. The Cursor replay and recovery suites pass (12); the adapter, testkit, contract, and recorder-workspace tests pass (26). The hermeticity and negative checks were re-run and hold. Servertscshows no TS errors or warnings.vp run knip:checkpasses (the earlier unused-export failure was fixed on the base by chore(web): drop the right panel sheet class left unused by the V2 rebase #13502).Model: Claude Opus 5.5 (Claude Code)
🤖 Generated with Claude Code