Skip to content

fix(server): sandboxed Cursor threads keep working after a Full access thread - #13805

Merged
juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
fix/v2-cursor-sandbox-verdict
Sep 26, 2026
Merged

juliusmarminge merged 1 commit into
t3code/codex-turn-mappingfrom
fix/v2-cursor-sandbox-verdict

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

After one Full access Cursor thread runs, every later Supervised, Auto-accept edits, or Auto Cursor turn in that server fails with "Provider error: The provider could not start this turn" until the server restarts.

Evidence

Seen live on the V2 head with gpt-5.4-nano. A Full access thread ran a turn. Then that thread was switched to Supervised, and separately a new Supervised thread was created. Both failed the next turn instantly. The server trace shows the cause:

CursorAgentSdkRunnerError: Cursor Agent SDK run.start failed.
  [cause]: ConfigurationError: Local SDK sandboxing was requested, but sandboxing is not supported in this environment.

Sandboxing does work on that machine: the same server ran sandboxed turns fine when the first Cursor run after startup was sandboxed, and cursorsandbox --preflight-only exits 0.

Cause (upstream, @cursor/sdk 1.0.31, same code in 1.0.32)

  • The SDK decides once per process whether local sandboxing works (isSandboxSupported), then caches the answer.
  • Only the sandboxed setup path points it at the cursorsandbox helper (configureSandboxPrereqs) before asking.
  • An unsandboxed run still triggers the check when its executor is built. With no helper path set, the SDK caches "unsupported".
  • From then on, every sandboxed run in the process is rejected.

Standalone repro against the SDK, with no T3 code involved. Each block is a fresh process with the same repo and model:

unsandboxed send (no tools), then sandboxed create+send  -> sandboxed: ConfigurationError
sandboxed send, then unsandboxed, then sandboxed         -> all finished
unsandboxed Agent.create only (no send), then sandboxed  -> sandboxed: finished

Fix

Before the first unsandboxed Cursor agent opens, CursorAgentSdkRunner.open warms a bare sandboxed executor once per process. It uses the SDK's public createAgentPlatform().prewarmLocalWorkspace() with no setting sources and no MCP servers, then releases it right away. That configures the helper path, so the SDK caches the real answer. Sandboxed opens skip this because they configure the helper themselves.

  • Warming is best effort. Where sandboxing really is unsupported, it fails quietly, the SDK caches "unsupported", and sandboxed turns report that as they do today.
  • The cost is about one executor build (~1–2s) before the first Full access Cursor turn after the server starts. Later opens reuse the settled promise.
  • Replays are unaffected: the replay runner does not go through makeCursorAgentSdkRunner.

Verification

  • New env-gated live test in CursorOrchestratorV2.live.test.ts: "runs a sandboxed thread after a full access thread in the same server". It runs a Full access thread, then a Supervised thread, through the real orchestrator and SDK in one process (T3_CURSOR_LIVE_ORCHESTRATOR=1, CURSOR_API_KEY set).
    • Before (base runner): expected [ 'failed' ] to deeply equal [ 'completed' ] for the Supervised thread. Ran twice, failed both times.
    • After: passes.
  • vp test run src/orchestration-v2/Adapters/CursorAgentSdk.test.ts: 3 passed.
  • vp test run src/orchestration-v2/testkit/OrchestratorReplayFixtures.integration.test.ts -t cursor: 10 passed.
  • vp exec tsc --noEmit -p . in apps/server: no error TS or warning TS.
  • vp lint and vp fmt on the three touched files: clean. vp run knip:check: clean.
  • Not run: the full server suite, a packaged desktop or CLI build, macOS (its sandbox check is a separate sandbox-exec path), and the other live Cursor tests.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

…s thread

The Cursor SDK caches whether local sandboxing works the first time any run
starts, and only sandboxed runs point it at its cursorsandbox helper first.
After one Full access run it caches "unsupported", and every later
Supervised, Auto-accept edits, or Auto Cursor turn fails to start until the
server restarts. Warm a bare sandboxed executor before the first unsandboxed
agent opens so the SDK caches the real answer.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added size:M 30-99 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. labels Sep 26, 2026
@github-actions

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.7 KiB — 29.3 KiB ✅
Claude Live turn messages — 1 — 8 ✅

Baseline: unavailable · PR result: 9a2b13d · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@macroscopeapp

macroscopeapp Bot commented Sep 26, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Approved at 9a2b13d

Macroscope's review found this PR approvable — This is a focused Cursor SDK bug fix that adds one best-effort, one-time sandbox capability warm-up before Full access agents and verifies the affected sequence with an opt-in live test. Existing policy values and unsupported-environment behavior remain unchanged, with only bounded first-use setup latency added.

You can add or adjust custom eligibility rules. Learn more.

@juliusmarminge
juliusmarminge merged commit 4fafbbe into t3code/codex-turn-mapping Sep 26, 2026
24 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the fix/v2-cursor-sandbox-verdict branch September 26, 2026 18:32
juliusmarminge added a commit that referenced this pull request Sep 26, 2026
Brings in the V2 bug-hunt fixes merged while this PR was refreshed
(#13787, #13789, #13790, #13797, #13805, #13806). No conflicts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M 30-99 changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant