test(server): remove skipped, duplicate, and constant-restating V2 tests - #13459
Conversation
Drops tests that never run, repeat coverage that a replay fixture or a stronger sibling test already provides, or only restate a constant or a fake's call log. Moves the one assertion unique to the removed OpenCode clean-EOF test (threadDisposition "broken") into the surviving one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This PR removes skipped and redundant server tests and unused test-harness code across test-only files; the sole assertion relocation preserves existing coverage. No production runtime behavior, product defaults, schema, or static-analysis configuration is changed. You can add or adjust custom eligibility rules. Learn more. |
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
4ed185e
into
t3code/codex-turn-mapping
…sts (#13459) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…sts (#13459) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…sts (#13459) Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Some V2 server tests never run, some repeat coverage that a recorded replay fixture or a stronger sibling test already provides, and some only check a constant or a fake's call log. They cost review and maintenance time and prove nothing extra. This PR deletes them after checking each claim against the source.
Removed (one line each)
testkit/ThreadFork.integration.test.tsit.effect.skip("merges a fork delta back into the source thread through context handoff"): never runs.ThreadMergeBack.integration.test.tscovers merge-back for both Codex and Claude using recorded transcripts: fork-delta summary text, thefork_delta_summaryhandoff, a consumed transfer withfork_delta_context, the source thread's native id preserved, handoff text kept out of the visible conversation, and sibling merges.makeExpectedForkDeltaSummary,transcriptWithMergeBackContinuation, and the helpers only they used (compactExpectedText,findCompletedAgentMessageText, theProviderReplayEntryimport) are removed too. The skipped test also asserted that a stale pending merge-back gets superseded. That assertion never ran, so no active coverage is lost (see Notes).testkit/ClaudeReplayFixtures.integration.test.tsentries.length >= 3.scripts/record-claude-agent-sdk-replay-fixture.ts --scenario simpleis the real recorder.emit_inboundframes out of the transcript (prompts and model were ignored).OrchestratorReplayFixturesrunssimple/claudeAgentandmulti_turn/claudeAgentthrough the full orchestrator and asserts the same assistant text plus the projection shape. The helperreplayClaudeAgentSdkTranscriptis now unused and is removed fromClaudeAdapterV2.testkit.ts.EffectWorker.test.ts["interrupt","detach","start"]on string-recording fakes, which restates the executor'sandThenchain.SelectionRestart.integration.test.ts"detaches the old provider session after an active provider handoff" drives the samedetachtransition through the real orchestrator, session manager, and effect worker, and asserts the old session closes and the target starts once.(kind, requestId). The case is a direct field pass-through that the service signature checks at compile time. The title service's real behaviour is covered inThreadTitleRegenerationService.test.tsandThreadLaunchService.test.ts.CheckpointRollbackService.test.tsmapErrorfallback's reason, message template, and cause. That mirrors a four-line wrapper line for line, and nothing branches on this reason.ThreadLaunchService.test.tsruntimeLayer.test.tsit.layer(TestLayer)block dispatchesthread.createand reads a projection through the same composition, then asserts more.Adapters/CursorAdapterV2.test.ts,Adapters/OpenCodeAdapterV2.test.ts,Adapters/ClaudeAdapterV2.test.tsas constcapability object and asserts the same literals.Adapters/OpenCodeAdapterV2.test.tstransport_error) and also checks that a laterstartTurnfails. The removed test also assertedthreadDisposition: "broken", so that assertion moves into the surviving test.Adapters/ClaudeAdapterV2.test.ts,Adapters/CodexAdapterV2.test.tsif (logger === undefined) return undefinedguard.Adapters/ClaudeAdapterV2.test.tssummary: null. The SDK typessummaryasstring, and every recordedtask_notificationin the fixtures has a non-empty summary. I deleted it instead of changing the frame: with a realistic summary it becomes "buffers wake output and requests a single continuation run" (same frames, assertsdetail === WAKE_SUMMARY) plus "terminalizes a continuation turn from a task-notification origin wake result". The adapter's defensivetypeof === "string"guard stays.provider/acp/XAiAcpExtension.test.tsKept after checking
EffectWorker.test.ts"uses durable deadlines, notifications, and a slow liveness poll": not a duplicate ofFoundationPersistence.test.ts"executes a retry at its durable deadline instead of the liveness interval". That test only covers a near deadline beating the liveness poll. This one also covers a far deadline capped by the liveness poll, a notification waking the loop early, and polling with no deadline.CheckpointRollbackService.test.ts"rejects a non-ready checkpoint before opening a session or restoring files": theruntimeLayer.test.tsrollback test covers the orchestrator's dispatch-time guard, which is a different check. This one covers the service's own execution-time guard, which matters because the service itself marks checkpointsstaleafter a rollback and an effect can run after the checkpoint changes.ThreadLaunchService.test.ts"deduplicates retried launch side effects in-process": not the same as "replays a server-allocated launch". It uses a client-supplied thread id and two concurrent calls, so it is the only test that covers the concurrent receipt-replay path for client-id launches.ClaudeReplayFixtures"classifies every fixture tool use" and "keeps unregistered native conversation-state transcripts reviewable": kept as asked.Verification
vp test runon EffectWorker, CheckpointRollbackService, ThreadLaunchService, runtimeLayer, Cursor/OpenCode/Claude/Codex adapter tests, and XAiAcpExtension: 9 files, 443 tests passed.vp test runon ThreadFork, ClaudeReplayFixtures, and ThreadMergeBack integration tests: 3 files, 12 tests passed.vp test run OrchestratorReplayFixtures.integration.test.ts -t "runs (simple|multi_turn)/claudeAgent": both passed. These are the fixtures cited above as coverage.vp exec tsc --noEmit -p .inapps/server: noerror TS.vp linton the 12 touched files: two warnings, both already on the base branch and not caused by this diff (an unusedmakeClaudeAgentSdkReplayQueryRunnerLayerin the Claude testkit, and an optional-chaining warning in an untouched ThreadLaunchService test).vp fmtis clean.Notes
Orchestrator.ts~3241 and ~4630/5786). The only test for it was the skipped one removed here. Worth a recorded fixture if that path matters.Line delta: +3 / −983 across 12 files.
Model: Claude Opus 5.5 (Claude Code)
🤖 Generated with Claude Code