Skip to content

test(server): record Codex replay fixtures on gpt-6-luna and Codex 0.156.1 - #13533

Merged
juliusmarminge merged 3 commits into
t3code/codex-turn-mappingfrom
v2/codex-gpt6-luna-rerecord
Sep 25, 2026
Merged

juliusmarminge merged 3 commits into
t3code/codex-turn-mappingfrom
v2/codex-gpt6-luna-rerecord

Conversation

@juliusmarminge

@juliusmarminge juliusmarminge commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

The hub no longer serves gpt-5.4, the model every Codex replay fixture was pinned to, and most of those transcripts came from Codex 0.120–0.137. This re-records every Codex fixture the recorder can reproduce, live on Codex 0.156.1, and moves the default Codex test model to gpt-6-luna.

What changed

Default model. CODEX_MODEL_SELECTION and the recorder's default are now gpt-6-luna. The per-fixture gpt-5.6-sol (subagent_v2, subagent_v2_nested) and gpt-5.6-luna (thread_rollback) overrides from #13505/#13508 are gone, because all three now record on the default.

Recorder (record-codex-app-server-replay-fixture.ts):

  • thread/start, thread/resume and thread/fork now send cwd and model, the way the adapter does. The old {} started threads on the hub's default model, and a revert then reloaded them with a model-mismatch warning. Replay ignores these fields, so the transcripts still match. Resume also sends excludeTurns: true, and fork sends the source's latest lastTurnId, again matching the adapter.
  • Scenarios can set their own model; --model still overrides it. Plan mode now takes the collaboration-mode model from the run instead of a hardcoded gpt-5.4.
  • Two new scenarios, queued_turn and web_search. Both fixtures were 0.120 probe captures the recorder couldn't reproduce.
  • After turn/interrupt, the recorder terminates any command still running (thread/backgroundTerminals/terminate), as the adapter does.
  • Codex no longer inherits VP_TOOL_RECURSION from the vp node shim. With it set, every node the agent ran failed with "Recursion detected". That turned the mid-tool interrupt recording into a command that failed instantly instead of one that was interrupted.
  • codexReplayRecordingRecords.ts scrubs every frame before writing: the checkout root becomes /home/replay-user/t3code, the home directory /home/replay-user, the hostname replay-host, any installationId value 00000000-0000-4000-8000-000000000000, and any planType value unknown. The adapter reads none of these.
  • The same substitutions were applied to every committed Codex transcript, including the three not re-recorded. That removes this machine's Codex installation id (it was in 23 transcripts, 3 of them already on the base from test(server): record the Codex thread rollback fixture live on 0.156.1 #13505/test(server): record Codex multi-agent V2 subagents live on 0.156.1 #13508), the pro plan type, and old /Users/… and worktree paths. Outbound frames were checked to be byte-identical before and after, so replay matching is unaffected.

Per-fixture models. Each one is set in both the recorder scenario and the fixture's modelSelection:

Fixture Model Why
subagent, subagent_continue gpt-5.6-luna Keeps the multi-agent v1 (collabAgentToolCall) path covered. gpt-6-luna runs v2, and these are the only v1 recordings.
tool_call_read_only_on_request gpt-6-sol Under a read-only sandbox, gpt-6-luna refuses the write without trying it, so no approval is ever requested.
tool_call_restricted_granular gpt-5.6-terra The gpt-6 models write through the shell, which raises a command approval. gpt-5.6-terra uses apply_patch, which raises the file-change approval this fixture exists for.
thread_fork_native_siblings, thread_merge_back_continue, thread_merge_back_siblings gpt-6-sol On the recall turn, gpt-6-luna often repeats an earlier acknowledgement ("merge source stored") instead of the markers. ThreadMergeBack.integration.test.ts uses gpt-6-sol for its Codex variant to match.

Multi-agent versions, checked in the recorded frames:

  • subagent: v1. collabAgentToolCall spawnAgent ×2 and wait, no subAgentActivity.
  • subagent_continue: v1. spawnAgent, then resumeAgent and sendInput into the same child thread.
  • subagent_v2: v2. subAgentActivity started/completed. The child's turn/started comes after the parent's activity, so assertSubagentActivityPrecedesChildTurns still expects 1 child turn.
  • subagent_v2_nested: v2 at three levels, 3 child turns.

Recorder config for two scenarios:

  • todo_list sets tools.update_plan.enabled: true. Codex 0.156 only registers update_plan when this is on (resolve_update_plan_enabled defaults to false). Without it the model reports "Plan update unavailable" and emits no turn/plan/updated. The adapter doesn't set it either, so for real users the todo list is empty on 0.156. That's a separate decision and is not changed here.
  • subagent sets skills.include_instructions: false. Without it, the children read the recording user's skill files into the transcript.

Assertion changes

Where an assertion pinned wording the model picked, I loosened it to the behaviour. None of these hides an adapter change.

  • subagent checks that each child got one file and reported something only that file contains (effect-codex-app-server, ../../tsconfig.base.json). It used to match gpt-5.4's exact prompt and result phrasing.
  • subagent_v2 / _nested check that each agent path extends its parent's by one segment. They used to pin the names the old model picked (/root/hello, relay_one/relay_two). The nested answers are Hello. at every level now.
  • subagent_continue uses a follow-up prompt without @hooke, a nickname the old recording happened to get, and its CodexReplayFixtures expectations use the recorder's method labels instead of hand-written role labels.
  • web_search now expects user, assistant, assistant. gpt-6 posts a commentary message ("I'll check current ticket pricing…") before the answer, and the adapter projects commentary as an assistant message by design.
  • proposed_plan checks that the plan mentions replay and fixtures instead of the heading Deterministic Replay Fixtures. The expected incoming label is now item/plan/delta: the model streams the plan and gives no separate agent message.
  • plan_questions: the model named its question schema_preference. The Codex variant asserts that id, and the input answers it. Grok and ACP keep their recorded schema_vs_ui_flexibility through the shared helper.
  • thread_rollback no longer expects two reasoning items in the post-rollback turn, since this recording has none. The rollback checks are unchanged: turn 2 hidden, the recall has turn 1 and not turn 2.
  • tool_call_restricted_granular no longer requires a command_execution item; the file-change approval is the behaviour under test. The turn/diff/updated and item/fileChange/outputDelta expectations are gone because the recorder opts out of turn/diff/updated and 0.156 sends no output delta for apply_patch.
  • turn_interrupt_mid_tool checks that the terminate request names the thread and processId of the command that was running. It used to hardcode the old recording's ids. It now also accepts the terminate going out before turn/completed: the adapter runs termination alongside the completion wait, and in this recording it lands first.
  • provider_thread_resume: since the adapter resumes with excludeTurns, the resume response has no turns. The check now asserts that the request sets excludeTurns, that the response is empty, and that the second answer still repeats the first. That last check proves the history reached the model.

Not re-recorded

  • thread_fork_native: hand-authored, with synthetic thread ids. ThreadFork.integration.test.ts depends on those ids.
  • thread_fork_native_prior_turn: a normalized 0.124 live probe that forks at an earlier turn. The recorder has no step that forks at a chosen turn.
  • delegated_task_status: constructed from composed frames for OrchestratorMcpToolkit.integration.test.ts. The recorder can't drive the MCP delegation it needs.
  • queued_cancelled_while_active: reuses the queued_turn transcript, which is re-recorded.

ThreadFork.integration.test.ts and OrchestratorMcpToolkit.integration.test.ts still pin gpt-5.4, because their three transcripts do.

Reviewer note: the adapter's excludeTurns on resume is only covered by OrchestratorReplayRecovery. If it regresses, that test fails with a 60s await timeout instead of a frame mismatch.

Bugs found

No adapter bug. Two findings for follow-up, not changed here:

  • update_plan is off by default in Codex 0.156, so T3 users on 0.156 get no todo lists unless the adapter enables it (see above).
  • The vp node shim sets VP_TOOL_RECURSION, and any child that inherits it breaks node inside agent shells. The recorder now strips it. A T3 server started through vp likely passes it on to Codex the same way. I did not check that.

Line delta: +2637 / −2841 across 41 files. Code and tests account for +281 / −127 in 16 files; the other 25 files are transcripts.

Verification

  • Recorded all 23 fixtures live with T3_CODEX_BIN=<codex 0.156.1> node apps/server/scripts/record-codex-app-server-replay-fixture.ts --scenario <name>, run from packages/effect-codex-app-server. None needed a rate-limit retry.
  • vp test run in apps/server on OrchestratorReplayFixtures.integration (all providers), OrchestratorReplayRecovery.integration, CodexReplayFixtures.integration, ThreadFork.integration, ThreadMergeBack.integration, OrchestratorReplayFixtures.contract, CodexAdapterV2.test, OrchestratorMcpToolkit.integration, ProviderSwitch.integration, ClaudeReplayFixtures.integration: 10 files, 287 tests passed.
  • After the scrub commit, the same suite minus ProviderSwitch and ClaudeReplayFixtures (which don't read Codex transcripts): 8 files, 220 tests passed. A scratch live recording of simple after the scrub carried the placeholder installation id, planType: "unknown", replay-host and /home/replay-user/t3code paths, and no local strings.
  • vp exec tsc --noEmit -p . in apps/server: no error TS or warning TS.
  • vp run knip:check: passes.
  • vp lint on the touched files: only the unused assertUserMessagesExclude in fixtures/shared.ts, which is already on the base branch.
  • Grepped all Codex transcripts (not only the diff) for home and /Users/ paths, .t3/worktrees, the installation id, "planType":"pro", hub and tailnet hostnames and IPs, the machine name, emails, JWTs and bearer strings. The only hit is git@github.com in a recorded git remote. Codex also reports the checked-out branch, the commit sha and https://github.com/pingdotgg/t3code.git in gitInfo; all of that is public.

Model: Claude Opus 5.5 (Claude Code)

🤖 Generated with Claude Code


Devin Review

juliusmarminge and others added 2 commits September 24, 2026 18:09
The recorder started threads with `{}`, so Codex used the hub's default
model and a later revert reloaded the thread with a model-mismatch
warning. It now sends `cwd` and `model` on thread/start, resume and fork
the way the adapter does, resumes with `excludeTurns`, forks at the
source's latest turn, and terminates commands an interrupt left running.
Replay ignores the thread params, so recordings still match.

The default model is gpt-6-luna. Scenarios can pin their own model where
gpt-6-luna does not exercise the path (multi-agent v1, apply_patch
approvals, marker recall). New queued_turn and web_search scenarios
replace 0.120 probe captures. todo_list enables `tools.update_plan`,
which Codex 0.156 leaves off by default.

Recordings now replace the home directory and hostname with neutral
values, and Codex no longer inherits the vp shim's VP_TOOL_RECURSION,
which made every `node` the agent ran fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…156.1

The hub no longer serves gpt-5.4, and most Codex transcripts came from
Codex 0.120-0.137. Re-record the 23 fixtures the recorder supports live
on Codex 0.156.1. The shared Codex model is now gpt-6-luna, and the
per-fixture overrides from the rollback and subagent_v2 recordings are
gone.

subagent and subagent_continue stay on gpt-5.6-luna so the multi-agent
v1 collabAgentToolCall path keeps live coverage. A few fixtures pin
gpt-6-sol or gpt-5.6-terra where gpt-6-luna skips the behaviour under
test.

Assertions that pinned the old model's wording (agent paths, plan
heading, question id, subagent phrasing, reasoning item count) now check
the behaviour instead. The resume check follows the adapter's
excludeTurns request, and the interrupt check derives the terminated
process from the recording.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting).

This review would cost an estimated $99.19, which exceeds your per-review limit of $15.00.

The top 3 files driving up this estimate:

File Diff Size Estimate
apps/server/src/orchestration-v2/testkit/fixtures/subagent/codex_transcript.ndjson 408.04KB $20.40
apps/server/src/orchestration-v2/testkit/fixtures/proposed_plan/codex_transcript.ndjson 319.53KB $15.98
apps/server/src/orchestration-v2/testkit/fixtures/thread_merge_back_siblings/codex_transcript.ndjson 127.61KB $6.38

Tip

To get this pull request reviewed, you can:

  1. Comment @macroscope-app on this PR to request a manual review (monthly spend limits still apply).
  2. Exclude the file(s) above from review by adding a pattern to your .macroscope/ignore.md — note that creating this file replaces Macroscope's built-in default ignores rather than extending them.
  3. Raise your cost limit in your workspace billing settings.

Turn off this reminder going forward

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Sep 25, 2026
@macroscopeapp

macroscopeapp Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Would Approve

Macroscope's review found this PR approvable — This PR refreshes Codex replay fixtures and improves the test-only recording harness for deterministic machine identity, model selection, thread lifecycle, and new scenarios. It does not alter production request paths or product defaults.

Not approved because:

  • Per-review cost limit exceeded (workspace setting). Approvability relies on correctness review in order to determine eligibility

Review your spending limits in Billing settings, or comment @macroscope-app review this PR to bypass the limit and review now. You can add or adjust custom eligibility rules. Learn more.

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 4.9 KiB — 6.8 KiB ✅
Codex Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Codex Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Codex Live turn WebSocket decoded — 20.4 KiB — 29.3 KiB ✅
Codex Live turn messages — 2 — 8 ✅
Claude Total thread wire — 4.9 KiB — 6.8 KiB ✅
Claude Thread snapshot wire — 3.7 KiB — 4.9 KiB ✅
Claude Live turn WebSocket wire — 1.2 KiB — 2.0 KiB ✅
Claude Live turn WebSocket decoded — 20.8 KiB — 29.3 KiB ✅
Claude Live turn messages — 2 — 8 ✅

Baseline: unavailable · PR result: 4136fc4 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

…play fixtures

The recorded remoteControl/status/changed frames carried this machine's
Codex installation id, in 23 transcripts. Account/updated frames carried
the plan type, and inbound cwd values named the local worktree layout.

The recorder now replaces any installationId with a fixed placeholder
UUID, any planType with "unknown", and the checkout root with
/home/replay-user/t3code, alongside the home and hostname it already
scrubbed. The same substitutions are applied to every committed Codex
transcript, including the hand-authored and constructed ones, which also
drops their old /Users paths. Outbound frames are byte-identical, so
replay matching is unchanged; the adapter reads none of these values.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@juliusmarminge
juliusmarminge merged commit 9604f00 into t3code/codex-turn-mapping Sep 25, 2026
24 of 25 checks passed
@juliusmarminge
juliusmarminge deleted the v2/codex-gpt6-luna-rerecord branch September 25, 2026 02:16
juliusmarminge added a commit that referenced this pull request Sep 25, 2026
…156.1 (#13533)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant