Conversation
|
Macroscope skipped reviewing this pull request. Per-review cost limit exceeded (workspace setting). This review would cost an estimated $17.93, which exceeds your per-review limit of $15.00. The top 3 files driving up this estimate:
Tip To get this pull request reviewed, you can:
|
ApprovabilityVerdict: Not approved Macroscope's review found this PR not approvable — This PR adds a large GPT Live capability spanning server brokering, secrets, billed OpenAI sessions, WebRTC, authenticated orchestration actions, UI automation, and persistent client history. Its broad runtime and security-sensitive surface, plus product-default changes and a diagnostic suppression, require human review. Not approved because:
Review your spending limits in Billing settings, or comment |
ea97483 to
9abcf4d
Compare
|
Additional UI consistency finding in Suggested fix: No single-hunk diff; this requires adding the Posted via Macroscope — UI Consistency |
|
Additional UI consistency finding in Suggested fix: No single-hunk diff; this requires adding the Posted via Macroscope — UI Consistency |
|
Effect service convention follow-up: Suggested fix: No diff — both declarations and consumers need coordinated renaming. Posted via Macroscope — Effect Service Conventions |
|
Effect Service Conventions found 3 blocking issues in Posted via Macroscope — Effect Service Conventions |
|
Correction to the summary above: the detailed review contains two inline findings, while the constructor/layer naming finding is in the immediately preceding PR-level comment. Posted via Macroscope — Effect Service Conventions |
6ca6a24 to
3e4ca4c
Compare
9abcf4d to
968deab
Compare
This comment has been minimized.
This comment has been minimized.
1bd44f2 to
3b9c885
Compare
968deab to
05d8dde
Compare
This comment has been minimized.
This comment has been minimized.
|
Effect Service Conventions found one new blocking issue in Posted via Macroscope — Effect Service Conventions |
|
Correction: the prior explicit Posted via Macroscope — Effect Service Conventions |
fe4f6ad to
87c67bd
Compare
Talk to T3 Code: a GPT Live (OpenAI Realtime over WebRTC) voice session that can list, open, read, start and follow up threads through the V2 orchestration APIs, with an overlay panel and a Settings > Integrations voice section. - Server: an authenticated voice broker under /api/voice mints Live sessions with the environment's stored OpenAI key (never sent to clients), plus session close, usage and the delegation backend. Every route requires orchestration:operate; settings also require access:write. The environment descriptor advertises voiceLive only while a key is configured. - Contracts: packages/contracts/src/voice.ts for the broker and tool schemas. - Web: the Live client, V2-backed tool executor (message.dispatch with commandId idempotency, run-status readouts), navigation, history, and the overlay UI. Mobile is not supported yet. - Docs: docs/user/voice*.md and docs/internals/voice-live.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
05d8dde to
850a8a7
Compare
|
All clear Posted via Macroscope — UI Consistency |
|
The latest change preserves the underlying settings error as Posted via Macroscope — Effect Service Conventions |
What Changed
Adds GPT Live voice to Orchestrator V2. You talk to T3 Code, and the voice agent lists, opens, reads, starts and follows up threads through the V2 orchestration APIs. The agent runs on OpenAI Live (Realtime over WebRTC).
/api/voicemints Live sessions with the environment's stored OpenAI key, and also handles session close, usage, and the Responses delegation backend. Every route requiresorchestration:operate, and settings also requireaccess:write. The key never leaves the server. The environment descriptor advertisesvoiceLiveonly while a key is configured, so clients show voice only where it works.packages/contracts/src/voice.tsholds the broker and tool schemas.message.dispatchwith a deterministiccommandId, run-status readouts from thread projections). Also navigation, a voice history, an overlay panel (corner button, expand, minimize), and a Settings → Integrations → Voice section for the key and models.docs/user/voice*.mdanddocs/internals/voice-live.md.This is the V2 counterpart of #12034, which targets
main.Why
Voice lets you steer threads hands-free: check on a worker, open a thread, or start a follow-up without leaving what you're doing. Porting it to V2 meant replacing the V1 session-status and turn model with V2 runs. Examples: delivery is keyed by run id, a follow-up readback detects a stale run by run identity, and dispatch mode comes from the active run's provider capabilities.
Surfaces: web and desktop. Mobile is not supported yet; the overlay only mounts in the web client. The voice agent drives threads through orchestration commands, so it works with every provider.
Relation to #13168
#13168 adds per-thread voice conversations on Codex's built-in realtime voice. The two are complementary: #13168 lets you talk to one thread's Codex session, and this PR adds an app-level agent that works across threads, projects and environments. Codex realtime can't host this agent's tools, because app-server realtime has no custom tool registry, so the agent stays on OpenAI Live.
Once #13168 lands, I plan to rebuild this PR on top of its transport. The agent would become a second session kind on its streaming voice-session RPC and shared client controller. That gives one voice stack, one entry point and one Settings section, and drops this PR's separate HTTP session endpoints and WebRTC handling. Until then, this PR stands on its own.
UI Changes
Before (V2 base without this PR): no voice entry point, and no Voice settings.
After (this PR): a Voice button appears in the corner and opens the panel, which minimizes back to the button. Settings → Integrations gains a Voice section with activation, fast commands, the key status, and the models.
Still captures
Before: landing view without a voice button.
After: landing view with the Voice button in the bottom-right corner.
After: the expanded panel, idle.
Before: Settings → Integrations.
After: Settings → Integrations with the Voice section.
Capture conditions: the GIFs and stills were recorded on the previous V2 head
d6dcd10, before the branch was rebased. At that verification point, the voice diff was byte-identical on the rebased branch, and I rechecked it there: fresh panel and Settings stills onf876a47match the ones below. Headless Chrome at 1280×800, both revisions against their own isolated dev home seeded with two synthetic projects. The candidate home holds an OpenAI key, which is what makesvoiceLivetrue. GIFs are real-time recordings at 10 fps. Each includes a full dev-mode page load when navigating to Settings, and a provider update toast dismissed with its own close button. The landing project differs between the two runs because of prior navigation state, which is unrelated to this change.Verification
ea97483on basef876a47): I used headless Chrome with a fake microphone. Connect went live in 2.9 s, End reached "Ended" 0.8 s later, and Connect was re-enabled, with no page errors.POST /api/voice/sessionsandPOST /api/voice/sessions/closeboth returned 200 against the real OpenAI Live API. The same check on the previous base took 5.9 s to go live. The historical head05d8ddediffered fromea97483only by the knip, import-order and test-dependency cleanup below, with no runtime behavior change.05d8ddeon base67a2be0:knip:check, covering files, dependencies and all workspace exports, is clean. The broker's secret name moved into its own leaf module (voice/secretNames.ts), so the environment descriptor no longer imports the broker; that import created a module cycle that broke the server memory tests. The PR adds no dependencies: the only jsdom-based voice test was removed, since nothing else in the repo uses a DOM test environment. The DOM control adapter keeps its logic tests indomControls.test.ts.vp test run src/voice src/missingThreadRedirects.test.ts src/components/settings→ 865 passed. The full web suite (437 files, 5,858 tests) also passed locally on the previous head.vp test run src/voice src/environment/ServerEnvironment.test.ts src/orchestration-v2/ProviderTurnStartService.memory.test.ts src/rpcInitialItems.memory.test.ts→ 41 passed.vp test run→ 524 passed.vp fmt --checkis clean.session.closedwith local WebRTC closed before broker accounting, stale-readback detection by run identity, restored DOM UI controls, and ending the session when the selected broker changes.Not exercised: a spoken conversation that drives tools (the fake mic only sends a tone), the Electron shell (desktop wraps the same web UI), and remote or tunnel connections, where the browser talks to OpenAI directly and the broker only exchanges SDP.
Known follow-ups: the research bridge's
delegate()has no production caller yet, so "start research and report back" launches the thread but records no delivery. Its polling stays dormant until that is wired.Checklist
apps/web/src/voicewith its tests, plus docs. It adds no dependencies.Model: Claude Opus 5.5 in T3 Code via Claude Code (V2 port fixes, review fixes, verification, PR); GPT-6 Astra via Codex (review); initial V2 port by GLM 5.3 Flash via OpenCode.
🤖 Generated with Claude Code
Latest review-fix verification (6fa193b):
vp test run apps/server/src/voice/broker.test.ts apps/web/src/voice/ui/VoiceControls.test.tsx apps/web/src/voice/ui/VoiceHistory.test.tsxpassed 44 tests; both UI files were rerun after the final control adjustment (18 passed).pnpm run typecheckpassed in apps/server and apps/web. Earlier screenshots and videos predate these control changes; no new browser verification was performed. Earlier head-specific test statements describe their historical runs.Final settings-error follow-up (
fa0ad2da83):vp test run apps/server/src/voice/broker.test.tspassed 27 tests;pnpm run typecheckin apps/server passed. Direct SWE-2 Max review exited 0 with no actionable findings.