Skip to content

feat(desktop): add free local prompt dictation - #136

Merged
NoahHendrickson merged 9 commits into
customfrom
t3code/research-native-speech-to-text
Sep 11, 2026
Merged

NoahHendrickson merged 9 commits into
customfrom
t3code/research-native-speech-to-text

Conversation

@NoahHendrickson

@NoahHendrickson NoahHendrickson commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

What Changed

Add built-in macOS desktop prompt dictation: click the microphone or press Ctrl+Shift+Space, speak, and press it again (or click stop) to insert an editable transcript at the cursor. Escape cancels. Sending is blocked and the editor is frozen during recording and transcription; changing threads cancels the recording. A recording in flight survives an approval or question arriving mid-sentence, and covering or minimizing the window does not cancel it.

Bundle a pinned whisper.cpp helper (Metal, precompiled shaders) and download the English quantized Whisper Small model (ggml-small.en-q5_1, 181 MiB) once with SHA-256 verification; later launches trust a sidecar stamp so the first dictation of a session does not re-hash the file. Transcription stays on the client device, including for remote environments, and works offline after the download. Temporary recordings are deleted after processing, and the helper exits to release model memory.

Why

Desktop users need free prompt dictation without an API key, subscription, or per-minute fees. This reuses the shared VoiceInputController and adds device-local IPC; server protocols and existing mobile/browser-only behavior are unchanged.

Shape

Mirrors mobile: ChatComposer owns the session through useForkDictationController (in apps/web/src/custom/voice/) and reads blocksSubmission / freezesEditor directly. ForkDictationControl only renders state. Fork IPC types live in apps/desktop/src/fork/voice/ and a renderer-local interface, not in packages/contracts or client-runtime.

Platform

macOS only. whisper.cpp publishes no prebuilt macOS CLI, so the helper is compiled from pinned source. The desktop dev/build tasks skip it when CMake is missing (the app runs without dictation); only the packaged artifact requires the toolchain. Browser-only (app.t3.codes) gets no mic; mobile keeps its native path.

Validation

  • Desktop, web, and client-runtime typecheck clean; knip clean.
  • Engine, recorder, fork guard, manifest, and packaging tests pass (the one build-desktop-artifact failure locally is the pre-existing host-arch-dependent Windows probe test, green on the Linux runner).
  • Live microphone/UI interaction has not been verified yet.

UI Changes

Adds a composer microphone, recording/download/transcription status, and cancellation. Screenshots and video deferred; live desktop verification pending.

Implemented with GPT-6 in the Codex harness; restructured with Claude Fable 5.1 in Cursor.

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 11, 2026
@github-actions

github-actions Bot commented Sep 11, 2026 •

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 13.6 KiB 13.6 KiB +29 B (+0.2%) 15.1 KiB ✅
Codex Thread snapshot wire 7.0 KiB 7.0 KiB −2 B (−0.0%) 7.3 KiB ✅
Codex Live turn WebSocket wire 6.5 KiB 6.5 KiB +31 B (+0.5%) 7.8 KiB ✅
Codex Live turn WebSocket decoded 57.0 KiB 57.1 KiB +88 B (+0.2%) 66.4 KiB ✅
Codex Live turn messages 8 10 +2 (+25.0%) 21 ✅
Claude Total thread wire 13.6 KiB 13.6 KiB −31 B (−0.2%) 15.1 KiB ✅
Claude Thread snapshot wire 7.0 KiB 7.0 KiB +2 B (+0.0%) 7.3 KiB ✅
Claude Live turn WebSocket wire 6.6 KiB 6.5 KiB −33 B (−0.5%) 7.8 KiB ✅
Claude Live turn WebSocket decoded 57.9 KiB 57.8 KiB −44 B (−0.1%) 66.4 KiB ✅
Claude Live turn messages 10 9 −1 (−10.0%) 21 ✅

Baseline: 3642323 · PR result: 36aab78 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 113.9 KiB
  • Claude decoded thread snapshot: 114.6 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale comment

This inverts the architecture that already exists. Mobile’s composer holds useVoiceInputController and derives blocksSubmission / freezesEditor locally. This PR puts the controller in a child, punches a boolean back through onBusyChange, then pins that flag into six last-resort ChatComposer fences via state and a ref. ChatComposer is already 6089 lines. That is a busy-flag pinboard, not an integration.

The judo is obvious: a useForkDictationController in apps/web/src/custom/voice/ that matches mobile. Composer owns the hook. The button is dumb. onBusyChange, voiceInputBusy, voiceInputBusyRef, remount-via-key, and data-fork-dictation-composer all go away. Fork IPC types also do not belong in packages/client-runtime. Do not land this shape.

Open in Web View Automation 

Sent by Cursor Automation: Thermo nuke 4.6

Comment thread apps/web/src/components/chat/ChatComposer.tsx Outdated
Comment thread apps/web/src/custom/voice/ForkDictationControl.tsx Outdated
Comment thread packages/client-runtime/src/voice-input/desktop.ts Outdated
Comment thread apps/desktop/src/fork/voice/VoiceInputIpc.ts Outdated

@NoahHendrickson NoahHendrickson left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review — local prompt dictation

Reviewed every hunk, traced the shared VoiceInputController contract into the Lexical editor's disabled path, and walked the electron-builder staging/entitlements flow. Each finding below was verified against the source rather than inferred from the diff.

One thing I checked and want to explicitly clear: Schema.Uint8Array over IPC is correct. In Effect 4 it is instanceOf<Uint8Array>, not a number-array transform, so the decode of a structured-cloned Uint8Array succeeds. The WAV encode/validate round-trip, the expanded-vs-collapsed cursor mapping in commitDraft, the model SHA-256/size verification, and the spawn-failure event ordering in runWhisper all check out too.

Findings 1, 2 and 3 are the ones I'd want fixed before merge — all three are small and all three are in ForkDictationControl.tsx.


On the direction

The core call is right. whisper.cpp running locally is the correct answer for a BYO-subscription app: free, no API key, offline, and the audio never touches the wire — which is what makes it work when the environment is remote but the microphone is local. Reusing mobile's existing VoiceInputController instead of writing a second state machine is exactly the right instinct. The shape of this should not change.

Four things worth pushing on before this is "just works":

Whisper small with auto-detect is the wrong default

466 MiB is a rough first-run tax, and multilingual small is worse at English than small.en at identical size. For English coding prompts — short, technical, dense with identifiers — base.en is 142 MiB and roughly 3x faster, and at dictation clip lengths the accuracy gap to small is small. Defaulting to base.en would cut the download by ~70%, cut the startup re-hash cost (finding 9), and cut transcription latency. Paired with finding 8 this is the highest-leverage quality change available here, and it is nearly a one-line change.

Windows and Linux have no acceleration path at all

The CMake flags set GGML_NATIVE=OFF, GGML_BLAS=OFF, no CUDA, no Vulkan; Metal is macOS-only. Add hardcoded --threads 4 and a 16-core Windows box runs four baseline-ISA CPU threads with no GPU. The PR description is candid that Windows/Linux were never executed. On an M3 Max, 11s of audio in ~1s is great; the same clip on a mid-range Windows laptop under these flags could plausibly run slower than realtime, at which point the feature is worse than typing. At minimum, scale --threads to core count, and treat Windows/Linux performance as unverified-until-measured rather than assumed.

Build-from-source is the wrong dependency to take on

It taxes every desktop contributor (finding 6), makes CI compile whisper.cpp per platform, and requires the Metal toolchain on macOS. whisper.cpp publishes prebuilt release binaries — fetching a pinned asset with SHA-256 verification is the same pattern this PR already implements correctly for the model, and would be smaller, faster and more reproducible. If compiling must stay, at least make it lazy so it does not block dev.

No streaming means perceived latency is the full stop-then-wait

Acceptable at ~1s on Apple Silicon; not acceptable at 8s. Worth naming the free alternative: macOS SFSpeechRecognizer with requiresOnDeviceRecognition streams partial text live with no download at all. But its technical/code vocabulary is meaningfully worse than Whisper's, so I'd keep Whisper and buy the latency back with a smaller model rather than switch engines. Not worth the scope.

Surface coverage

Since it is a stated repo constraint: mobile is already handled natively via apps/mobile/src/features/voice-input/, desktop is this PR, and browser-only (app.t3.codes) silently gets nothing — ForkDictationControl returns null with no bridge. That is a defensible deferral rather than a gap, but it should be stated in the PR description instead of left implicit.

if (
!(target instanceof Element) ||
!root.current ||
target.closest("[data-fork-dictation-composer]") !==

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1. Ctrl+Shift+Space can start dictation but never stop it.

This scope check compares event.target.closest("[data-fork-dictation-composer]") against the control's own composer. But starting dictation sets voiceInputBusy -> ChatComposer.tsx:6026 passes disabled -> ComposerPromptEditor.tsx:1703 calls editor.setEditable(false) -> Lexical renders contentEditable={false} with no tabIndex, so the div becomes unfocusable and activeElement falls back to <body>.

On the next press target is document.body, body.closest(...) is null, the comparison fails, and the handler returns early. The stop branch on line 187 is unreachable from the keyboard — the advertised shortcut is one-way, and the user has to click the mic or the X to stop.

Cheapest fix: when instance.currentState.phase === "recording", skip the scope check (only one dictation can be active app-wide anyway), or capture the owning composer element at start and compare against that instead of against event.target.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f209dd3/90692f778: the window-level Ctrl+Shift+Space handler in useForkDictationController stops when recording and starts when idle/error, so it toggles both ways.

const instance = controller.current;
if (!instance || event.repeat || event.defaultPrevented) return;
const busy = voiceInputBlocksSubmission(instance.currentState);
if (event.key === "Escape" && busy) {

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2. Escape is swallowed app-wide while dictation is busy.

This branch has none of the composer-scope guarding the Space branch below it has, and it is a window capture-phase listener that calls stopPropagation(). busy includes preparing, so during the 466 MiB first-run download the command palette, dialogs and popovers cannot be dismissed for the entire duration of the download.

Apply the same data-fork-dictation-composer scope check the Space branch uses, or at minimum only stopPropagation() when the event originated inside the composer.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f209dd3/90692f778: Escape only cancels when the keydown target is inside the composer or is body/document, so dialogs and other editors keep their Escape.

};
controller.current = instance;
const visibility = () => {
if (document.hidden) void instance.appMovedToBackground();

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

3. Minimizing the window destroys the model download with no resume.

visibilitychange -> appMovedToBackground() -> controller.ts:308 invalidates the operation during preparing, which aborts the fetch. LocalSpeechEngine.ts:182 then rms the partial .download file in its finally, so there is nothing to resume from and retry restarts at 0%.

This is the single worst hit to "just works" in the PR: a 466 MiB download is exactly the moment a user is most likely to minimize the window and go do something else.

Two independent fixes, both worth doing: don't treat preparing as background-cancellable (the download does not need the window), and keep the partial file plus a Range request so a retry resumes.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f209dd3: dictation no longer cancels on visibilitychange; the recorder keeps going in the background and only microphone loss (onInterrupted) ends a recording.

Comment thread scripts/build-desktop-artifact.ts Outdated
<key>com.apple.security.cs.allow-jit</key>
<true/>
<!-- fork:begin fork-local-dictation — see .fork/customizations.yaml#fork-local-dictation -->
<key>com.apple.security.device.audio-input</key>

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

4. The macOS microphone entitlement is attached to the wrong branch.

com.apple.security.device.audio-input was added only inside renderMacPasskeyEntitlements, and macEntitlementsPath is only set when macPasskeySigning resolves (line 3799). A signed build without passkey configuration falls through to electron-builder's default template — I read it at app-builder-lib/templates/entitlements.mac.plist, and it contains only allow-jit, allow-unsigned-executable-memory and disable-library-validation. No audio-input, so getUserMedia is denied under the hardened runtime.

Note the contrast with NSMicrophoneUsageDescription, which this PR correctly places on the unconditional extendInfo at line 2745. The entitlement belongs at that same altitude — applied whenever the build is signed, not only on the passkey path.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d9b21ea + 65dfcf3. Signed mac builds now always get an entitlements plist: renderMacEntitlements() for the non-passkey path, and both renderers share MAC_SIGNED_ENTITLEMENT_KEYS (the three electron-builder defaults + com.apple.security.device.audio-input) so the passkey variant cannot drop the mic key. Covered by a new test in build-desktop-artifact.test.ts.

"!apps/desktop/prod-resources/browser-secret",
"!apps/desktop/prod-resources/browser-secret/**/*",
// fork:begin fork-local-dictation — see .fork/customizations.yaml#fork-local-dictation
"!apps/desktop/prod-resources/voice-input",

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

5. The helper ships twice — once in app.asar, once as an extra resource.

The helper is staged into stageResourcesDir = apps/desktop/resources/voice-input (line 3752), then copied to prod-resources (line 3780). These exclusions only cover prod-resources/voice-input, so apps/desktop/resources/voice-input/** is still packed into app.asar and emitted as an extra resource.

Compare browser-secret on the four lines directly above, which excludes both resources/ and prod-resources/ paths. Roughly 8 MB duplicated on darwin-arm64 (2.3 MB binary + 6.1 MB metallib).

Adding the matching !apps/desktop/resources/voice-input and !apps/desktop/resources/voice-input/**/* entries fixes it.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f209dd3: MAC_FILE_EXCLUSIONS now excludes both apps/desktop/resources/voice-input/** and apps/desktop/prod-resources/voice-input/**, so the helper ships once as an extra resource.

Comment thread apps/desktop/vite.config.ts Outdated

const repoEnv = loadRepoEnv();
/* fork:begin fork-local-dictation — see .fork/customizations.yaml#fork-local-dictation */
const voiceInputBuild = "node scripts/build-voice-input.mjs && ";

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

6. Desktop dev now hard-requires CMake and a C++ toolchain.

voiceInputBuild is prepended unconditionally to build, dev and dev:bundle. Anyone running the desktop app for a completely unrelated reason now downloads a whisper.cpp tarball from codeload and compiles it before Electron will start — or gets a hard failure if CMake (or, on macOS, the Xcode Metal toolchain) is missing.

For the dev tasks specifically this should be lazy or opt-in: the app runs fine without the helper, since ForkDictationControl already returns null when the bridge is absent. See the direction notes in the review body on fetching prebuilt whisper.cpp binaries instead of compiling.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f209dd3 and tightened in d9b21ea: desktop dev/dev:bundle run the helper build with --optional, which warns and skips when CMake or the Metal toolchain (xcrun -f metal) is missing; only artifact builds require them.

owner: string,
operation: (signal: AbortSignal) => Promise<T>,
): Promise<T> {
if (this.active) throw new Error("Voice input is still finishing. Try again shortly.");

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

7. The engine's single-slot lock is app-wide, but the session is per-renderer.

operate() throws when this.active is set, but exactly one LocalSpeechEngine is constructed per app in installVoiceInputIpc, while owner is scoped \${sender.id}:${requestId}``. With two desktop windows open, a transcription started in the second window is rejected outright rather than queued, and the user gets "Voice input is still finishing" with no way to retry that recording.

Either key the slot per sender.id, or queue rather than reject.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 65dfcf3: the engine queues instead of rejecting, so a second window's transcription waits for the slot rather than failing.

NodePath.join(this.options.cacheDirectory, "ggml-small.bin"),
"--file",
audio,
"--language",

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

8. --language auto is hardcoded even though the locale is already known.

prepare returns locale: navigator.language (ForkDictationControl.tsx:92), but controller.ts:394 only consumes it for English spacing heuristics in resolveTranscriptCommit — it never reaches whisper.

Auto-detection costs an extra detection pass and misfires on short clips, which is precisely the shape of a dictated prompt. Threading the known locale through to --language would be both faster and more accurate. This pairs with the model-choice note in the review body.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f209dd3: the helper is invoked with --language en (this fork targets English dictation; the ggml-base.en model is English-only anyway), so no detection pass runs.

});
}

private async verifyModel(path: string, signal: AbortSignal): Promise<boolean> {

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

9. The first dictation of every app launch re-hashes 466 MiB.

verifyModel streams the entire 487 MB file through SHA-256. It is memoized via modelVerified, so it runs once per main-process lifetime — but that once lands on the first dictation after each launch, under the "Preparing..." label the user is actively waiting on. On a slow disk that is several seconds of dead time on the exact action being requested.

Writing a sidecar stamp (size + mtime + verified hash) at download time and re-hashing only on mismatch removes the cost while keeping the integrity guarantee.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in f209dd3: verifyModel checks size plus a <model>.sha256 sidecar stamp written at download/verify time and only streams the full hash on a stamp mismatch.

@NoahHendrickson NoahHendrickson left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review

Verdict: solid foundation, not mergeable as-is. The architecture is right for the fork (device-local IPC, reuses upstream VoiceInputController, everything in custom//fork/ with fences and a manifest entry), and the engine code is careful. But there are three behaviors that will lose the user's speech in ordinary use, and the PR has never been exercised with a real microphone. Against the stated intent (free dictation that feels fast and good to use) it is currently correct but not fast-feeling: no live feedback while speaking, a cold whisper process per utterance, and a per-launch 466 MiB hash before the first recording can start.

What was verified locally

  • Checked out the PR head, vp i, ran the 3 PR test files: 15/15 pass.
  • apps/desktop and apps/web typecheck: 0 errors. Lint on new files: clean.
  • The snapshot.value + snapshot.expandedCursor pairing in readDraft is correct; it matches the existing addTerminalContext path in ChatComposer.tsx.
  • Did not build whisper.cpp or record audio (the PR itself says live mic/UI is unverified).

Blocking — these lose speech

1. The editor isn't frozen during recording, so any keystroke discards the transcript.
Upstream's controller rejects the transcript if the draft text or revision changed (resolveTranscriptCommit → "stale"). Mobile guards this with readOnly={voiceInput.freezesEditor} in both ThreadComposer.tsx and NewTaskDraftScreen.tsx. This PR only disables the send button. Dictate for two minutes, tap a key to fix a typo → "The draft changed while voice input was running. The transcript was not added." voiceInputFreezesEditor already exists in client-runtime; wire it to the editor like mobile does. (Letting the transcript insert at the captured cursor even if text changed elsewhere would be nicer on desktop, but that's an upstream controller change, so the freeze is the honest minimum.)

2. visibilitychange → appMovedToBackground() cancels recording when the window is merely covered.
On macOS, Electron/Chromium reports document.hidden === true when a window is fully occluded, not just minimized. A user who starts dictating and then brings a doc or terminal in front of T3 Code loses the recording. This is a mobile behavior (backgrounded app = audio session gone) that doesn't apply on desktop. Remove it, or restrict it to the preparing phase.

apps/web/src/custom/voice/ForkDictationControl.tsx L139–142:

const visibility = () => {
  if (document.hidden) void instance.appMovedToBackground();
};
document.addEventListener("visibilitychange", visibility);

3. useEffect(() => { if (props.disabled) controller.current?.cancel(); }, [props.disabled]) — disabled includes pendingUserInputs.length > 0 and isComposerApprovalState. If the agent asks a question or requests approval mid-dictation, the recording is silently cancelled. This should only prevent starting, not kill an in-flight recording.

Should fix before merge

4. vp run dev for desktop now hard-requires CMake + Xcode. build-voice-input.mjs throws when cmake is missing, and it's prepended to the dev and dev:bundle tasks. LocalSpeechEngine.prepare already handles a missing binary gracefully ("The local speech engine is missing"), so the dev task should warn-and-skip; only build-desktop-artifact should fail hard.

5. Per-launch full SHA-256 of the model delays the first recording. prepare() runs before recording starts (controller: preparing → recording), and verifyModel hashes 466 MiB on the first dictation of every app session. That's a visible "Preparing…" pause before the mic even opens — the wrong place to spend latency. Verify once at download and write a sidecar stamp; afterwards, size-check only. If tamper detection matters, do it in the background at IPC install time, not on the click path.

6. Escape is captured app-wide while busy. The keydown listener on window (capture phase) swallows Escape anywhere in the app during recording/transcribing, so an open dialog's Escape cancels dictation instead of closing the dialog. Scope it to the composer like the Ctrl+Shift+Space branch already does.

7. Whisper flags leave speed and quality on the table:

  • --language auto costs an extra detection pass and is the classic source of short-clip misdetection/hallucination. navigator.language is already available; pass its 2-letter code and only fall back to auto if unsupported.
  • No --flash-attn (-fa), a meaningful Metal speedup in whisper.cpp ≥ 1.7.
  • No --suppress-nst / --no-speech-thold. The silence guard in recordingToWav (> 0.0001) only catches a muted mic; a real noise floor passes through and Whisper Small emits "Thank you." or "Subtitles by…" on it.

8. Docs say "hiding the app cancels an active recording" — if #2 is fixed, update docs/user/composer.md.

Minor / nits

  • Electron.webContents.fromId(event.sender.id) — event.sender is already the WebContents.
  • prepare's signal abort listener isn't removed on the success → recording → cancel path (harmless, GC'd with the controller).
  • recordingToWav decodes into a 16 kHz OfflineAudioContext (which already resamples) and then renders again through a second one for the mono downmix. Correct, just two passes.
  • --threads 4 is fine on Metal; on Windows/Linux CPU it underuses a big box. Consider os.availableParallelism() capped.
  • whisper.cpp source tarball is pinned by commit SHA but not checksummed. Acceptable for a personal fork; noting for completeness.
  • Guard test uses source toContain matching — that's the established convention in __fork_guards__, no objection.
  • Screenshots/video deferred and live mic unverified. For a feature whose entire value is feel, one real recording pass before merge is warranted.

Intent review: "quality, efficient, free, feels fast and good to use"

Free & private: fully delivered. No key, no fees, audio never leaves the device, works offline after one download.

Quality: Whisper Small is a reasonable middle. For dictated coding prompts (identifiers, file names) it's noticeably weaker than large-v3-turbo. Options in the same HF repo: ggml-small-q5_1.bin (~190 MiB, ~same quality, faster load — recommended as the new default), or ggml-large-v3-turbo-q5_0.bin (~574 MiB, materially better, still fast on Metal).

Efficient: memory-wise yes (helper exits). Speed-wise the design pays a cold start every utterance: process spawn + model load + Metal init before inference begins. The PR's own numbers (0.6–1.3 s for 11 s of audio) are probably ~half startup.

Feels fast / good to use — this is where it falls short:

  1. No live feedback while speaking. Status is a static "Recording" label. Mobile has a waveform meter (voiceInputMetering.ts); desktop has nothing — no level, no elapsed time, and the 5-minute limit stops silently. At minimum add an elapsed counter (1 Hz interval, not a continuous animation) and an AnalyserNode level indicator from the stream that already exists.
  2. Nothing happens until you stop. Users perceive dictation speed as "how soon do words appear," not total wall time. A streaming/partial transcript is the single biggest feel improvement. whisper-stream exists but is fiddly; the cheap version is to transcribe on stop but keep the process warm (below).
  3. Cold start per utterance. Keep a whisper-server (ships in whisper.cpp) alive with a ~60 s idle timeout: first dictation pays the load, follow-ups are near-instant, memory is still released when you walk away. This preserves the "helper exits, frees memory" intent while removing the tax on rapid back-and-forth.
  4. First-run friction: 466 MiB before the first word. The quantized small model halves that.

Alternative worth a serious look: the mobile app already uses Apple's on-device SpeechAnalyzer / SpeechTranscriber (iOS 26), and the identical API exists on macOS 26. A small Swift helper (there's precedent in native/ for compiled helpers) would give: zero model download, OS-managed models, true streaming partial results, and no CMake/Xcode-Metal build step. whisper.cpp would remain as the Windows/Linux/older-macOS path. That likely gets closer to "feels fast" than any amount of whisper tuning, but it's a bigger change and a per-platform decision, so it belongs in its own PR, not bolted onto this one.

Suggested landing order: fix #1–#3 and #5–#6 in this PR (they're small), switch to small-q5_1 + explicit language + -fa, land it. Follow up with the warm server and recording feedback. Evaluate the Apple Speech helper separately.


Reviewed with Claude Fable 5.1 in Cursor.

xxxxxxxxxxxxx and others added 4 commits September 11, 2026 12:32
Restructure local dictation per PR review before more is layered on it.

Web: the composer now holds `useForkDictationController` (mirroring mobile's
`useVoiceInputController`) and reads `blocksSubmission` / `freezesEditor`
directly. The mic button only renders state. This removes the
`onBusyChange` callback, the mirrored `useState` + `useRef` busy flag, the
`key=` remount, and the `data-fork-dictation-composer` attribute from
ChatComposer, and fixes four behaviors on the way:

- covering or minimizing the window no longer cancels a recording or aborts
  the model download (desktop has no background audio session to lose)
- an approval or question arriving mid-dictation no longer silently kills
  the recording; `disabled` gates starting only
- Ctrl+Shift+Space can stop as well as start (the frozen editor drops focus
  to <body>, so stop/cancel accept that target)
- Escape only cancels dictation from the composer or an unfocused page, so
  dialogs keep their own Escape during the download

Desktop: fork IPC types move out of `client-runtime` into the desktop fork
and a renderer-local interface, and the one generic dispatcher becomes three
explicit handlers. The engine switches to `ggml-small.en-q5_1` (181 MiB,
English-only, same quality tier), passes `--language en`, `--flash-attn`,
`--suppress-nst`, scales threads to the host, and verifies the model once
with a sidecar stamp instead of re-hashing on the first dictation of every
launch.

Build: macOS only (whisper.cpp publishes no prebuilt macOS CLI, so compile
stays). The desktop `dev`/`build` tasks skip the helper when CMake is
missing; only the packaged artifact requires it. Task commands are literal
strings again so knip can see the script entries. The helper is a mac-only
extraResource and is excluded from app.asar at both staging paths.

Implemented with Claude Fable 5.1 in Cursor.

Co-authored-by: Cursor <cursoragent@cursor.com>
The composer-owned hook read refs during render, which the fork-owned lint
gate rejects, and its editor-disable fence sat ahead of isConnecting where
the fork-new-agent-draft guard expects it.

Read the latest composer closures through useEffectEvent, construct the
controller in a lazy useState initializer, drop the redundant prompt
revision (the controller already compares owner and text), and move the
freezesEditor fence after isComposerApprovalState.

Claude Fable 5.1 via Cursor

Co-authored-by: Cursor <cursoragent@cursor.com>
…l meter

First live run: fetching the recording's blob: URL failed in the Electron
renderer (and the packaged CSP has no blob: in connect-src), so every
transcription ended in "Failed to fetch". The recorder now keeps the Blob it
built and reads it directly.

Composer controls reworked from that session: the mic becomes a check with an
X beside it while recording, a ten-bar white level meter (dB-mapped so quiet
speech registers, DOM-written at 20 Hz so the composer does not re-render)
sits to their left, the check becomes a spinner while transcribing, attach
hides while dictation is in flight, and the transient "Preparing…" label is
gone. The dictation buttons take the fork's 24px composer-action box so the
prompt row stays 44px, and the ghost actions render pure white in dark mode.

Claude Fable 5.1 via Cursor

Co-authored-by: Cursor <cursoragent@cursor.com>
Swaps the composer's paperclip for Phosphor's plus via the lucide shim,
inside the fork-composer-shell fence.

Claude Fable 5.1 via Cursor

Co-authored-by: Cursor <cursoragent@cursor.com>

@NoahHendrickson NoahHendrickson left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the current head, 1ff09618b. Changes needed before merge. Three inline findings cover transcript loss/misdirection when a question or approval arrives, an unusable dictation control on Windows/Linux, and the remaining mandatory Metal-toolchain dependency in desktop dev.

The previous microphone-entitlement finding also remains unresolved: renderMacPasskeyEntitlements is only used when passkey signing is configured. Signed macOS builds without that configuration fall back to electron-builder's entitlement template, which lacks com.apple.security.device.audio-input. I checked the installed electron-builder 26.15.6 template and signing selection; Apple documents this entitlement as permitting audio input under the hardened runtime. Apply it independently of passkey configuration.

Validation: checked out the PR head in an isolated worktree and ran vp test run apps/desktop/src/fork/voice/LocalSpeechEngine.test.ts apps/web/src/custom/voice/BrowserVoiceRecorder.test.ts apps/web/src/__fork_guards__/forkLocalDictation.test.ts scripts/build-desktop-artifact.test.ts. Result: 85 passed, 1 failed. All 16 dictation tests passed; packaging had 69 passes and the Windows cross-architecture native-probe assertion failure already documented in the PR. I did not build a signed artifact or exercise a real microphone/UI.

Surface review: traced web/desktop composer integration, local IPC and remote-environment behavior, macOS packaging, and unsupported desktop platforms. Browser-only and mobile do not receive the preload bridge. Earlier fixes for editor freezing, window visibility, and shortcut scope are present; the pending-question/approval guarantee still fails at the draft binding described inline.

Reviewed with GPT-6 in the Codex harness.

Comment on lines +1783 to +1786
readDraft: () => readComposerSnapshot(),
commitDraft: (text, cursor) => {
const collapsedCursor = collapseExpandedComposerCursor(text, cursor);
onPromptChange(

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Keep dictation bound to the original prompt when pending input arrives

Start recording while an agent is running, then let a question or approval arrive before stopping. ComposerPromptEditor.value switches to the question's customAnswer or an empty approval value (lines 5964–5969), so readComposerSnapshot() now reads a different draft despite the unchanged owner key. With a nonempty original prompt, resolveTranscriptCommit reports stale and discards the entire transcript. With an empty original prompt and an empty question answer, the stale check passes, but onPromptChange routes the transcript into the question answer (lines 2764–2778), or silently drops it for a choice-only question. Gating only new starts does not preserve an in-flight dictation. Capture the original prompt target/cursor and read/commit that target independently of the editor's temporary approval/question mode.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed and fixed in d9b21ea. While an approval or question is showing, readDraft/commitDraft now read and write the prompt draft directly (promptRef/setPrompt/setComposerCursor) instead of going through the editor snapshot and onPromptChange, so a dictation that started before the pending input arrived lands in the original prompt and passes the stale check.

Comment thread apps/desktop/src/preload.ts Outdated
Comment on lines +387 to +389
voiceInput: {
prepare: (requestId: string): Promise<void> =>
invokeVoiceInput<void>("fork:voice-prepare", { requestId }),

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Expose voice input only on supported desktop platforms

This bridge is exposed unconditionally, and readForkVoiceInputBridge() treats the presence of transcribe as availability. Consequently Windows and Linux desktop users get the microphone and Ctrl+Shift+Space handler even though build-voice-input.mjs skips those platforms and packaging only includes the helper on macOS. Every attempt freezes the composer briefly and fails with 'The local speech engine is missing. Rebuild or reinstall the desktop app.' Reinstalling cannot fix an unsupported platform. Gate this bridge with the existing client platform (or return an explicit unavailable capability) so these clients do not advertise dictation.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d9b21ea: voiceInput is undefined unless clientPlatform === "darwin", and readForkVoiceInputBridge() already treats a missing bridge as no dictation, so Windows/Linux never see the mic or the shortcut.

Comment on lines +53 to +56
try {
NodeChildProcess.execFileSync("cmake", ["--version"], { stdio: "ignore" });
} catch {
if (values.optional) {

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Include the Metal toolchain in the optional-build preflight

--optional only handles a missing cmake executable. On a macOS contributor machine with CMake and Command Line Tools but no full Xcode/Metal toolchain, that check succeeds and the later uncaught cmake --build fails: the pinned ggml-metal-lib target runs xcrun ... metal and metallib. Since this script is prepended with && to desktop build, dev, and dev:bundle, none of those tasks reaches the app. Preflight the required Apple toolchain too and warn/skip when optional, while keeping artifact packaging strict.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d9b21ea: the preflight now checks xcrun -f metal alongside cmake --version; --optional warns and skips when either is missing, non-optional (artifact) builds still fail loudly naming the missing tool.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thermo-nuclear review

Do not approve.

The last review's four structural findings are gone. ChatComposer now holds useForkDictationController and reads blocksSubmission / freezesEditor locally. ForkDictationControl only renders. Fork IPC lives in the desktop fork plus a renderer-local bridge. Three typed handlers replaced the generic dispatcher.

What remains is two lying boundaries. The skill bar is not met until those contracts match the code.

  1. Microphone entitlement is bolted onto renderMacPasskeyEntitlements. A signed Mac build without passkey config never writes that plist. Hardened runtime then denies getUserMedia while NSMicrophoneUsageDescription still claims the app records.
  2. LocalSpeechEngine.operate() is a process-global mutex dressed as a per-window owner. withWindow builds `${sender.id}:${requestId}`; the slot ignores that owner and rejects anyone. Pick one invariant.

Not findings: ChatComposer already ≫1k so the crossing rule does not newly fire; the detail overlay; the PlusIcon swap; already-fixed CMake / asar / language / hash / shortcut / Escape / visibility issues.

Open in Web View Automation 

Sent by Cursor Automation: Thermo nuke 4.6

Comment thread scripts/build-desktop-artifact.ts Outdated
Comment on lines +1354 to +1356
<key>com.apple.security.device.audio-input</key>
<true/>
<!-- fork:end fork-local-dictation -->

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is feature logic leaking into a shared path. NSMicrophoneUsageDescription sits on the unconditional Mac extendInfo. com.apple.security.device.audio-input does not. It was bolted into renderMacPasskeyEntitlements, and macEntitlementsPath is only written when macPasskeySigning resolves (~3807). A signed Mac build without passkey config falls through to electron-builder's default template (jit / unsigned-exec-memory / disable-library-validation). Hardened runtime then denies getUserMedia while the usage string still claims the app records. The function name now lies about what it contains, and a later passkey-only edit can drop the mic key by accident.

Write a signed-Mac entitlements plist at the same altitude as extendInfo, always include audio-input, and merge the passkey keys on top when that config exists. Do not hide a runtime capability inside a function that only runs for Clerk associated-domains.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 65dfcf3 (see the thread above): one shared MAC_SIGNED_ENTITLEMENT_KEYS constant feeds both plists, and the non-passkey signed path writes its own plist at staging time. renderMacPasskeyEntitlements keeps its name because it still renders the passkey-specific keys on top of the shared set.

Comment on lines +116 to +128
private async operate<T>(
owner: string,
operation: (signal: AbortSignal) => Promise<T>,
): Promise<T> {
if (this.active) throw new Error("Voice input is still finishing. Try again shortly.");
const active = { owner, abort: new AbortController() };
this.active = active;
try {
return await operation(active.abort.signal);
} finally {
if (this.active === active) this.active = null;
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This owner model lies. withWindow builds ${sender.id}:${requestId} and comments that engine work is owned by the window that asked for it. operate() then ignores that owner and rejects anyone if this.active is set. Cancel is owner-scoped; the lock is process-global. Two desktop windows get "Voice input is still finishing" with no retry of the second recording, and a reader cannot tell whether the contract is per-window, per-request, or app-wide.

Pick one invariant and make the types match it. If whisper is a single process-wide slot, drop the owner fiction: operate() takes no owner, active is just { abort }, and cancel is "abort the current slot if this request still holds it." If windows really own work, key the slot by sender.id (a Map, not one nullable). Do not keep a per-window owner string in front of an app-wide mutex.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, fixed in 65dfcf3. The invariant is now explicit: whisper is one process-wide slot, requests queue FIFO behind it, and each request is keyed (${sender.id}:${requestId}) so cancel — from the renderer or its window going away — aborts only that request, running or waiting. owner is renamed requestKey throughout and the "still finishing" rejection is gone; tests updated for the queue semantics.

xxxxxxxxxxxxx and others added 2 commits September 11, 2026 16:41
Review follow-ups on PR #136:

- A dictation started before an approval or question arrived read the
  editor, which by then held the answer field, so the transcript was judged
  stale or routed into the answer. The composer now reads and commits the
  prompt draft directly while that mode is showing.
- The preload bridge is only exposed on macOS, so Windows and Linux builds
  never advertise a mic they cannot back.
- The optional helper build preflights Xcode's Metal toolchain as well as
  CMake, since Command Line Tools alone cannot compile the shader library.
- Signed macOS builds without passkey signing get an entitlements plist too,
  carrying the microphone entitlement alongside electron-builder's defaults.

Claude Fable 5.1 via Cursor

Co-authored-by: Cursor <cursoragent@cursor.com>
The engine rejected any request while another was running, so a second
window's transcription failed outright with no way to retry it. Whisper
stays one process at a time app-wide; requests now queue FIFO and each is
cancellable by its own key, running or waiting.

Both macOS entitlement plists also render their shared hardened-runtime and
microphone keys from one constant so the passkey variant cannot drop the
microphone by accident.

Claude Fable 5.1 via Cursor

Co-authored-by: Cursor <cursoragent@cursor.com>
@NoahHendrickson

Copy link
Copy Markdown
Owner Author

Code review — max level

Reviewed at head 65dfcf33. Note that d9b21ea98 and 65dfcf338 landed while this review was running; findings that those two commits already fixed are listed at the bottom rather than above, so nothing here should be stale. Every item below was re-verified against 65dfcf33.


Blocking

1. The dictation controls unmount mid-recording — apps/web/src/components/chat/ChatComposer.tsx:5140

ForkDictationControl lives only inside composerPrimaryActionSlot, which ComposerPromptRow renders only when showInlinePrimaryAction — and resolveComposerShellVisibility (apps/web/src/custom/ComposerShell.tsx:21) defines that as !approvalPending && !mobilePendingActionsVisible.

Start dictating; the agent raises an approval. The whole action cluster unmounts, but the hook lives in ChatComposer, so the recording keeps running: blocksSubmission stays true, the editor stays disabled, and send reads "Finish or cancel dictation before sending." — with no mic, no check, no X. The macOS mic indicator stays lit until the five-minute auto-stop. d9b21ea98 fixed the transcript path for exactly this case (commitDraft now writes promptRef directly) but not the control path, so the manifest's "settled by the user, not silently dropped" still fails. The only exit left is the undocumented window-level Escape.

2. modelVerified is latched for the process lifetime — apps/desktop/src/fork/voice/LocalSpeechEngine.ts:153

Nothing resets it and nothing re-stats the file. Dictate once (model downloads, flag latches), then delete the model to reclaim the 181 MiB the tooltip advertises — or let any cleaner do it. Every later dictation: prepare returns at :153 without downloading, transcribe's if (!this.modelVerified) gate at :237 also passes, whisper-cli exits non-zero on the missing --model path, and the recording is destroyed with "Local transcription failed. Please record again." Repeats forever; only an app restart recovers. A stat before the early return, or clearing the flag on a whisper failure, closes it.

3. The cancel X is hidden exactly during transcribing — apps/web/src/custom/voice/ForkDictationControl.tsx:85

busy && !transcribing hides the X in the one phase where the mic is inert (pointer-events-none, :122) and the composer is frozen. A five-minute dictation takes whisper Small tens of seconds. The only escape is Escape, which appears nowhere in the UI (the tooltip reads "Transcribing") and works only because focus happened to fall to <body>. CLAUDE.md: "If you added a way in, add the way out and the way to see it."

4. --optional still does not cover the download or the build — apps/desktop/scripts/build-voice-input.mjs:84

65dfcf33 added the xcrun -f metal probe, which closes the missing-toolchain case. But the fetch at :84 and both cmake invocations (:98, :117) throw unguarded, and apps/desktop/vite.config.ts:23 runs this script first in an && chain ahead of vp pack. Offline, or any compile error, kills vp run dev for the desktop app before it starts — which is the outcome the file's own comment says --optional exists to prevent. Wrapping :80-133 in a try/catch that honors values.optional fixes it.

5. Scope: two unrelated restyles ride along — apps/web/src/components/chat/ChatComposer.tsx:5116, apps/web/src/theme.custom.css:722

<PlusIcon /> replaces <PaperclipIcon /> under the fork-composer-shell fence, and the new dark-mode rule forces --control-icon-color: #ffffff on a selector list that includes [data-fork-composer-action="attach"] — changing an existing control's icon color. The head commit is titled "style(web): attach button uses a plus glyph." Neither change is in fork-composer-shell's intent: nor asserted by forkComposerShell.test.ts, so the next upstream sync restores the paperclip with CI fully green. CLAUDE.md: "Do not mix unrelated fixes into one change"; "One concern per PR."


Correctness

6. errorAction is dropped, so a denied microphone is a dead end — apps/web/src/custom/voice/useForkDictationController.ts:241

The hook returns error but never errorAction. With mic access denied in System Settings, queryMicrophonePermission returns {granted: false, canAskAgain: false}, the controller sets "Microphone access is required for voice input." with errorAction: "settings", and getUserMedia (whose message does mention system settings) is never reached, so detail stays null. The popover renders a bare sentence plus Dismiss; clicking the mic again reproduces it forever. Mobile — the stated model for this hook — handles it at ComposerDictationControl.tsx:403 with an Open-microphone-settings action.

7. Escape is swallowed app-wide while an error banner is up — apps/web/src/custom/voice/useForkDictationController.ts:210

The Escape branch fires for every non-idle phase, including the sticky error phase, then calls preventDefault() + stopPropagation() on a capture-phase window listener. After one error ("No speech was detected.") the phase stays error indefinitely, so the next Escape pressed with focus in the composer or on <body> — the normal target, since freezing the editor drops focus there — is consumed to dismiss the dictation banner. ComposerBannerStack's Escape (:171) and any Base UI popover anchored in the composer never see it. Restricting the branch to the blocking phases fixes it.

8. The deferred focusAt has no staleness guard — apps/web/src/components/chat/ChatComposer.tsx:1809

focusAt is not a pure focus call — ComposerPromptEditor.tsx:1802 ends by re-committing its own snapshot through onChange. The only other deferred focusAt in this file (:2883-2888) guards it with if (promptRef.current !== next.text) return; and a comment saying exactly why: "it drags the caret back behind what was typed since." Transcript lands, the editor unfreezes in the same commit, the user types within the ~16 ms window, and one frame later the caret is yanked back to the end of the transcript and onPromptChange re-fires.

9. Focus is restored only on the successful-commit path — apps/web/src/components/chat/ChatComposer.tsx:1796

The approval/question branch returns before the rAF, and every cancel or error exit leaves focus on <body> with the caret lost. That is precisely why the key handler has to be window-level in the first place — so it is worth closing the loop the same way.

10. Raw Node errors reach the composer, with absolute home paths — apps/desktop/src/fork/voice/VoiceInputIpc.ts:77

If whisper-cli exits 0 without writing the transcript, readFile rejects with ENOENT: … open '/Users/<name>/Library/Application Support/…/recording-a1b2c3/transcript.txt'. That becomes VoiceInputError.message → the IPC error field → new Error(result.error) in the preload → onDetail, and since the hook prefers detail over state.error it wins over the curated "Could not transcribe this recording." Same for mkdtemp EACCES/ENOSPC. Every other engine failure has a hand-written message, so this looks unintended — and it leaks the user's home directory name.

11. Temporary recordings survive a crash — apps/desktop/src/fork/voice/LocalSpeechEngine.ts:238

mkdtemp(… "recording-") is cleaned only by the finally at :271, and the scope finalizer (VoiceInputIpc.ts:50) is a synchronous engine.dispose() that aborts without awaiting the rm. Force-quit or crash during inference and .../voice-input/recording-XXXXXX/audio.wav persists with up to 9.6 MB of the user's voice. No startup sweep exists anywhere in apps/desktop, and docs/user/composer.md ships the opposite promise: "temporary recordings are deleted after processing."

12. The model cache bypasses --home-dir isolation — apps/desktop/src/fork/voice/VoiceInputIpc.ts:48

Electron.app.getPath("userData") is the only getPath( call in all of apps/desktop/src. Every other desktop-managed directory goes through environment.stateDir (DesktopWslServerTree.ts:127, DesktopSnapShot.ts:726, DesktopConnectionCatalogStore.ts:386). Line 38 already binds environment. As written, a --home-dir-isolated dev run silently shares the real install's 181 MiB cache instead of its own.

13. new AudioContext() sits outside the best-effort try — apps/web/src/custom/voice/BrowserVoiceRecorder.ts:145

The comment at :141 says "Best effort: the recording must not depend on the level meter," but the try begins at :148 and covers only createMediaStreamSource(...).connect(analyser). stopMeter fires close() without awaiting, so back-to-back start/cancel cycles can hold several unretired contexts against Chromium's per-document cap. The next constructor throw propagates out of record() (called at :124, after recorder.start() at :122) into the controller's catch — the session dies with "Could not start voice recording." for a purely decorative meter, leaving a live MediaRecorder that release() never stops.

14. Every MediaRecorder error is reported as a mic disconnect — apps/web/src/custom/voice/BrowserVoiceRecorder.ts:120

recorder.addEventListener("error", () => this.onInterrupted()) discards event.error, and onInterrupted is wired to the fixed string "Microphone disconnected. Please record again." A mid-recording encoder failure (SecurityError, InvalidModificationError, disk pressure) sends the user to check cables on a healthy mic. The class already has an onError(message) channel for this, used only for getUserMedia at :86.

15. Per-session recorder state resets only on the success path — apps/web/src/custom/voice/BrowserVoiceRecorder.ts:93

this.uri = null sits after the getUserMedia try/catch, whose catch rethrows at :91, so a denied mic on a later session leaves this.uri pointing at an already-revoked blob URL that releaseResources() re-adds and double-deletes. Benign today, but finishRecording (completedUri ?? recorder.uri ?? this.recordingUri) would happily adopt that dangling handle. this.stopped has the same shape — never reset in prepareToRecordAsync. Both belong at the top of prepareToRecordAsync and in release().

16. The five-minute bound is hardcoded twice with one byte of slack — apps/desktop/src/fork/voice/LocalSpeechEngine.ts:22, apps/web/src/custom/voice/BrowserVoiceRecorder.ts:48

44 + 16_000 * 2 * 300 vs Math.min(16_000 * 300, …). They agree today: the max recording is exactly 9,600,044 bytes and the check is >, so it passes by one byte. But both 300s are literals while VOICE_RECORDING_LIMIT_SECONDS (packages/client-runtime/src/voice-input/controller.ts:5) is what actually arms the recorder. Raise the shared limit upstream and every full-length recording is rejected in the main process after the user has spent five minutes talking.

17. downloadPercent is never cleared on completion — apps/web/src/custom/voice/useForkDictationController.ts:155

set.download(null) runs only on the transition into preparing. After bridge.prepare resolves at 100% the controller is still preparing while it runs configureRecording() and recorder.prepareToRecordAsync() — and the latter raises the macOS mic permission dialog on first use. "Downloading model 100%" sits on screen for as long as the user takes to answer it.


UX / layout

18. The right inset shrinks as the action cluster grows — apps/web/src/components/chat/ChatComposer.tsx:2252

showComposerAttachAction doubles as the resting composer's right-inset reservation (:5967). Press the mic on a resting composer with the context meter off: blocksSubmission flips true, the flag goes false, and the prompt row drops from pr-20 to pr-12 — 32 px less — while the cluster simultaneously gains the 10-bar meter, the X, and a max-w-36 "Downloading model 42%" label. Prompt text runs under the controls for the whole session, then snaps back.

19. The composer is frozen for the entire one-time model download — apps/web/src/custom/voice/useForkDictationController.ts:242

voiceInputBlocksSubmission includes preparing (controller.ts:15), and freezesEditor is the same predicate. On a slow connection one press of the mic — or an accidental Ctrl+Shift+Space — disables the prompt editor and sets sendDisabledReason for minutes. The user cannot type or send at all until they notice the X. On mobile preparing is instantaneous, so upstream's semantics were never exercised this way. Related: the download is awaited before the microphone opens, though only transcribe needs the model.

20. Ctrl+Shift+Space bypasses the keybinding registry — apps/web/src/custom/voice/useForkDictationController.ts:213

apps/web/src/keybindings.ts:220 resolves shortcuts and runs a conflict pass surfaced in KeybindingsSettings.tsx; ChatComposer.tsx already imports it and uses it in its own capture-phase window keydown effect. This chord registers none of it, so it is unrebindable, invisible in Settings, and silently wins over a user binding on the same chord (capture phase + stopPropagation). Matching on event.code === "Space" also skips the layout normalization in shortcutKeyFromEvent that every other shortcut gets, and the label is hardcoded in three places, so it lies after any remap.


Performance

21. A never-idle CSS transition on ten elements for up to five minutes — apps/web/src/custom/voice/ForkDictationControl.tsx:48

BrowserVoiceRecorder.ts:155 runs a 50 ms setInterval for the whole recording; each tick rewrites transform on all ten bars, and each bar carries transition-transform duration-75 — so a new compositor animation starts before the previous finishes and a transition is always in flight. On a 120 Hz display a five-minute dictation is ~36,000 frames across ten promoted layers, plus 200 style mutations/sec and a second AudioContext opened purely for the meter. CLAUDE.md: "No continuously repainting animations; they peg the GPU on high-refresh displays." Dropping the transition leaves the meter reading fine as discrete 50 ms steps. The setInterval also keeps sampling while the window is occluded, which requestAnimationFrame would not.

22. The download percent re-renders the composer ~101 times — apps/web/src/custom/voice/useForkDictationController.ts:171

setDownloadPercent is a useState slice in a hook called from ChatComposer's body, and LocalSpeechEngine.ts:177 emits one IPC event per distinct integer percent — each a separate macrotask with no React batching. That is ~101 full re-renders of a 4,600-line component during the one-time download. Coalescing on the emitter (every 5%) or routing the percent through the same DOM-write channel the meter already uses both fix it.

23. A redundant second OfflineAudioContext render pass — apps/web/src/custom/voice/BrowserVoiceRecorder.ts:50

The decode context is constructed at 16 kHz, so decodeAudioData already resampled to mono 16 kHz (the capture is constrained to channelCount: 1); the second graph then re-renders it to mono 16 kHz. For a five-minute clip that is an extra ~19 MB buffer and 4.8 M frames through the graph for a copy. decoded.getChannelData(0).subarray(0, length) is zero-copy and already clamps; keep the render path as the fallback for a device that ignores the mono constraint.


Fork bookkeeping and tests

24. Manifest gaps — .fork/customizations.yaml:41, :62

files: omits apps/web/src/theme.custom.css even though the mic's sizing, hover and icon color are implemented there — every one of the eleven other entries that styles through that file lists it. And apps/desktop/src/fork/voice/LocalSpeechEngine.test.ts is listed only under verify:, so .fork/lint-owned.mjs never sees it (apps/desktop/src/fork is not a fork-owned directory, and the gate reads entry.files). Compare fork-clerk-launch-resilience, which lists its desktop test under files:. The CSS omission passes silently because customizationsManifest.test.ts only validates fence→manifest and those rules carry no fence — leaving theme.custom.css:732-740 with zero guard coverage and no recorded intent.

25. The guard pins expression fragments — apps/web/src/__fork_guards__/forkLocalDictation.test.ts:66

toContain("dictation.blocksSubmission ?") and toContain("dictation.freezesEditor ||") break when prettier reflows the ?? chain at ChatComposer.tsx:1813, yet still pass if the expression moves into a branch that never evaluates — the actual regression the guard exists to catch. Lines 74-75 have the inverse problem: expect(runtimeExports).not.toContain("fork") fails the moment upstream writes the word "fork" in a comment. Other fork guards pin values (forkAppIdentity.test.ts:58) or parsed CSS selectors (forkComposerShell.test.ts:651).

26. deleteRecording is stubbed in both suites — apps/web/src/custom/voice/BrowserVoiceRecorder.test.ts:54

Production wires deleteRecording: releaseRecording, which does recordings.delete(uri) and URL.revokeObjectURL(uri). The recorder suite substitutes (uri) => URL.revokeObjectURL(uri) and forkLocalDictation.test.ts:38 a no-op — delete recordings.delete(uri) from the source and every test still passes, while each dictation retains a ~9.6 MB Blob in the module Map for the renderer's life. Neither suite calls resetVoiceInputGlobalsForTests(), which upstream's own controller.test.ts:153 uses to guard the module-level activeSession singleton; without it one test that throws before releaseResources() cascades into "Another voice recording is already active." for the rest of the file.


Worth a check, not asserted

27. Use-before-declaration in the new hook argument, under React Compiler — apps/web/src/components/chat/ChatComposer.tsx:1778

The four closures reference composerFormRef (:2018), setPrompt (:2384), onPromptChange (:2771) and readComposerSnapshot (:2903). Plain-JS TDZ is not hit — they are only called later — and this is the only such site in the file. But apps/web/vite.config.ts:190 enables reactCompilerPreset(), and the compiler places memo-block dependency checks at the closure's creation point. If it emits if ($[k] !== readComposerSnapshot) at :1778 the component throws on every render; if its validation pass catches the hoist instead, it bails out of compiling ChatComposer entirely, silently dropping memoization on the app's hottest component. Cheap to settle by watching for a bailout in a dev build. Moving the hook down is not the fix — sendDisabledReason (:1813) reads dictation.blocksSubmission at render time — ref mirrors are.

28. useEffectEvent stored in a mount-time object — apps/web/src/custom/voice/useForkDictationController.ts:177

React's contract is that the returned function is called directly from an effect or handler and never stored or passed as a value; here it is handed to a plain class graph inside a useState initializer. BranchToolbar.tsx:299-301 carries the comment "A render-synced mirror instead of useEffectEvent: the compiler memoizes the event callback, which left observers reading the first render's null element forever," and the mobile hook this one mirrors (useVoiceInputController.ts:87) uses a latestInputRef. If the compiler memoizes here, readDraft/commitDraft resolve against the mount-time readComposerSnapshot and onPromptChange. The useState initializer also double-invokes under StrictMode, building and orphaning a second recorder/controller pair per composer mount.


Already fixed by the two commits that landed during review

Listed so they are not re-raised: the non-macOS mic button (preload.ts now gates voiceInput on clientPlatform !== "darwin"); the transcript being discarded on a choice-only question (commitDraft now writes promptRef/setPrompt directly instead of routing through onPromptChange); cancel-then-retry reporting "Voice input is still finishing" (the engine now queues behind a per-request AbortController map rather than one global slot, which also softens the renderer-reload orphan to a wait instead of a rejection); and a missing Metal toolchain hard-failing dev (the probe now checks xcrun -f metal).

Checked and clean

The pinned whisper.cpp CLI accepts every flag passed (against the pinned revision's examples/cli/cli.cpp); the model URL, byte count and SHA-256 match Hugging Face exactly; Schema.Uint8Array in effect v4 is instanceOf and survives both clone hops without walking 9.6 M elements; the WAV offsets in encodeVoiceWav and validateVoiceWav agree field-for-field and the sample clipping gives the correct asymmetric int16 range; download metering handles overrun, truncation and percent dedup correctly; the .sha256 sidecar stamp correctly avoids re-hashing 181 MiB per session; the download streams to disk rather than buffering; MAC_VOICE_INPUT_EXTRA_RESOURCES resolves correctly through the staging copy and native/voice-input/build/ is gitignored; the new manifest entry passes customizationsManifest.test.ts; all 25 fences are format-exact and balanced; and the PlusIcon swap resolves correctly through the fork's lucide→Phosphor alias.


Reviewed by Claude Opus 5 (1M context) via Claude Code, /code-review max.

…d-mic exit, lighter meter

Addresses the max-level review on PR #136:

- A recording whose controls unmount (approval replaces the action cluster)
  is transcribed on the spot instead of running on with nothing to end it.
- Denied microphone permission offers an Open microphone settings action
  via a fork IPC channel; the renderer's openExternal allows no OS schemes.
- Escape only cancels during blocking phases, so a lingering error banner
  no longer eats Escape from other composer UI. Cancel and errors return
  focus to the editor; the deferred focusAt after a commit is skipped when
  typing has already moved on.
- Level bars drop their CSS transition (a new value every 50 ms kept one in
  flight for the whole recording) and pause sampling while hidden.
- Recorder: the level meter is fully best-effort, MediaRecorder errors say
  what failed instead of blaming the mic, per-session state resets on entry,
  the length bound comes from VOICE_RECORDING_LIMIT_SECONDS, and mono 16 kHz
  audio skips the redundant OfflineAudioContext render.
- Engine: the model is re-verified each session (a stat plus the sidecar
  stamp) so a deleted model downloads again instead of failing until
  restart; stale recording-* scratch dirs are swept on the first session;
  the cache lives under the desktop stateDir; raw Node errors with paths
  never reach the composer.
- Optional helper builds skip on download or compile failure too.
- Manifest lists theme.custom.css and the engine test as owned files and
  records the plus glyph and white ghost actions under fork-composer-shell,
  with guards that pin them.

Claude Fable 5.1 via Cursor

Co-authored-by: Cursor <cursoragent@cursor.com>
@NoahHendrickson

Copy link
Copy Markdown
Owner Author

Thanks — verified all 28 against 65dfcf33. Fixed in e96f9f892 unless noted.

Blocking

  1. Controls unmount mid-recording — real. ChatComposer now mirrors showInlinePrimaryAction (dictationControlsVisible) and, when it goes false, calls the hook's settleWithoutControls(): a recording is transcribed on the spot (it lands in the prompt draft via the d9b21ea98 path), a pending prepare is cancelled. Guarded.
  2. modelVerified latched — real. The flag is gone; prepare and transcribe both call verifyModel, which is a stat plus the sidecar stamp read, so a deleted model downloads again. Test added.
  3. X hidden while transcribing — not changed; this is the maintainer's explicit design (spinner replaces the check, X hidden). Escape still cancels, and the label/tooltip now reads "Transcribing (Escape to cancel)" so the way out is visible.
  4. --optional didn't cover download/build — real. The whole download + configure + build is wrapped; optional warns and exits 0, artifact builds still throw.
  5. Plus glyph / white ghosts riding along — both were requested by the maintainer in this PR, so they stay. The bookkeeping gap was real: fork-composer-shell's intent now records the plus glyph and the dark-mode white ghost actions, and forkComposerShell.test.ts pins <PlusIcon /> after the attach stamp, rejects <PaperclipIcon />, and finds the --control-icon-color: #ffffff rule.

Correctness
6. errorAction dropped — real. Hook returns it; the popover shows "Open microphone settings" for "settings", served by a new fork channel fork:voice-open-microphone-settings (shell.openExternal on the Privacy_Microphone deep link — the renderer's openExternal only allows web/editor schemes).
7. Escape swallowed in error phase — real. The Escape branch is restricted to preparing/recording/transcribing.
8. focusAt staleness — real. The rAF now bails if promptRef.current !== text.
9. Focus only on success — the hook takes focusEditor (the composer's focusComposer); called on cancel (button, Escape) and on the error transition. The approval/question commit branch intentionally does not refocus: the editor is showing the answer field.
10. Raw Node errors with paths — real. withWindow's catch passes through only errors without a code (engine-authored messages); Node errors become "Local dictation failed."
11. Temp recordings survive a crash — real. sweepStaleRecordings() removes recording-* under the cache on the first session of each launch. Test added.
12. Cache bypasses --home-dir — real. Now environment.stateDir/voice-input.
13. new AudioContext() outside the try — real. Constructor and graph wiring are inside one try; a half-built context is closed on failure.
14. Every MediaRecorder error = disconnect — real. onInterrupted now carries a message; track ended says disconnected, error reports event.error's message.
15. Per-session state — uri and stopped reset at the top of prepareToRecordAsync. Not in release(): the controller releases the session before it reads recorder.uri (controller.ts:358-360), so nulling it there would break every transcription.
16. Five-minute bound duplicated — both sides now derive from VOICE_RECORDING_LIMIT_SECONDS (desktop already depends on client-runtime).
17. downloadPercent never cleared — real. Cleared as soon as bridge.prepare resolves.

UX / layout
18. Right inset shrinks — real. showComposerAttachAction is back to upstream's definition (it is the inset reservation); the attach button is hidden at its render site instead, and the resting inset is pr-36 while dictating.
19. Composer frozen during download — not changed. It is one-time, the X is on screen, and it is upstream's voiceInputFreezesEditor semantics; letting typing race a pending draft snapshot is what the stale check exists to prevent.
20. Keybinding registry — deferred. Keybinding ids are packages/contracts schema, which this fork does not change from a frontend PR. Nothing in the default map uses Ctrl+Shift+Space; it stays a fixed chord for now.

Performance
21. Never-idle transition — real. transition-transform removed; sampling also pauses while document.hidden.
22. ~101 composer re-renders — real. Percent is floored to 5% steps before it hits state.
23. Redundant render pass — real. Mono 16 kHz output uses getChannelData(0).subarray(0, length); the graph render remains as the fallback.

Bookkeeping
24. Manifest gaps — theme.custom.css and LocalSpeechEngine.test.ts are under files:; intent updated for the approval/unmount behaviour and the cache location.
25. Guard pins fragments — replaced with whitespace-tolerant regexes anchored on sendDisabledReason = and the editor disabled={…} chain; the not.toContain("fork") check now only matches code lines, not comments.
26. deleteRecording stubbed — both suites use the real releaseRecording, and the lifecycle test asserts recordingToWav rejects afterwards (registry emptied). resetVoiceInputGlobalsForTests is not exported from the voice-input subpath, and adding it means editing client-runtime, so the suites stay as they are on that point.

Worth a check
27. React Compiler bailout — checked, not an issue: running babel-plugin-react-compiler over ChatComposer.tsx yields the identical four bailouts on the custom base (two Cannot access refs during render, two upstream try/finally blocks at the component's top level), so the main component was never compiled before this PR and the hook argument changes nothing. The fork's own hook and control compile clean.
28. useEffectEvent stored in a mount-time object — not changed. React's useEffectEvent returns a stable trampoline whose implementation is swapped at every commit, and it throws only when invoked during render; here it is invoked from controller callbacks and key handlers after commit, so it always sees the latest closure. The BranchToolbar note was about the event return value being read during render, a different shape. StrictMode's double initializer builds an idle recorder/controller pair with no acquired resources; nothing leaks.

Reviewed with Claude Fable 5.1 in Cursor.

…estion is up

While a question is active the composer repurposes promptRef as a mirror of
the answer field, so the pending-mode readDraft compared the captured prompt
against the answer text, judged the transcript stale, and dropped it for any
nonempty prompt. Read the draft store's prompt instead and commit through
setPrompt alone; the ref and composerCursor belong to the answer editor then.

A controller test documents the stale rule the wiring has to satisfy, and
the guard pins the pending-mode branch to prompt/setPrompt.

Claude Fable 5.1 via Cursor

Co-authored-by: Cursor <cursoragent@cursor.com>
@NoahHendrickson

Copy link
Copy Markdown
Owner Author

Correction to item 1 / the earlier P1 thread: the question case was not fixed by d9b21ea98, only the approval case was. While a question is active the composer's pending-input sync effect repurposes promptRef as a mirror of the answer field (promptRef.current = nextCustomAnswer), so the pending-mode readDraft compared the captured prompt against the answer text, judged the transcript stale, and dropped it for any nonempty prompt. Approvals never write promptRef, which is why that path worked and hid the problem.

Fixed in 36aab784c: the pending-mode branch reads the draft store's prompt and commits through setPrompt alone — promptRef and composerCursor belong to the answer editor while a question is up, so neither is touched. Insertion still uses the selection captured at start(); the commit-time read only feeds the stale check.

Tests: a controller test in forkLocalDictation.test.ts documents the stale rule (same text at start and commit → transcript lands; answer field read instead → discarded), and the guard pins the pending-mode branch to prompt/setPrompt and rejects promptRef.current inside the hook argument.

@NoahHendrickson
NoahHendrickson merged commit 3a68a82 into custom Sep 11, 2026
20 checks passed
@NoahHendrickson
NoahHendrickson deleted the t3code/research-native-speech-to-text branch September 11, 2026 23:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants