Repository navigation
Add Apple Speech as a transcription engine choice - #1776
Merged
Merged
Conversation
Adds macOS 26's on-device SpeechAnalyzer/SpeechTranscriber as a fifth model choice next to Parakeet and Whisper. It works for meetings and imported files (with the meeting language picker) and for dictation (Mac's language). macOS downloads each language's assets on first use; unsupported languages fail with a message that says what to change. The default engine stays Parakeet V3. Never compiled: this session has no Swift toolchain. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
The languageCode parameter shadowed the static languageCode(ofIdentifier:) helper. Also silence the AVAudioPCMBuffer Sendable warning in the converter input block the same way WhisperEngine imports WhisperKit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
10 of 13 tasks
…ngine-ykvmp9 # Conflicts: # Sources/Speech/STTRouter.swift
Contributor
Author
|
Generated by Claude Code |
…[skip ci] A meeting-language download after the engine was ready published .downloading, which made isModelLoaded false and sent dictation back through initialize(), waiting on a language it doesn't use. Later downloads now stay quiet once the engine is ready. Also only clear an install-task entry if it's still the caller's own task, so cleanup() during an install can't drop a newer one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
Setup used to install the saved meeting language, so dictation in the Mac's language downloaded its files at stop time with no progress shown. Setup now installs the Mac's language (what dictation uses) and fetches a different meeting language quietly in the background. If Apple can't transcribe the Mac's language, dictation falls back to a supported meeting language, and otherwise says plainly that the Mac's language isn't supported instead of pointing at the meeting language setting. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…gine id [skip ci] TranscriptValidator rejected transcription_engine: apple_speech_local, so validate-all and the import smokes would fail on any Apple Speech meeting. Adds it to the allowlist, the QA test, the router's pinned identifier suite and the capture-format engine table. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
A new SpeechAnalyzer per diarized segment with the default model retention may reload Apple's model hundreds of times for a long meeting. Ask for .lingering retention so it stays loaded across segments. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…n [skip ci] A Mac set to zh-Hant with a region Apple has no Chinese locale for fell through to the zh home region (CN) and got Simplified. The locale policy now takes the Mac language's script into account, and the engine reads region and script from the Mac's first preferred language before the region setting. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…code [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
8 of 9 tasks
…p9 [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
Cheap insurance: if macOS checks for NSSpeechRecognitionUsageDescription on the SpeechAnalyzer path, a missing string aborts the app instead of returning an error. Harmless when never used. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…vations [skip ci] The installed-language cache lived for the whole process, so if macOS removed or updated a language's assets every later segment failed until relaunch. A transcription failure now clears that mark so the next segment or retry asks Apple again. assetInstallationRequest reserves locales itself and throws past AssetInventory.maximumReservedLocales. Before installing a new language at the limit, release the least recently used reserved language that isn't dictation's, the saved meeting language, or one mid-install. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…K joins [skip ci] 1cd9456 made the Mac language entry's region apply to every language, so an English (US) Mac in Mexico got es_US for Spanish meetings instead of es_MX. The entry's region and script now apply only when resolving the Mac's own language. Final result pieces were joined with spaces, which put spaces mid-sentence for Chinese, Japanese, Cantonese and Thai. Those now join without one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…ess [skip ci] Once Apple Speech was ready, a newly chosen meeting language downloaded silently inside the next meeting's transcription, with no progress. Now changing the Meeting language (or finishing setup) starts that download right away, and the Meeting language row shows "Downloading Spanish from Apple… 40%" or a plain failure line. Per-language progress lives in a new languageDownload value so it never flips the engine to not-loaded. Also from review: the background prefetch is tracked, generation-checked and cancelled by cleanup(); only the install that owns the progress may restore state, so a cancelled install can't overwrite a newer one's progress; setup re-checks its generation before downloading; and the Settings row no longer caches an empty locale list (it checks SpeechTranscriber.isAvailable instead). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…opy [skip ci] Apple's "can't transcribe Finnish" error didn't match any failure classifier, so a meeting with a saved language Apple can't handle showed generic "needs another pass" copy and every retry failed the same way. The router now rethrows it with the "select a whisper model" wording the classifiers already route to the Choose a Whisper model guidance, and that guidance says the selected model can't use the saved language. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…loads [skip ci] The failed card showed any Apple message verbatim, including raw download errors; now it shows Apple Speech's own errors (which name the fix) and plain retry copy otherwise. The loading card no longer claims the files are already on the Mac, since Apple may still download them. The Settings button reads Try Again instead of Retry Download for Apple Speech. Adds fast tests for the Apple card copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
transcribe, makePCMBuffer, convert and makeTranscriber were statics of a @mainactor class, so the sample copy, AVAudioConverter pass and result loop ran on the main thread for every meeting segment. They're now nonisolated. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…skip ci] Apple's speech engines have historically used zh-HK/zh-MO for Cantonese. The app lists Cantonese separately, so "zh" skips those regions in the Mac-region step and a zh-Hant-HK Mac gets zh_TW. zh_HK is still used when it's the only Chinese locale. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…p ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…ested policy [skip ci] Which reserved language to release and the Meeting language row's download line were only reachable through the Speech framework. Both now live in the Foundation-only AppleSpeechLocalePolicy file, so the fast tests cover the limit/variant/in-use/LRU rules and the caption clamping. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…n races [skip ci] From a review of today's fixes: - Only one install drives the model card's setup progress, and only it restores the state, so a second install can't freeze or strand the card on a stale percentage. - Choosing a new meeting language mid-download cancels the old language's download unless a recording is waiting on it, so it stops holding bandwidth and a reservation slot. - Freeing a reservation slot and reserving the new language happen one install at a time, so two installs at Apple's limit can't both count on the same freed slot. - A new language choice clears an old language's failure note. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…opy [skip ci] - The Meeting language row only shows a download note for the language it's about (the saved one, or the Mac's for Auto). - An empty supported-language list is asked for again a few times instead of only on the next appear. - An Auto meeting that fails on the Mac's language keeps Apple's own message; only a saved language gets the Choose a Whisper model copy. - Apple's downloading card names the Try Again button it actually has. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…peech-engine-ykvmp9 Lifts the CI hold: this push runs CI on everything since 00fcf03. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
The final review found cleanup() clearing foregroundInstallWaiters could under-count a recording that joins after re-init, letting a language change cancel an install a recording waits on. Each waiter's defer already balances its own count, so stop clearing it. Also correct the comments on cancelled prefetch reservations and the Auto-path error. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…ine-ykvmp9 Keep both engine cases. Grouped non-Whisper lists include both, Ultra maps to the .ultra variant and Apple Speech to none, and both engine ids stay in the QA validator. Both PRs bumped the model count from 4 to 5, so the model-choice and QA engine counts are now 6. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
claude Bot
pushed a commit
that referenced
this pull request
Sep 24, 2026
Sources/CLAUDE.md: keep the one-meeting Boost wording and add Apple Speech to the model list. Also time the "can't hear the call" stretch by wall clock. recordingDuration reads 0 on an unexpected stop, so a long loss there would not be marked degraded (#1811 hit the same reset for meeting length). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SgXMcLRGBtiSiybdgyD1VJ
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Justin · project thread
Before: the Model picker offered Parakeet V3, Parakeet V2, and two Whisper models. People who find Apple's own speech engine more accurate for their language had no way to use it.
After: the Model picker has a fifth choice, "Apple Speech (built into macOS)". It runs on the Mac using macOS 26's SpeechAnalyzer/SpeechTranscriber. It works for meetings and imported files (with the Meeting language picker) and for dictation (in the Mac's language). The first time a language is used, macOS downloads its files from Apple and the Model files card shows progress. Parakeet V3 stays the default.
It has never run on a Mac. It merges into main once CI is green and its deep review has nothing open. Merging doesn't ship it: it goes out in a later release after the Mac test below. It's not part of 1.1.62.
Why
User feedback (Chris, emailed 2026-09-23): Apple's on-device engine is the most accurate for his native language, and he asked for SpeechAnalyzer/SpeechTranscriber by name. His main use is importing hour-long recorded lectures and meetings.
Product Impact
meetings,dictationmeeting reliabilityWhat changed
Sources/Speech/AppleSpeechEngine.swift(new): wrapsSpeechTranscriber+SpeechAnalyzer. Each diarized segment is converted to the analyzer's preferred format, analyzed, and its final results joined. Language assets install throughAssetInventory.assetInstallationRequest, with progress published as.downloading. Finished text runs through the custom dictionary, like the other engines.Sources/Speech/AppleSpeechLocalePolicy.swift(new, Foundation-only, fast-tested): maps the app's bare language codes ("es") onto Apple's per-region locales. The Mac's region wins, then the region its script implies (Traditional Chinese → Taiwan), then the language's home region, then a stable first match.TranscriptionModelChoice.appleSpeech(raw valueapple-speech) with its own warmup runtime, so it never shares a lease with Parakeet or Whisper.supportsMeetingLanguageChoicereplacesisWhisperfor the language setting.STTRouterroutes every engine switch to the new engine. Dictation uses the existing Whisper-style external-engine path over Parakeet-recorded samples.-framework Speechadded to the shared swiftc args.SpeechAnalyzeris created with.lingeringmodel retention so Apple keeps the model loaded across a meeting's segments.transcription_engine: apple_speech_local, anddocs/capture-format.mdlists it.NSSpeechRecognitionUsageDescriptionadded to Info.plist as insurance in case macOS checks it on this path.apple_speech.*EventReporter events aren't on either allowlist, so they stay in local logs only.How I checked it
scripts/dev/agent-preflight.shpython3 scripts/dev/check-build-source-lists.py(passed)bash -non the edited scripts (passed)bash build.sh --no-openandbash run-tests.sh: green in CI (app-build, checks, spm-tests on 00fcf03). The Speech API calls compiled first try.Tests/AppleSpeechLocalePolicyTests.swiftand extended the model/language/router policy tests.Mac test (before the release that carries it)
Do one step at a time. Stop and report at the first thing that looks wrong.
If anything fails, grab
~/Library/Application Support/Transcripted/logs/app.jsonland search forapple_speech.Risk Review
.agent-review/visuals/evidence: none, never rendered.Notes
Known limits, left for measurement: one analyzer per diarized segment (the STT shootout in #1788 measures Apple Speech on an hour-long lecture, and if it's slow, a per-job analyzer is the next step); setup needs the Mac's language, so if that download fails, the engine fails even when the meeting language is installed. Language reservations go through one queue, so a hung Apple reservation call would stall later installs too (accepted; it's what stops two installs from releasing each other's slot).
How: one new engine behind the existing
STTRouterswitch. Nothing inTranscriptedCorechanged. The pipeline already asks the engine to resolve the meeting language once per job, then passes that language to every segment.Agent handoff
COORD_DONE: BRIEF | this PR | Apple Speech engine choice for meetings/imports/dictation | none | Mac test before its release (steps in PR body) | preflight + source-list check + CI green on 00fcf03 | CI green, then merge into main🤖 Generated with Claude Code
https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2