Skip to content

Add Apple Speech as a transcription engine choice - #1776

Merged
claude[bot] merged 26 commits into
mainfrom
claude/apple-speech-engine-ykvmp9
Sep 24, 2026
Merged

claude[bot] merged 26 commits into
mainfrom
claude/apple-speech-engine-ykvmp9

Conversation

@claude

@claude claude Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Requested by Justin · project thread

Before: the Model picker offered Parakeet V3, Parakeet V2, and two Whisper models. People who find Apple's own speech engine more accurate for their language had no way to use it.

After: the Model picker has a fifth choice, "Apple Speech (built into macOS)". It runs on the Mac using macOS 26's SpeechAnalyzer/SpeechTranscriber. It works for meetings and imported files (with the Meeting language picker) and for dictation (in the Mac's language). The first time a language is used, macOS downloads its files from Apple and the Model files card shows progress. Parakeet V3 stays the default.

It has never run on a Mac. It merges into main once CI is green and its deep review has nothing open. Merging doesn't ship it: it goes out in a later release after the Mac test below. It's not part of 1.1.62.

Why

User feedback (Chris, emailed 2026-09-23): Apple's on-device engine is the most accurate for his native language, and he asked for SpeechAnalyzer/SpeechTranscriber by name. His main use is importing hour-long recorded lectures and meetings.

Product Impact

  • Affects: meetings, dictation
  • Lane: meeting reliability
  • Why this matters: gives people a third engine without changing anything for users who keep the default.

What changed

  • Sources/Speech/AppleSpeechEngine.swift (new): wraps SpeechTranscriber + SpeechAnalyzer. Each diarized segment is converted to the analyzer's preferred format, analyzed, and its final results joined. Language assets install through AssetInventory.assetInstallationRequest, with progress published as .downloading. Finished text runs through the custom dictionary, like the other engines.
  • Sources/Speech/AppleSpeechLocalePolicy.swift (new, Foundation-only, fast-tested): maps the app's bare language codes ("es") onto Apple's per-region locales. The Mac's region wins, then the region its script implies (Traditional Chinese → Taiwan), then the language's home region, then a stable first match.
  • Language rules: Apple's engine can't detect a language, so Auto means the Mac's first preferred language. An explicit meeting language is honored. When a language isn't supported, the error says so and names what to change ("Apple Speech can't transcribe Finnish yet. Choose another meeting language or another model in Settings."). The Settings row says this before recording too, and the picker only lists languages macOS reports as supported.
  • TranscriptionModelChoice.appleSpeech (raw value apple-speech) with its own warmup runtime, so it never shares a lease with Parakeet or Whisper. supportsMeetingLanguageChoice replaces isWhisper for the language setting.
  • STTRouter routes every engine switch to the new engine. Dictation uses the existing Whisper-style external-engine path over Parakeet-recorded samples.
  • Model files card: Apple-specific download copy ("Downloading from Apple") and the real failure message instead of the generic retry text.
  • -framework Speech added to the shared swiftc args.
  • Setup downloads the Mac's language first, since that's what dictation uses. A different saved meeting language downloads right after, with progress in the Meeting language row. Once the engine is ready, later language downloads don't flip it back to "not loaded", so dictation never waits on a meeting language.
  • SpeechAnalyzer is created with .lingering model retention so Apple keeps the model loaded across a meeting's segments.
  • TranscriptedQA accepts transcription_engine: apple_speech_local, and docs/capture-format.md lists it.
  • Changing the Meeting language starts that language's download right away, and the Meeting language row shows "Downloading Spanish from Apple… 40%" (or a plain failure line). This progress is separate from the model state, so the engine never looks unloaded because of it.
  • Apple caps how many languages an app can reserve. At the cap, the engine releases the least recently used one that isn't in use. A transcription failure makes the next attempt re-check the language with Apple instead of trusting a stale "installed" mark.
  • NSSpeechRecognitionUsageDescription added to Info.plist as insurance in case macOS checks it on this path.
  • A saved meeting language Apple can't transcribe now gets the "Choose a Whisper model" failure copy instead of generic "needs another pass" copy.
  • Per-segment audio conversion runs off the main thread. Chinese, Japanese, Cantonese and Thai results join without spaces. Traditional Chinese Macs get zh_TW, and Mandarin never lands on a Hong Kong (Cantonese) locale.
  • No new Sentry/analytics events: the new apple_speech.* EventReporter events aren't on either allowlist, so they stay in local logs only.

How I checked it

  • scripts/dev/agent-preflight.sh
  • python3 scripts/dev/check-build-source-lists.py (passed)
  • bash -n on the edited scripts (passed)
  • bash build.sh --no-open and bash run-tests.sh: green in CI (app-build, checks, spm-tests on 00fcf03). The Speech API calls compiled first try.
  • Everything after 00fcf03 (two in-thread reviews, the release-wide deep review, and a final review, all findings fixed), merged with main after 1.1.62: CI running.
  • Apple API names checked against Apple's published docs (SpeechAnalyzer.Options(priority:modelRetention:), AssetInventory reserve/release/maximumReservedLocales, SpeechTranscriber.isAvailable).
  • Added Tests/AppleSpeechLocalePolicyTests.swift and extended the model/language/router policy tests.
  • Mac test: not done yet (before the release that carries it, steps below).

Mac test (before the release that carries it)

Do one step at a time. Stop and report at the first thing that looks wrong.

  1. Open Transcripted from Applications (this PR's build). Go to Settings and set Model to Apple Speech (built into macOS).
    • Expect: no crash. If an "allow speech recognition" popup appears, click Allow and note it. The Model files card says it's getting ready, or shows a download from Apple.
  2. Wait for the Model files card to say ready.
    • Expect: it finishes, or shows a clear message. It shouldn't spin forever.
  3. Set Meeting language to Spanish.
    • Expect: Spanish is in the list. Right away, the note under it says "Downloading Spanish from Apple… N%" and then goes back to the normal note when it's done.
  4. Import a short Spanish audio or video file (1–5 minutes).
    • Expect: a Spanish transcript that saves and opens.
  5. Set Meeting language back to Auto. Record a short meeting (30 seconds, just talk).
    • Expect: a transcript in your Mac's language.
  6. Dictate one sentence into any text field.
    • Expect: the text pastes. It may be a bit slower than Parakeet on the first try.
  7. Switch Model back to Parakeet V3 and dictate once.
    • Expect: it works exactly like before.

If anything fails, grab ~/Library/Application Support/Transcripted/logs/app.jsonl and search for apple_speech.

Risk Review

  • Privacy / local-first behavior reviewed: audio stays on the Mac. Only Apple's language assets are downloaded, by macOS.
  • Storage path or migration impact reviewed: new raw value only. Unknown values already fall back to Parakeet, so downgrading is safe.
  • Public-facing copy stays concrete and matches current product scope
  • Release/update impact reviewed: not part of 1.1.62.
  • Agent PRs link the issue/workpad and stay draft until human review
  • UI changes include sanitized .agent-review/visuals/ evidence: none, never rendered.
  • No private transcripts, audio, tokens, personal paths, or customer data are included

Notes

Known limits, left for measurement: one analyzer per diarized segment (the STT shootout in #1788 measures Apple Speech on an hour-long lecture, and if it's slow, a per-job analyzer is the next step); setup needs the Mac's language, so if that download fails, the engine fails even when the meeting language is installed. Language reservations go through one queue, so a hung Apple reservation call would stall later installs too (accepted; it's what stops two installs from releasing each other's slot).

How: one new engine behind the existing STTRouter switch. Nothing in TranscriptedCore changed. The pipeline already asks the engine to resolve the meeting language once per job, then passes that language to every segment.

Agent handoff

COORD_DONE: BRIEF | this PR | Apple Speech engine choice for meetings/imports/dictation | none | Mac test before its release (steps in PR body) | preflight + source-list check + CI green on 00fcf03 | CI green, then merge into main

🤖 Generated with Claude Code

https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2

Adds macOS 26's on-device SpeechAnalyzer/SpeechTranscriber as a fifth
model choice next to Parakeet and Whisper. It works for meetings and
imported files (with the meeting language picker) and for dictation
(Mac's language). macOS downloads each language's assets on first use;
unsupported languages fail with a message that says what to change.
The default engine stays Parakeet V3.

Never compiled: this session has no Swift toolchain.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
@claude claude Bot assigned r3dbars Sep 23, 2026
@claude
claude Bot requested a review from r3dbars September 23, 2026 17:33
The languageCode parameter shadowed the static languageCode(ofIdentifier:)
helper. Also silence the AVAudioPCMBuffer Sendable warning in the
converter input block the same way WhisperEngine imports WhisperKit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
@claude claude Bot mentioned this pull request Sep 23, 2026
10 of 13 tasks
…ngine-ykvmp9

# Conflicts:
#	Sources/Speech/STTRouter.swift
@claude

claude Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

build-and-test is red on 1a17e3a because its three Mac jobs (app-build, checks, spm-tests) were cancelled on purpose. Nothing failed. CI on this draft is on hold so the 1.1.62 release PRs get the Mac runners first. The last full run (00fcf03) was green. Once 1.1.62 ships I'll merge main in again and push without [skip ci], which re-runs everything.


Generated by Claude Code

…[skip ci]

A meeting-language download after the engine was ready published
.downloading, which made isModelLoaded false and sent dictation back
through initialize(), waiting on a language it doesn't use. Later
downloads now stay quiet once the engine is ready. Also only clear an
install-task entry if it's still the caller's own task, so cleanup()
during an install can't drop a newer one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
Setup used to install the saved meeting language, so dictation in the
Mac's language downloaded its files at stop time with no progress shown.
Setup now installs the Mac's language (what dictation uses) and fetches
a different meeting language quietly in the background. If Apple can't
transcribe the Mac's language, dictation falls back to a supported
meeting language, and otherwise says plainly that the Mac's language
isn't supported instead of pointing at the meeting language setting.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…gine id [skip ci]

TranscriptValidator rejected transcription_engine: apple_speech_local, so
validate-all and the import smokes would fail on any Apple Speech
meeting. Adds it to the allowlist, the QA test, the router's pinned
identifier suite and the capture-format engine table.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
A new SpeechAnalyzer per diarized segment with the default model
retention may reload Apple's model hundreds of times for a long meeting.
Ask for .lingering retention so it stays loaded across segments.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…n [skip ci]

A Mac set to zh-Hant with a region Apple has no Chinese locale for fell
through to the zh home region (CN) and got Simplified. The locale policy
now takes the Mac language's script into account, and the engine reads
region and script from the Mac's first preferred language before the
region setting.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…code [skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
@claude claude Bot mentioned this pull request Sep 23, 2026
8 of 9 tasks
…p9 [skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
Cheap insurance: if macOS checks for NSSpeechRecognitionUsageDescription
on the SpeechAnalyzer path, a missing string aborts the app instead of
returning an error. Harmless when never used.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…vations [skip ci]

The installed-language cache lived for the whole process, so if macOS
removed or updated a language's assets every later segment failed until
relaunch. A transcription failure now clears that mark so the next
segment or retry asks Apple again.

assetInstallationRequest reserves locales itself and throws past
AssetInventory.maximumReservedLocales. Before installing a new language
at the limit, release the least recently used reserved language that
isn't dictation's, the saved meeting language, or one mid-install.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…K joins [skip ci]

1cd9456 made the Mac language entry's region apply to every language, so
an English (US) Mac in Mexico got es_US for Spanish meetings instead of
es_MX. The entry's region and script now apply only when resolving the
Mac's own language.

Final result pieces were joined with spaces, which put spaces mid-sentence
for Chinese, Japanese, Cantonese and Thai. Those now join without one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…ess [skip ci]

Once Apple Speech was ready, a newly chosen meeting language downloaded
silently inside the next meeting's transcription, with no progress. Now
changing the Meeting language (or finishing setup) starts that download
right away, and the Meeting language row shows "Downloading Spanish from
Apple… 40%" or a plain failure line. Per-language progress lives in a new
languageDownload value so it never flips the engine to not-loaded.

Also from review: the background prefetch is tracked, generation-checked
and cancelled by cleanup(); only the install that owns the progress may
restore state, so a cancelled install can't overwrite a newer one's
progress; setup re-checks its generation before downloading; and the
Settings row no longer caches an empty locale list (it checks
SpeechTranscriber.isAvailable instead).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…opy [skip ci]

Apple's "can't transcribe Finnish" error didn't match any failure
classifier, so a meeting with a saved language Apple can't handle showed
generic "needs another pass" copy and every retry failed the same way.
The router now rethrows it with the "select a whisper model" wording
the classifiers already route to the Choose a Whisper model guidance,
and that guidance says the selected model can't use the saved language.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…loads [skip ci]

The failed card showed any Apple message verbatim, including raw
download errors; now it shows Apple Speech's own errors (which name the
fix) and plain retry copy otherwise. The loading card no longer claims
the files are already on the Mac, since Apple may still download them.
The Settings button reads Try Again instead of Retry Download for Apple
Speech. Adds fast tests for the Apple card copy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
transcribe, makePCMBuffer, convert and makeTranscriber were statics of
a @mainactor class, so the sample copy, AVAudioConverter pass and result
loop ran on the main thread for every meeting segment. They're now
nonisolated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…skip ci]

Apple's speech engines have historically used zh-HK/zh-MO for Cantonese.
The app lists Cantonese separately, so "zh" skips those regions in the
Mac-region step and a zh-Hant-HK Mac gets zh_TW. zh_HK is still used when
it's the only Chinese locale.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…p ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…ested policy [skip ci]

Which reserved language to release and the Meeting language row's
download line were only reachable through the Speech framework. Both now
live in the Foundation-only AppleSpeechLocalePolicy file, so the fast
tests cover the limit/variant/in-use/LRU rules and the caption clamping.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…n races [skip ci]

From a review of today's fixes:
- Only one install drives the model card's setup progress, and only it
  restores the state, so a second install can't freeze or strand the
  card on a stale percentage.
- Choosing a new meeting language mid-download cancels the old
  language's download unless a recording is waiting on it, so it stops
  holding bandwidth and a reservation slot.
- Freeing a reservation slot and reserving the new language happen one
  install at a time, so two installs at Apple's limit can't both count
  on the same freed slot.
- A new language choice clears an old language's failure note.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…opy [skip ci]

- The Meeting language row only shows a download note for the language
  it's about (the saved one, or the Mac's for Auto).
- An empty supported-language list is asked for again a few times
  instead of only on the next appear.
- An Auto meeting that fails on the Mac's language keeps Apple's own
  message; only a saved language gets the Choose a Whisper model copy.
- Apple's downloading card names the Try Again button it actually has.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…peech-engine-ykvmp9

Lifts the CI hold: this push runs CI on everything since 00fcf03.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
The final review found cleanup() clearing foregroundInstallWaiters could
under-count a recording that joins after re-init, letting a language
change cancel an install a recording waits on. Each waiter's defer
already balances its own count, so stop clearing it. Also correct the
comments on cancelled prefetch reservations and the Auto-path error.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
…ine-ykvmp9

Keep both engine cases. Grouped non-Whisper lists include both, Ultra
maps to the .ultra variant and Apple Speech to none, and both engine ids
stay in the QA validator. Both PRs bumped the model count from 4 to 5,
so the model-choice and QA engine counts are now 6.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013GAxhzGyuaTxNTv8tpj9j2
@claude
claude Bot marked this pull request as ready for review September 24, 2026 04:30
@claude
claude Bot merged commit b7cba6c into main Sep 24, 2026
7 checks passed
@claude
claude Bot deleted the claude/apple-speech-engine-ykvmp9 branch September 24, 2026 04:30
claude Bot pushed a commit that referenced this pull request Sep 24, 2026
Sources/CLAUDE.md: keep the one-meeting Boost wording and add Apple Speech
to the model list.

Also time the "can't hear the call" stretch by wall clock. recordingDuration
reads 0 on an unexpected stop, so a long loss there would not be marked
degraded (#1811 hit the same reset for meeting length).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SgXMcLRGBtiSiybdgyD1VJ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants