Utter is a macOS menu bar app for AI-powered voice input and dictation. It supports both fully local on-device inference and remote LLM APIs. Press a hotkey to start recording, release to transcribe, and the result is typed directly into whatever app you're using.
Three output modes are available:
- Verbatim — raw transcription, lowest latency
- Smart Format — transcription cleaned up by an LLM (contextual filler removal, grammar fixes, structured formatting)
- Voice Command — speak a command and get an AI-generated response based on screen context
Watch on the website or download the MP4s: Feature tour (0:51) · A day at the office (0:58) · Local voice input (0:28). All three have Mandarin voice-over. The UI is an animated reconstruction of the real app, and the sample texts are illustrative.
| Feature | Description |
|---|---|
| Multiple Speech Engines | Local Qwen3-ASR, Confucius4-R2T2, FireRedASR2, Mega-ASR, WhisperKit, or on-device Apple Speech; Doubao ASR as a remote option |
| Smart Text Processing | Local MLX models (Qwen3.5 / Qwen3 / Gemma) or a remote LLM infers spoken intent — contextual cleanup, "scratch that" restarts, self-correction handling, spoken punctuation, technical terms, numbers/ranges/units, and structured formatting |
| LLM-Owned Spoken Formatting | Spoken casing, no-space dictation, identifiers, file paths, shortcuts, emoji, Markdown tasks, dates/times, quantities, units, formulas, fractions, and digit sequences are handled by the Smart Format / Voice Command prompts instead of local hardcoded rewrite rules |
| Voice Edit Commands | In Voice Command mode, an LLM classifies safe structured actions for replacing, undoing, proofreading, titling, summarizing, drafting replies, making meeting notes, extracting key points/decisions/questions/risks/deadlines/owners/action items, rewriting tone, expanding, making tables/lists, or deleting the previous Utter insertion or selected text |
| Verbatim & Preview Boundary | Verbatim mode, streaming HUD, integration partials, and instant-insert drafts keep ASR text close to raw output with only dictionary, whitespace, duplicate, and non-speech-artifact cleanup |
| Remote LLM Support | OpenAI, Claude (Anthropic format), Gemini, OpenRouter, SiliconFlow, Doubao, Bailian, MiniMax (CN & Global) |
| Global Hotkey | Configurable key (Fn/Ctrl/Shift/Option) with long-press, double-tap, or single-tap activation |
| Translation Dictation | Use a dedicated hotkey chord to speak in one language and insert an English, Chinese, Japanese, Korean, Spanish, French, or German translation |
| Screen Context OCR | Captures on-screen text via ScreenCaptureKit + Vision to help the LLM correct homophones |
| Voice Command Mode | Screen-aware voice assistant — summarize, reply, translate based on what's on screen |
| Input Memory | Recent input history injected as LLM context for better continuity |
| Industry Vocabulary | Choose Medical, Legal, Finance & Accounting, or Software Technology terms for Apple Speech / Whisper biasing and output normalization; personal terms take priority |
| Edit Rules | Personal text replacement rules applied on every output |
| Language Style Presets | Casual / Professional / Custom prompt |
| Input History & Stats | Full history with raw vs. processed comparison, word count stats, configurable retention |
| Bilingual UI | Chinese and English interface, independent of recognition language |
| Sound Feedback | Audio cues on recording start and stop |
| Guided Onboarding | Step-by-step setup: permissions, model download, and first use |
- OS: macOS 26 (Tahoe) or later
- Chip: Apple Silicon (M1 / M2 / M3 / M4)
- Disk: ~0.7 GB minimum (Apple Speech + Qwen3.5 0.8B); ~1.8 GB for the default setup (Apple Speech + Qwen3.5 2B); larger speech and text models need several GB more
Grab the latest .dmg from Releases, open it, and drag Utter.app to Applications.
"Cannot verify the developer" on first launch? The app is not notarized by Apple. Before first run, execute in Terminal:
xattr -cr /Applications/Utter.appOr go to System Settings → Privacy & Security and click "Open Anyway".
# Build .app bundle + .dmg installer
bash scripts/build-app.sh
# Or for development
swift build
swift run OpenType
# Build, sign, and launch the development .app bundle
bash scripts/build-and-run.sh --verify
# Or open in Xcode
open Package.swiftMaterial changes follow the artifact-driven workflow in
docs/sdlc/README.md. Before opening a pull request, run:
bash scripts/sdlc-checks.sh
bash scripts/ci-basic-checks.sh
swift testThe public app is Utter.app. The Swift package product remains OpenType so existing source integrations and upgrade paths continue to work.
- Launch Utter — it appears as a waveform icon in the menu bar
- The onboarding wizard guides you through permissions and model setup
- Grant Microphone and Accessibility permissions (required)
- Wait for the default text model (Qwen3.5 2B, ~1.7 GB) to download — one time only
- Hold Fn to start dictating, release to stop and insert text
| Permission | Purpose | Required |
|---|---|---|
| Microphone | Audio capture | Yes |
| Accessibility | Global hotkey, text insertion (clipboard + simulated ⌘V), and reading the focused field for context | Yes |
| Speech Recognition | Apple on-device ASR engine | Only if using Apple Speech |
| Screen Recording | OCR for screen context and Voice Command mode | Optional |
| Network | Model downloads; remote LLM API calls | First run / remote LLM mode |
Everything below runs on your Mac. Download a model once and dictation keeps working with the network off.
| Stage | Models | Notes |
|---|---|---|
| Speech recognition | Qwen3-ASR 1.7B (recommended), Confucius4-R2T2, FireRedASR2-AED, Mega-ASR, WhisperKit (large-v3-turbo … tiny), Apple Speech (on-device) | Qwen3-ASR and Confucius run through native Swift + MLX; FireRed and Mega-ASR through MLX; Apple Speech is forced on-device |
| Text cleanup (LLM) | Qwen3.5 0.8B / 2B (default) / 9B / 35B-A3B, Qwen3 0.6B–30B-A3B, Qwen2.5, Gemma 4 E2B / E4B, Gemma 3 1B / 4B / 12B | MLX on Apple Silicon; an experimental ANE-LM runtime can run a local Qwen3 model; you can also add your own MLX model |
Remote options are opt-in: Doubao (Volcengine) ASR and the remote LLM providers below send audio or text to that provider.
Utter supports both OpenAI-compatible and Anthropic API formats:
| Provider | API Format | Base URL |
|---|---|---|
| OpenAI | OpenAI | https://api.openai.com/v1 |
| Anthropic Claude | Anthropic | https://api.anthropic.com/v1 |
| Google Gemini | OpenAI | https://generativelanguage.googleapis.com/v1beta/openai |
| OpenRouter | OpenAI | https://openrouter.ai/api/v1 |
| SiliconFlow | OpenAI | https://api.siliconflow.cn/v1 |
| Volcengine Doubao | OpenAI | https://ark.cn-beijing.volces.com/api/v3 |
| Alibaba Bailian | OpenAI | https://dashscope.aliyuncs.com/compatible-mode/v1 |
| MiniMax (China) | OpenAI | https://api.minimax.chat/v1 |
| MiniMax (Global) | OpenAI | https://api.minimaxi.chat/v1 |
Sources/
├── App/ # Entry point, AppDelegate, AppState, VoicePipeline, AppIcon
├── Audio/ # Microphone capture (AVAudioEngine), sound playback
├── Config/ # AppSettings, ModelCatalog, RemoteModelConfig, Localization
├── Hotkey/ # Global hotkey via CGEvent tap
├── LLM/ # LLMEngine (MLX), RemoteLLMClient (OpenAI/Anthropic)
├── Output/ # Text insertion (clipboard + simulated ⌘V; Accessibility for selection and context)
├── Processing/ # TextProcessor, InputHistory, MemoryStore, personal and industry vocabulary
├── Prompts/ # PromptBuilder, prompt catalogs, style prompt presets
├── Screen/ # Screen OCR (ScreenCaptureKit + Vision)
├── Speech/ # SpeechEngine protocol, WhisperKit, Apple Speech, Doubao ASR, local ASR engines
├── UI/ # SwiftUI: MenuBar, Settings, Onboarding, Overlay, History, Models
└── Resources/ # Localization strings (en/zh-Hans), sounds, app icon
docs/ # Website (GitHub Pages) incl. demo videos, SDLC artifacts, research
marketing/ # Promo video sources, render engine, voice-over scripts
scripts/
├── build-and-run.sh # Build, sign, and launch a development .app bundle
├── build-app.sh # Build release .app bundle and .dmg installer
├── ci-basic-checks.sh # CI guardrails for linked files and resources
├── create-signing-cert.sh # Generate self-signed code signing certificate
├── evaluate-voice-quality.py # Score ASR / formatting output (CER, terms, numbers, latency)
├── generate-icon.swift # Generate AppIcon.icns from source PNG
├── render-promo.sh # Render a promo video from marketing/<promo>/
├── sdlc-checks.sh # Validate SDLC artifact stages and approvals
├── test-industry-lexicons.sh # Validate vocabulary recall and non-target preservation
├── unit-test-coverage.sh # Run unit tests with coverage thresholds
└── validate-volc-asr.swift # Validate Volcengine ASR configuration manually
- WhisperKit — offline Whisper speech recognition
- mlx-swift-lm — local LLM inference on Apple Silicon (Qwen3.5 / Qwen3 / Gemma)
- SwiftUI + AppKit — native macOS UI
- ScreenCaptureKit + Vision — screen OCR
- AVAudioEngine — low-latency microphone capture
- Apple Speech Framework — on-device speech recognition
- Qwen3-ASR / FireRedASR2 / Mega-ASR on MLX — local speech recognition
Made with care for Apple Silicon


