Repository navigation
Speech-to-text model shootout harness - #1788
Merged
Merged
Conversation
Benchmark harness for comparing on-device speech-to-text models against the Parakeet V3 model the app ships: downloads a human-captioned test video, runs each model in its own process, and reports speed, latency, peak memory and word error rate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Runs Whisper large-v3-turbo through the same WhisperKit revision and decode options the app's Whisper model choice uses. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
… STT shootout [skip ci] Adds Whisper turbo and Distil-Whisper (mlx-whisper), Parakeet V3/V2 (parakeet-mlx), Canary 1B v2 and 180M Flash (onnx-asr + Silero VAD), Moonshine base and medium (moonshine-voice), whisper.cpp turbo (pywhispercpp), Granite Speech 4.0 and Nemotron streaming (mlx-audio). APIs checked against each package's source. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
From a review against each package's source: - parakeet-mlx: hand its loader float32 (bf16 broke the mel spectrogram) - a failed video download moves on to the next video; prefer AAC audio afconvert can read; give yt-dlp a JavaScript runtime (deno extra) - drop uv venv --clear (older uv lacks it); cap setup steps at 45 min - keep results a model wrote before crashing in teardown - read transcripted-cli results from per-file JSON, not stdout - pin WhisperKit's swift-transformers/swift-jinja to its own resolved set - strip caption speaker labels mid-line; guard report math Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
From the deep review of #1788: - run the Parakeet V3/Ultra rows on an APFS clone of the model folder, since FluidAudio deletes and re-downloads a folder it fails to load; fail the Ultra row if its marker is gone after the run - key the cached video by its URL list so --url can't reuse another video - keep every download (HF_HOME, uv cache and Python, whisper.cpp, Moonshine) under ~/stt-shootout; stop below 20 GB free; print cleanup - pin pip packages; record packages, HF snapshot hashes, normalizer, and run conditions (app running, battery) in report.json - label peak memory honestly for Core ML and Apple Speech rows - reuse results only when clip settings match; scrub home paths - add a test-matrix rule for scripts/stt-shootout Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
hillclimb_bench.py speaks the lab's transcripted.hillclimb v1 protocol: knob stt.engine picks the model, each suite item is one recording, and it reports time, speed, latency, load, memory and WER per item. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
claude Bot
pushed a commit
that referenced
this pull request
Sep 23, 2026
Adds the stt-shootout bench (script lives in PR #1788), a stt.engine knob with the shootout's engine list, a Mac-local stt-clips suite, and a speech-model-accuracy objective: fewer wrong words, at most ~35% slower, no engine failures or empty text. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS
12 of 15 tasks
… [skip ci] Every Transcripted meeting's audio is named microphone.m4a or system_audio.m4a, so naming the converted WAV by stem alone let a later hill-climb item be scored on the first item's audio. The converted copy is now keyed by source path, size and mtime. report.json also records which WhisperKit model files were measured (Hub commit when available, plus a size fingerprint). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
7 of 9 tasks
claude Bot
pushed a commit
that referenced
this pull request
Sep 24, 2026
… lab typecheck [skip ci] Second deep-review round on #1791: - N1: only the outermost run_group starts a new session; it marks the child env so an adapter's own run_group keeps the app/CLI/harness in the same group. The climber's timeout now kills the real work too. Two-level tests (the outer one fails on the old code). - S2 follow-up: lab_control launch turns analytics and crash reporting off for the launched process only, as NSArgumentDomain launch arguments, so the person's saved Settings are never needed or changed. - N2: scripts/dev/typecheck-lab-build.sh type-checks the app with -D TRANSCRIPTED_LAB_CONTROL; app-build CI runs it after the normal build. - N3: a holdout check writes a started row before it runs, so a crash still uses up budget. - N4/N5: guide notes on #1788's audio naming and LabKnobOverrides' env var. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
- onnx-asr's Canary decoder keeps words generated after end-of-text when segments are batched, which made Canary 180M write 15,489 words for a 7,191-word answer key. Decode one VAD segment at a time. - onnxruntime 1.30 rejects a .onnx.data file symlinked out of the model folder (the Hugging Face cache layout), so Canary 1B v2 never loaded. onnx-asr models now download into plain folders under models/onnx-asr, and their Hub commits still land in the report. - Granite Speech needs jinja2 for its chat template. - Rows whose word count is far off the answer key (or WER over 50%) are marked broken and left out of the pick. - Record whether Transcripted was running or the Mac was on battery right before and after each model, not only at report time. - WhisperKit's model record looked for the bare variant name, but the folder is openai_whisper-<variant>. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
`--rerun canary-1b-v2,granite-speech` redoes only those models and reuses every other saved result, so re-running a few fixed rows still produces a complete report. Plain `--rerun` still redoes everything. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Brings in Parakeet Ultra (#1783) so the shootout worktree on the Mac has its install script. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
7 tasks done
Picks up the VM test socket-path fix (#1883) and everything merged today. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
9 of 19 tasks
r3dbars
marked this pull request as ready for review
September 28, 2026 02:53
r3dbars
added a commit
that referenced
this pull request
Sep 28, 2026
* Add hill-climb lab core: knob registry, held-out splits, paired stats, climber [skip ci] Python stdlib tool under scripts/hillclimb that tunes app knobs against scored benches. Dev/holdout splits are hash-stable, verdicts need a bootstrap-CI win above min_effect with no guardrail or hard-gate regression, timing benches interleave A/B runs, and holdout checks are budgeted. A synthetic demo registry backs the self-tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * Let the lab override 8 meeting-pipeline knobs from a file [skip ci] LabKnobOverrides reads TRANSCRIPTED_LAB_KNOBS_FILE once per process. With no env var nothing changes: every call returns today's default and there is no file I/O. Wired: diarizer clustering threshold, VBx Fa/Fb, min segment duration, and same-voice consolidation / small-cluster absorb per embedder. Not compiled yet (no Swift toolchain in session). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * Add dictation stop-latency bench adapter and 64-phrase suite [skip ci] Synthesizes fixtures with say, runs the real DictationStopBenchmarkRunner once per repetition in an isolated HOME, and reports per-phrase latency, word error rate (from the saved Markdown), and missing-text, silence-text and unstable-output gates. Self-test now also runs bench adapter tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * Add real knob registry, objectives, speaker naming bench, and lab guide [skip ci] 101 knobs (10 live, 21 bench-only, 70 mapped but hardcoded) with verified defaults and source lines; three objectives (dictation stop latency, meeting turnaround, speaker naming across calls); a SpeakerEvalHarness autoeval adapter with per-cache items and every safety counter as a hard gate; docs/hill-climb-lab.md; a test-matrix rule and a repo-hygiene step that runs the lab self-test and registry validation. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * Add meeting turnaround bench adapter and corpus suite builder [skip ci] Runs transcripted-cli import-audio per corpus item with a fresh empty speaker database, times Stop-to-transcript per second of audio, scores word recall and speaker count against truth, and passes lab knob overrides through TRANSCRIPTED_LAB_KNOBS_FILE (erroring if the CLI did not confirm them). One untimed warmup import absorbs model load. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * Add a lab control channel so an agent can drive the real app [skip ci] Off unless the app is launched with TRANSCRIPTED_LAB_CONTROL_DIR. Commands dropped as JSON into <dir>/inbox (ping, status, start/stop dictation, start/stop meeting, import audio) call the same entry points the menus use; responses go to <dir>/responses.jsonl. Nothing from the channel goes off-device. scripts/hillclimb/lab_control.py launches the app and sends commands. Swift not compiled yet (no toolchain in session). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * Plug the speech model shootout into the lab [skip ci] Adds the stt-shootout bench (script lives in PR #1788), a stt.engine knob with the shootout's engine list, a Mac-local stt-clips suite, and a speech-model-accuracy objective: fewer wrong words, at most ~35% slower, no engine failures or empty text. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * hillclimb: kill the whole process group when a bench adapter times out [skip ci] subprocess.run(timeout=) only kills the direct child, so the app, the CLI or the speaker harness kept running into the next trial and skewed timings. Adds hc_proc.run_group (new session + killpg) and routes every adapter subprocess through it, with grandchild-survival tests per adapter. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * hillclimb: exact sign-flip tests, minimum suite sizes, fail-closed guardrails [skip ci] Fixes the deep review's stats and holdout findings on #1791: - S4: verdicts use a one-sided paired sign-flip permutation test (exact up to 16 units) instead of the percentile bootstrap; the holdout check uses p < 0.01. Climb and confirm refuse suites under 10 dev / 8 holdout independent units. simulate-null reports false-accept rates at real sizes. - S5: a guardrail measured on fewer than half the primary's units, or that lost items, rejects. Metrics can declare the item field they need (truth, speakers, text) and the lab refuses suites that can't feed them. - S6: holdout peeks count by holdout item overlap (>50% = same holdout), so adding an item no longer resets the budget. Holdout per-item values are sealed out of trials.jsonl and bench work dirs go to holdout-sealed/. - S7: suite items can carry a cluster; clusters count once and never straddle the holdout line. Speaker items cluster by family and identity split, which blocks that objective until the harness emits per-person rows. Its recommendation now lists the per-bucket contract as required. - S8: CommandBench runs through hc_proc.run_group. - M1: the inconclusive re-measure pools with the first run at alpha/2. - M2: every repetition's build/host/OS must match. - M3: climb-result.json is checkpointed after each decision; climb --resume. Malformed bench results become item errors instead of crashing. - M4: forced holdout checks are recorded as forced. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * lab control: compile the channel only into local lab builds, lock its folders [skip ci] Fixes the deep review's B1, S2, S3, M6, M8 and M9 on #1791: - B1: LabControlChannel and its launch hook sit behind #if TRANSCRIPTED_LAB_CONTROL, set only by build.sh --lab. build-beta.sh refuses TRANSCRIPTED_LAB_BUILD and fails if the binary contains the channel's env var name. In lab builds the control dir, inbox/ and done/ must be real 0700 dirs owned by this uid, re-checked every poll. - M6: command files open O_NOFOLLOW|O_NONBLOCK and are fstat-checked (regular, ours, <= 64 KB) before reading; done/ moves use rename(2). - S3: stop_dictation pastes only when paste is explicitly true; start_dictation calls the session directly and never activates an app. - S2: lab_control.py launch needs --container (or --use-real-library), refuses a relocated capture library and telemetry-on unless overridden. - M9: LabKnobOverrides drops unknown ids with one stderr line. - M8: Support and Core CLAUDE.md list the lab files and Core's env var. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * hillclimb: #1789 merge note hooks, source revision in recommendations, models-dir note [skip ci] Review S1/M5/M7 on #1791: the clustering knob's notes say it's in cosine units and point at the merge note for #1789's FluidAudio 0.17 distance change; recommendation.json records the revision its source line numbers came from; the guide explains the shared FluidAudio model cache. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * hillclimb: nested timeouts kill the real work, telemetry off per run, lab typecheck [skip ci] Second deep-review round on #1791: - N1: only the outermost run_group starts a new session; it marks the child env so an adapter's own run_group keeps the app/CLI/harness in the same group. The climber's timeout now kills the real work too. Two-level tests (the outer one fails on the old code). - S2 follow-up: lab_control launch turns analytics and crash reporting off for the launched process only, as NSArgumentDomain launch arguments, so the person's saved Settings are never needed or changed. - N2: scripts/dev/typecheck-lab-build.sh type-checks the app with -D TRANSCRIPTED_LAB_CONTROL; app-build CI runs it after the normal build. - N3: a holdout check writes a started row before it runs, so a crash still uses up budget. - N4/N5: guide notes on #1788's audio naming and LabKnobOverrides' env var. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * hillclimb: add #1789's speaker-lab bench, knobs and objective [skip ci] Pasted from #1789's speaker_lab.README.md: the speaker-lab command bench, 6 new knobs (diarization backend, Nemotron preset, match mode and floor, replay dedup, write-path fixes), speaker-lab-recognition, and an identical copy of its suite. The adapter lives on #1789, so trials are item errors until it merges. The objective reports BLOCKED: 4 holdout series, need 8. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * hillclimb: take #1789's 24-series speaker-lab suite (8 holdout) [skip ci] Byte-for-byte copy of config/hillclimb/suites/speaker-lab-ami.json from #1789 at ffe5871 (salt speaker-lab-ami-v2, 16 dev / 8 holdout series), so speaker-lab-recognition is no longer BLOCKED. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * hillclimb: speaker-lab objective is no longer blocked; update notes [skip ci] Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS * Hill-climb test: compare resolved paths for the bench HOME The adapter resolves the request path, so on macOS a temp dir under /var comes back as /private/var and the startswith check failed. The suite had only run on Linux. Now 164/164 on macOS too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: r3dbars <r3dbars@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Justin · project thread
Before: there was no way to see how other on-device speech-to-text models compare with the Parakeet V3 model the app ships, on speed or accuracy.
After: one command on an Apple Silicon Mac (
bash scripts/stt-shootout/run.sh) downloads an hour-long, human-captioned YouTube lecture and runs 15 models on it. The report table gives time for the whole hour, speed vs real time, speed vs Parakeet V3, 10-second clip latency (warm and first use), load time, peak memory, and word error rate against the captions.Why
Justin wants to know which models are faster than Parakeet V3, and how accurate they are, before adding any to the app.
Product Impact
docs only(bench tooling underscripts/stt-shootout/; no app code changes)agent workflowWhat changed
scripts/stt-shootout/shootout.py: the runner.afconvert.proc_pid_rusage.report.md/json/csvand the transcripts.transcripted-clion an APFS clone of the model folder. FluidAudio deletes and re-downloads a folder it fails to load, and the clone keeps that from ever touching the app bundle or the Ultra install.~/stt-shootout: HF_HOME, uv's cache and Python, WhisperKit, whisper.cpp and Moonshine. The run stops below 20 GB free and prints a one-line cleanup command at the end.report.jsonrecords the installed packages, HF snapshot hashes, the normalizer, and run conditions (Transcripted running, on battery). Home paths are scrubbed.engines/apple_speech.swift: Apple SpeechAnalyzer runner, using the same calls as Add Apple Speech as a transcription engine choice #1776.engines/whisperkit-bench/: WhisperKit at the app's pinned revision, with the app's decode options.engines/py_engines.py, APIs checked against each package's source:hillclimb_bench.py: speaks the hill-climb lab's v1 bench protocol (Hill-climb lab: tune app settings against scored, held-out tests #1791). The knobstt.enginepicks the model, and it reports per-item metrics..agents/test-matrix.yml: a rule forscripts/stt-shootout/**.How I checked it
scripts/dev/agent-preflight.shbash -n scripts/stt-shootout/run.sh,py_compile,shootout.py --self-test,hillclimb_bench.py --self-testtest-matrix-checks.py,agent-context.pyandagent-check.pyRisk Review
~/stt-shootoutexcept Apple's OS-managed speech assets, and the app's model files are never written.Notes
[skip ci]while 1.1.62's release PRs hold the Mac runners. Justin will run it on his Mac after the release build.DownloadUtils.enforceOfflinearound every--models-dirload (Sources/Speech/ParakeetLocalModelLoader.swift), so a failed load can't delete or re-download the folder. The APFS clone here already protects the app's files without it. Keep the clone anyway, because the shootout runs whatever CLI is installed.🤖 Generated with Claude Code
https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Mac results (2026-09-28, Apple M5 Max)
Combined with main plus #1862 and #1791, these passed:
bash -n scripts/stt-shootout/run.sh,py_compilefor the shootout scripts,shootout.py --self-test,hillclimb_bench.py --self-test, andtest-matrix-checks.py --self-test. The harness itself (the full engine shootout) wasn't re-run here.