Skip to content

Speech-to-text model shootout harness - #1788

Merged
r3dbars merged 14 commits into
mainfrom
claude/stt-model-shootout-0647al
Sep 28, 2026
Merged

r3dbars merged 14 commits into
mainfrom
claude/stt-model-shootout-0647al

Conversation

@claude

@claude claude Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Requested by Justin · project thread

Before: there was no way to see how other on-device speech-to-text models compare with the Parakeet V3 model the app ships, on speed or accuracy.

After: one command on an Apple Silicon Mac (bash scripts/stt-shootout/run.sh) downloads an hour-long, human-captioned YouTube lecture and runs 15 models on it. The report table gives time for the whole hour, speed vs real time, speed vs Parakeet V3, 10-second clip latency (warm and first use), load time, peak memory, and word error rate against the captions.

Why

Justin wants to know which models are faster than Parakeet V3, and how accurate they are, before adding any to the app.

Product Impact

  • Affects: docs only (bench tooling under scripts/stt-shootout/; no app code changes)
  • Lane: agent workflow
  • Why this matters: picks the next STT model with real numbers from his Mac, and plugs into the hill-climb lab as a bench.

What changed

  • scripts/stt-shootout/shootout.py: the runner.
    • Fetches audio plus human captions with yt-dlp and converts the audio with afconvert.
    • Runs each model in its own process and its own env, and polls peak memory with proc_pid_rusage.
    • Scores WER with Whisper's English normalizer and jiwer.
    • Writes report.md/json/csv and the transcripts.
  • The Parakeet V3 and Ultra rows run the installed app's transcripted-cli on an APFS clone of the model folder. FluidAudio deletes and re-downloads a folder it fails to load, and the clone keeps that from ever touching the app bundle or the Ultra install.
  • Every download stays under ~/stt-shootout: HF_HOME, uv's cache and Python, WhisperKit, whisper.cpp and Moonshine. The run stops below 20 GB free and prints a one-line cleanup command at the end.
  • Pip packages are pinned. report.json records the installed packages, HF snapshot hashes, the normalizer, and run conditions (Transcripted running, on battery). Home paths are scrubbed.
  • engines/apple_speech.swift: Apple SpeechAnalyzer runner, using the same calls as Add Apple Speech as a transcription engine choice #1776.
  • engines/whisperkit-bench/: WhisperKit at the app's pinned revision, with the app's decode options.
  • engines/py_engines.py, APIs checked against each package's source:
    • mlx-whisper: turbo, distil
    • parakeet-mlx: v3, v2
    • onnx-asr: Canary 1B v2, 180M Flash
    • moonshine-voice: base, medium
    • pywhispercpp: turbo
    • mlx-audio: Granite Speech 4.0 1B, Nemotron streaming
  • hillclimb_bench.py: speaks the hill-climb lab's v1 bench protocol (Hill-climb lab: tune app settings against scored, held-out tests #1791). The knob stt.engine picks the model, and it reports per-item metrics.
  • .agents/test-matrix.yml: a rule for scripts/stt-shootout/**.

How I checked it

  • scripts/dev/agent-preflight.sh
  • bash -n scripts/stt-shootout/run.sh, py_compile, shootout.py --self-test, hillclimb_bench.py --self-test
  • Self-tests for test-matrix-checks.py, agent-context.py and agent-check.py
  • Linux dry run end to end with a fake model and a fake app bundle and CLI. It covered model staging, per-file CLI JSON, the report and the JSON output.
  • Two independent reviews against the package sources. Every finding is fixed.
  • Real run on Justin's Mac. The Swift helpers have never been compiled because this session has no Swift toolchain. A compile failure only drops that one row.

Risk Review

  • Privacy / local-first behavior reviewed: everything runs locally. The network is used only for the video, packages and model weights.
  • Storage path impact reviewed: nothing is written outside ~/stt-shootout except Apple's OS-managed speech assets, and the app's model files are never written.
  • No private transcripts, audio, tokens, personal paths, or customer data are included

Notes

  • Commits carry [skip ci] while 1.1.62's release PRs hold the Mac runners. Justin will run it on his Mac after the release build.
  • Once Try Parakeet Ultra as an experimental model #1783 merges, the app's CLI sets FluidAudio's DownloadUtils.enforceOffline around every --models-dir load (Sources/Speech/ParakeetLocalModelLoader.swift), so a failed load can't delete or re-download the folder. The APFS clone here already protects the app's files without it. Keep the clone anyway, because the shootout runs whatever CLI is installed.

🤖 Generated with Claude Code

https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr

Mac results (2026-09-28, Apple M5 Max)

Combined with main plus #1862 and #1791, these passed: bash -n scripts/stt-shootout/run.sh, py_compile for the shootout scripts, shootout.py --self-test, hillclimb_bench.py --self-test, and test-matrix-checks.py --self-test. The harness itself (the full engine shootout) wasn't re-run here.

Benchmark harness for comparing on-device speech-to-text models against
the Parakeet V3 model the app ships: downloads a human-captioned test
video, runs each model in its own process, and reports speed, latency,
peak memory and word error rate.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
@claude claude Bot assigned r3dbars Sep 23, 2026
@claude
claude Bot requested a review from r3dbars September 23, 2026 21:10
Runs Whisper large-v3-turbo through the same WhisperKit revision and
decode options the app's Whisper model choice uses.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
… STT shootout [skip ci]

Adds Whisper turbo and Distil-Whisper (mlx-whisper), Parakeet V3/V2
(parakeet-mlx), Canary 1B v2 and 180M Flash (onnx-asr + Silero VAD),
Moonshine base and medium (moonshine-voice), whisper.cpp turbo
(pywhispercpp), Granite Speech 4.0 and Nemotron streaming (mlx-audio).
APIs checked against each package's source.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
From a review against each package's source:
- parakeet-mlx: hand its loader float32 (bf16 broke the mel spectrogram)
- a failed video download moves on to the next video; prefer AAC audio
  afconvert can read; give yt-dlp a JavaScript runtime (deno extra)
- drop uv venv --clear (older uv lacks it); cap setup steps at 45 min
- keep results a model wrote before crashing in teardown
- read transcripted-cli results from per-file JSON, not stdout
- pin WhisperKit's swift-transformers/swift-jinja to its own resolved set
- strip caption speaker labels mid-line; guard report math

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
From the deep review of #1788:
- run the Parakeet V3/Ultra rows on an APFS clone of the model folder,
  since FluidAudio deletes and re-downloads a folder it fails to load;
  fail the Ultra row if its marker is gone after the run
- key the cached video by its URL list so --url can't reuse another video
- keep every download (HF_HOME, uv cache and Python, whisper.cpp,
  Moonshine) under ~/stt-shootout; stop below 20 GB free; print cleanup
- pin pip packages; record packages, HF snapshot hashes, normalizer,
  and run conditions (app running, battery) in report.json
- label peak memory honestly for Core ML and Apple Speech rows
- reuse results only when clip settings match; scrub home paths
- add a test-matrix rule for scripts/stt-shootout

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
hillclimb_bench.py speaks the lab's transcripted.hillclimb v1 protocol:
knob stt.engine picks the model, each suite item is one recording, and
it reports time, speed, latency, load, memory and WER per item.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
claude Bot pushed a commit that referenced this pull request Sep 23, 2026
Adds the stt-shootout bench (script lives in PR #1788), a stt.engine knob
with the shootout's engine list, a Mac-local stt-clips suite, and a
speech-model-accuracy objective: fewer wrong words, at most ~35% slower,
no engine failures or empty text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS
@claude claude Bot mentioned this pull request Sep 23, 2026
12 of 15 tasks
… [skip ci]

Every Transcripted meeting's audio is named microphone.m4a or
system_audio.m4a, so naming the converted WAV by stem alone let a later
hill-climb item be scored on the first item's audio. The converted copy
is now keyed by source path, size and mtime. report.json also records
which WhisperKit model files were measured (Hub commit when available,
plus a size fingerprint).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
claude Bot pushed a commit that referenced this pull request Sep 24, 2026
… lab typecheck [skip ci]

Second deep-review round on #1791:
- N1: only the outermost run_group starts a new session; it marks the
  child env so an adapter's own run_group keeps the app/CLI/harness in the
  same group. The climber's timeout now kills the real work too. Two-level
  tests (the outer one fails on the old code).
- S2 follow-up: lab_control launch turns analytics and crash reporting off
  for the launched process only, as NSArgumentDomain launch arguments, so
  the person's saved Settings are never needed or changed.
- N2: scripts/dev/typecheck-lab-build.sh type-checks the app with
  -D TRANSCRIPTED_LAB_CONTROL; app-build CI runs it after the normal build.
- N3: a holdout check writes a started row before it runs, so a crash still
  uses up budget.
- N4/N5: guide notes on #1788's audio naming and LabKnobOverrides' env var.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
- onnx-asr's Canary decoder keeps words generated after end-of-text when
  segments are batched, which made Canary 180M write 15,489 words for a
  7,191-word answer key. Decode one VAD segment at a time.
- onnxruntime 1.30 rejects a .onnx.data file symlinked out of the model
  folder (the Hugging Face cache layout), so Canary 1B v2 never loaded.
  onnx-asr models now download into plain folders under models/onnx-asr,
  and their Hub commits still land in the report.
- Granite Speech needs jinja2 for its chat template.
- Rows whose word count is far off the answer key (or WER over 50%) are
  marked broken and left out of the pick.
- Record whether Transcripted was running or the Mac was on battery right
  before and after each model, not only at report time.
- WhisperKit's model record looked for the bare variant name, but the
  folder is openai_whisper-<variant>.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
`--rerun canary-1b-v2,granite-speech` redoes only those models and reuses
every other saved result, so re-running a few fixed rows still produces a
complete report. Plain `--rerun` still redoes everything.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Brings in Parakeet Ultra (#1783) so the shootout worktree on the Mac has its install script.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012s4yhheGj3PGz1t9FCgvTr
Picks up the VM test socket-path fix (#1883) and everything merged today.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@r3dbars
r3dbars marked this pull request as ready for review September 28, 2026 02:53
r3dbars added a commit that referenced this pull request Sep 28, 2026
* Add hill-climb lab core: knob registry, held-out splits, paired stats, climber [skip ci]

Python stdlib tool under scripts/hillclimb that tunes app knobs against
scored benches. Dev/holdout splits are hash-stable, verdicts need a
bootstrap-CI win above min_effect with no guardrail or hard-gate
regression, timing benches interleave A/B runs, and holdout checks are
budgeted. A synthetic demo registry backs the self-tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* Let the lab override 8 meeting-pipeline knobs from a file [skip ci]

LabKnobOverrides reads TRANSCRIPTED_LAB_KNOBS_FILE once per process. With
no env var nothing changes: every call returns today's default and there is
no file I/O. Wired: diarizer clustering threshold, VBx Fa/Fb, min segment
duration, and same-voice consolidation / small-cluster absorb per embedder.
Not compiled yet (no Swift toolchain in session).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* Add dictation stop-latency bench adapter and 64-phrase suite [skip ci]

Synthesizes fixtures with say, runs the real DictationStopBenchmarkRunner
once per repetition in an isolated HOME, and reports per-phrase latency,
word error rate (from the saved Markdown), and missing-text, silence-text
and unstable-output gates. Self-test now also runs bench adapter tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* Add real knob registry, objectives, speaker naming bench, and lab guide [skip ci]

101 knobs (10 live, 21 bench-only, 70 mapped but hardcoded) with verified
defaults and source lines; three objectives (dictation stop latency,
meeting turnaround, speaker naming across calls); a SpeakerEvalHarness
autoeval adapter with per-cache items and every safety counter as a hard
gate; docs/hill-climb-lab.md; a test-matrix rule and a repo-hygiene step
that runs the lab self-test and registry validation.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* Add meeting turnaround bench adapter and corpus suite builder [skip ci]

Runs transcripted-cli import-audio per corpus item with a fresh empty
speaker database, times Stop-to-transcript per second of audio, scores
word recall and speaker count against truth, and passes lab knob
overrides through TRANSCRIPTED_LAB_KNOBS_FILE (erroring if the CLI did not
confirm them). One untimed warmup import absorbs model load.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* Add a lab control channel so an agent can drive the real app [skip ci]

Off unless the app is launched with TRANSCRIPTED_LAB_CONTROL_DIR. Commands
dropped as JSON into <dir>/inbox (ping, status, start/stop dictation,
start/stop meeting, import audio) call the same entry points the menus
use; responses go to <dir>/responses.jsonl. Nothing from the channel goes
off-device. scripts/hillclimb/lab_control.py launches the app and sends
commands. Swift not compiled yet (no toolchain in session).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* Plug the speech model shootout into the lab [skip ci]

Adds the stt-shootout bench (script lives in PR #1788), a stt.engine knob
with the shootout's engine list, a Mac-local stt-clips suite, and a
speech-model-accuracy objective: fewer wrong words, at most ~35% slower,
no engine failures or empty text.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* hillclimb: kill the whole process group when a bench adapter times out [skip ci]

subprocess.run(timeout=) only kills the direct child, so the app, the CLI or
the speaker harness kept running into the next trial and skewed timings.
Adds hc_proc.run_group (new session + killpg) and routes every adapter
subprocess through it, with grandchild-survival tests per adapter.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* hillclimb: exact sign-flip tests, minimum suite sizes, fail-closed guardrails [skip ci]

Fixes the deep review's stats and holdout findings on #1791:
- S4: verdicts use a one-sided paired sign-flip permutation test (exact up
  to 16 units) instead of the percentile bootstrap; the holdout check uses
  p < 0.01. Climb and confirm refuse suites under 10 dev / 8 holdout
  independent units. simulate-null reports false-accept rates at real sizes.
- S5: a guardrail measured on fewer than half the primary's units, or that
  lost items, rejects. Metrics can declare the item field they need
  (truth, speakers, text) and the lab refuses suites that can't feed them.
- S6: holdout peeks count by holdout item overlap (>50% = same holdout), so
  adding an item no longer resets the budget. Holdout per-item values are
  sealed out of trials.jsonl and bench work dirs go to holdout-sealed/.
- S7: suite items can carry a cluster; clusters count once and never
  straddle the holdout line. Speaker items cluster by family and identity
  split, which blocks that objective until the harness emits per-person
  rows. Its recommendation now lists the per-bucket contract as required.
- S8: CommandBench runs through hc_proc.run_group.
- M1: the inconclusive re-measure pools with the first run at alpha/2.
- M2: every repetition's build/host/OS must match.
- M3: climb-result.json is checkpointed after each decision; climb --resume.
  Malformed bench results become item errors instead of crashing.
- M4: forced holdout checks are recorded as forced.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* lab control: compile the channel only into local lab builds, lock its folders [skip ci]

Fixes the deep review's B1, S2, S3, M6, M8 and M9 on #1791:
- B1: LabControlChannel and its launch hook sit behind
  #if TRANSCRIPTED_LAB_CONTROL, set only by build.sh --lab. build-beta.sh
  refuses TRANSCRIPTED_LAB_BUILD and fails if the binary contains the
  channel's env var name. In lab builds the control dir, inbox/ and done/
  must be real 0700 dirs owned by this uid, re-checked every poll.
- M6: command files open O_NOFOLLOW|O_NONBLOCK and are fstat-checked
  (regular, ours, <= 64 KB) before reading; done/ moves use rename(2).
- S3: stop_dictation pastes only when paste is explicitly true;
  start_dictation calls the session directly and never activates an app.
- S2: lab_control.py launch needs --container (or --use-real-library),
  refuses a relocated capture library and telemetry-on unless overridden.
- M9: LabKnobOverrides drops unknown ids with one stderr line.
- M8: Support and Core CLAUDE.md list the lab files and Core's env var.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* hillclimb: #1789 merge note hooks, source revision in recommendations, models-dir note [skip ci]

Review S1/M5/M7 on #1791: the clustering knob's notes say it's in cosine
units and point at the merge note for #1789's FluidAudio 0.17 distance
change; recommendation.json records the revision its source line numbers
came from; the guide explains the shared FluidAudio model cache.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* hillclimb: nested timeouts kill the real work, telemetry off per run, lab typecheck [skip ci]

Second deep-review round on #1791:
- N1: only the outermost run_group starts a new session; it marks the
  child env so an adapter's own run_group keeps the app/CLI/harness in the
  same group. The climber's timeout now kills the real work too. Two-level
  tests (the outer one fails on the old code).
- S2 follow-up: lab_control launch turns analytics and crash reporting off
  for the launched process only, as NSArgumentDomain launch arguments, so
  the person's saved Settings are never needed or changed.
- N2: scripts/dev/typecheck-lab-build.sh type-checks the app with
  -D TRANSCRIPTED_LAB_CONTROL; app-build CI runs it after the normal build.
- N3: a holdout check writes a started row before it runs, so a crash still
  uses up budget.
- N4/N5: guide notes on #1788's audio naming and LabKnobOverrides' env var.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* hillclimb: add #1789's speaker-lab bench, knobs and objective [skip ci]

Pasted from #1789's speaker_lab.README.md: the speaker-lab command bench,
6 new knobs (diarization backend, Nemotron preset, match mode and floor,
replay dedup, write-path fixes), speaker-lab-recognition, and an identical
copy of its suite. The adapter lives on #1789, so trials are item errors
until it merges. The objective reports BLOCKED: 4 holdout series, need 8.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* hillclimb: take #1789's 24-series speaker-lab suite (8 holdout) [skip ci]

Byte-for-byte copy of config/hillclimb/suites/speaker-lab-ami.json from
#1789 at ffe5871 (salt speaker-lab-ami-v2, 16 dev / 8 holdout series), so
speaker-lab-recognition is no longer BLOCKED.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* hillclimb: speaker-lab objective is no longer blocked; update notes [skip ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M8VU2yBBnE8eH5f18mubhS

* Hill-climb test: compare resolved paths for the bench HOME

The adapter resolves the request path, so on macOS a temp dir under /var
comes back as /private/var and the startswith check failed. The suite
had only run on Linux. Now 164/164 on macOS too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: r3dbars <r3dbars@users.noreply.github.com>
@r3dbars
r3dbars merged commit 452053d into main Sep 28, 2026
8 checks passed
@r3dbars
r3dbars deleted the claude/stt-model-shootout-0647al branch September 28, 2026 08:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants