Skip to content

Meetings: Nemotron 3 speaker separation + sooner naming, on by default (FluidAudio 0.17) - #1887

Merged
r3dbars merged 41 commits into
mainfrom
claude/yodas-speaker-lab
Sep 28, 2026
Merged

r3dbars merged 41 commits into
mainfrom
claude/yodas-speaker-lab

Conversation

@r3dbars

@r3dbars r3dbars commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Why

Speaker separation is the foundation for everything else in meetings. Today a 7-person call comes back as about 4 voices, with roughly half the words under the wrong person. A 1:1 can show a naming sheet with 10+ rows for one person. And a person needs 5 confirmed meetings before we name them silently.

This PR adds a lab that measures all of that on simulated calls with an answer key (YODAS3, CC BY 3.0 YouTube speech). It then ships what won, as defaults:

  1. NVIDIA Nemotron 3 Diarization splits the call into speakers.
  2. Voiceprints come from the same WeSpeaker model the saved people came from, so existing people keep matching.
  3. Cleanup and naming on top. Tiny split-off voices fold back into a real person. A one-on-one invite caps the call at one voice. People on the invite (or, with no invite, the 12 heard most recently) get named after 2 confirmed meetings instead of 5.

It also brings in #1789's FluidAudio 0.17.0 upgrade, which Nemotron needs.

Product Impact

  • Affects: meetings, dictation (via FluidAudio 0.17)
  • Lane: meeting reliability
  • Why this matters: on 45 fresh holdout meetings, never used for tuning, run through the real TranscriptionTaskManager:
Holdout set Today: exactly right Now: exactly right Today: words right Now: words right
1:1 calls 80% 100% – –
3–4 people 83% 100% 80.5% 85.0%
6–8 people 0% (27 people missed) 80% (2 missed, 0 blended) 54.8% 82.8%
Hard mode (overlap, sound-alikes) 0% (88% blended) 62% (12% blended) 56% 84%

Stress sets:

Stress set Today: exactly right Now: exactly right Words right, today → now
Very short calls 55% 80% (0 blends) 81% → 88%

The noisy, 8/10-person, and hour-long sets are still running and will be added to YODAS_LAB_RESULTS.md.

Naming, end to end on multi-week company series: naming work is −31% with invites and −37% with none. Regulars are named automatically 20/21 times (vs 8/21), from their 3rd–4th meeting. 0 wrong names.

What changed

App defaults (no toggles):

  • DiarizationBackendPreferences.defaultChoice is .nemotron. The hidden switch back is defaults write com.justinbetker.draft diarization-backend-preference pyannote or TRANSCRIPTED_DIARIZATION_BACKEND=pyannote.
  • DiarizationService: if Nemotron can't load (offline on first use, a bad download), pyannote stands in for that session. activeBackend records which one ran, and cleanup() retries Nemotron next time.
  • New FluidOfflineWeSpeakerSegmentEmbedder: it embeds each Nemotron turn with the pyannote path's own FBank/Embedding models, in a real 10 s context window with a mask over the turn (zero padding skews the features). Lab parity against native pyannote vectors: 0.986–0.995 speaker-level cosine, 29/29 clusters matched top-1. It shares speakers.sqlite, so no migration is needed and saved people keep matching.
  • SpeakerSeparationOptions.tuned(for:invitedPeople:):
    • Nemotron: fold voices under 5 s, and cap at 1 only for one-on-one invites. No merge.
    • pyannote fallback: the lab-tuned 0.70 / fold / merge 0.6 / invite cap.
  • Lineup naming is always on (MeetingCalendarNaming).
  • Removed the two beta toggles and SpeakerSeparationPreferences/CalendarNamingPreferences.

From #1789: FluidAudio 0.17.0, DiarizationBackend, NemotronDiarizationRunner, NemotronTurnBuilder, and the FluidAudioCompatibility clustering conversions.

Lab: scripts/speaker_lab/* and speaker-eval-harness meeting-series, which now takes --backend and --sep-* flags and has stress families. ami_to_lab.py is included. The lab never touches the real speaker DB, library, stats, or prefs.

Dictation trade-off (FluidAudio 0.17)

  • Normal speech and silence decode about 40% faster.
  • A take that is only background noise returns "nothing heard" about 150 ms later (53 → 197 ms). That comes from 0.17's empty-decode recovery ladder: when Parakeet returns no words on audio that isn't silent (2 s or longer, RMS ≥ 0.003), it retries up to 5 more ways to rescue real speech 0.15 used to drop. On pure noise all 5 retries come up empty.
  • Short-clip WER is +0.67 pp.

Shipping it anyway.

How I checked it

  • bash build-deps.sh --force + bash build.sh --no-open + bash run-tests.sh (TZ=America/Chicago): 19,801 passed, 0 failed
  • bash run-integration-smoke.sh: pass
  • swift test --filter SpeakerTests: 390 passed, 0 failed (full swift test left to CI)
  • check-source-pins.py --changed-only: PASS
  • Lab: 45-meeting holdout, the short-call stress set, and companies end to end through the real pipeline

Mac or hardware test still needed? Yes. Record a 1:1 and a 5+ person call on this build. Expected: one row on the 1:1, and each person once on the big call. The log shows "Nemotron diarizer models loaded" with embedder: wespeaker.

Risk Review

  • Privacy / local-first: all local, no new analytics or Sentry keys
  • Storage: Nemotron shares speakers.sqlite (same embedding space); no migration
  • Release: the Nemotron fast128 model (193 MB) is not bundled yet. The first meeting on each Mac downloads it, and falls back to pyannote if the download fails. Bundling it in build-beta.sh is a follow-up.
  • Text pins checked
  • Independent deep review of the full diff: not done, merged on the owner's call

🤖 Generated with Claude Code

0.17.0 adds Nemotron 3 Diarization, which the speaker lab needs. The bump
itself should change nothing users see, so:

- build-deps: tools 6.2 manifest so FluidAudio can opt out of its default
  NemoTextProcessing trait (a prebuilt Rust static lib we never call and
  would not archive). swiftLanguageModes [.v5] keeps in-tree targets on
  the language mode they built under before.
- DiarizationService: clusteringThreshold is now a Euclidean cut distance,
  so the tuned 0.6 cosine becomes sqrt(0.8); constrainedAssignment pinned
  off to match 0.15.x assignment.
- Parakeet (app + CLI): keep the 0.15.x long-form chunking
  (melChunkContext on, no seam-gap repair).
- Integration fake FluidAudio grows the ASRConfig shape.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
DiarizationBackendPreferences reads TRANSCRIPTED_DIARIZATION_BACKEND, then
the diarization-backend-preference default, then falls back to pyannote.
No Settings UI. Not wired into meetings yet; that lands with the Core
Nemotron backend.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… ci]

DiarizationBackend (.pyannote default, .nemotron experimental) picks the
model. The pyannote path is unchanged. The Nemotron path loads a FluidAudio
Nemotron3 preset (fast128 by default, TRANSCRIPTED_NEMOTRON_PRESET for the
lab), runs it on a private serial queue, turns frame probabilities into
exclusive speaker turns with the pure NemotronTurnBuilder, and embeds each
turn with the injected embedder or a new FluidWeSpeakerSegmentEmbedder,
since Nemotron emits no voiceprints.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
NemotronTurnBuilder: exclusivity, argmax on overlap, ties, gap bridging,
min-duration drop and blip rejoin, first-appearance remapping, empty and
malformed input. DiarizationService: pyannote default, backend round-trip,
not-loaded error, thresholds, shared reembed path. Nemotron preset override
and the WeSpeaker fallback embedder's pure helpers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
MeetingSessionController passes the chosen backend to DiarizationService.
With Nemotron and no ERes2Net, voiceprints come from FluidAudio's online
WeSpeaker model, so they go to their own speakers_wespeaker-fluid-online
database (meetings and Settings > People agree) until the lab proves they
match the offline pipeline's vectors. Default pyannote is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… thresholds and fingerprint-update knobs [skip ci]

dump takes --backend pyannote|nemotron and --embedder native|eres2net
(--eres2net-model for a path-based model) and records backend, embedder,
dimension, wall time, and audio length. Old dumps still decode.

replay takes --thresholds auto|weSpeaker|eRes2Net, --match adaptive,
--same-voice, --dedup, and the write-back EMA knobs. Defaults reproduce
the old behavior for WeSpeaker dumps. Per meeting it now reports raw
cluster count and whether each cluster matched an existing profile.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… [skip ci]

speaker_eval_common.py now holds RTTM parsing, overlap helpers, the
identity metrics (fragmentation, false merge, re-ID curve), and a pure
Python DER/JER that matches pyannote.metrics to float precision
(collar convention, UEM, overlap, Hungarian mapping). It is about 40x
faster, which matters once the lab scores many replays.

score_speaker_eval.py uses it and no longer needs pyannote installed.
Its markdown and JSON output are unchanged on a synthetic fixture.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…ng-speaker recognition [skip ci]

score_speaker_lab.py scores a lab run directory into scores.json and
REPORT.md: per variant, raw diarizer DER/JER and count error, pipeline
DER after clustering and DB matching, and the recognition scoreboard
(recognized / wrong person / asked again / undetected, plus new people
false-matched to a known profile). It picks each variant's best knob
setting and shows which knob values moved recognition. Own-calls mode
reports behavior without ground truth, agreement against a baseline
variant, and writes a self-contained timeline.html.

Also has helper subcommands the driver uses: grid, dump-ok, and
own-calls-list. Unit tests run on Linux with synthetic fixtures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…ingle trial [skip ci]

run_speaker_lab.sh dumps each meeting once per variant (backend x
embedder x Nemotron preset) into its own cache, replays every knob
setting, and scores it all into reports/speaker-lab/<stamp>/. Every
knob is a flag with an env twin, the run is non-interactive, fails loud,
and prints the scores.json path as its last stdout line. --single runs
one variant at one setting for an outer optimizer. --own-calls runs the
same variants on saved meetings (call track read in place) and writes a
timeline page.

Tests drive the script end to end against a fake harness, so the
orchestration is covered on Linux too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…s [skip ci]

16 scenario series (8 Edinburgh, 4 Idiap, 4 TNO), 4 sessions each with
the same 4 people, for the speaker lab's cross-call recognition test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…[skip ci]

README covers what the lab measures, how to run it on AMI and on your
own saved calls, every knob, how an optimizer drives --single trials,
the scores.json schema, and how to add a diarizer or embedder. The
test matrix now runs the lab's syntax checks and unit tests when the
lab files change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
….17 [skip ci]

Review of #1789 found the bump was not as behavior-neutral as claimed:

- FluidAudio 0.17 pins speaker-diarization-coreml to one commit and
  deletes any cache without a matching revision marker. Every 0.15.x
  cache, and the offline-diarizer-models copy bundled in the app, has
  none, so the first meeting would delete inside the signed bundle (or
  fail offline). FluidAudioCompatibility.keepUnpinnedDiarizerCaches()
  resolves that repo at main, as 0.15.x did; DiarizationService,
  FluidWeSpeakerSegmentEmbedder and the CLI call it before loading.
- The tuned pyannote config moves to
  FluidAudioCompatibility.tunedOfflineDiarizerConfig() and a test pins
  the distance threshold, constrainedAssignment=false and every tuned
  value.
- CLI diarize/batch start from the 0.15.x default (cosine 0.6, no
  constrained assignment), and a config file's clusteringThreshold stays
  a cosine similarity, converted for 0.17.
- Parakeet comments now say only the two new defaults are pinned: 0.17
  also changed >15 s chunk merging with no switch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… [skip ci]

- THIRD_PARTY_LICENSES: FluidAudio 0.17's Japanese G2P (Misaki ports,
  UniDic) and Spanish/French lexicon notices; NeMo text processing is not
  linked.
- Parakeet comment: 0.17 retries blank v3 decodes up to five more times,
  so near-silent dictation can take longer and return text.
- build-deps: FluidAudio's resource bundle never ships, so LuxTts G2p must
  stay unused.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…and valueless flags [skip ci]

A dump with an unknown TRANSCRIPTED_NEMOTRON_PRESET recorded the typo while
Core's runner silently ran fast128, so the cached dump lied about its variant.
Validate with FluidAudio's Nemotron3Config.preset(named:) and record 'default'
for unset. --write-path-fixes now rejects anything but on|off (a typo used to
run the legacy path), and a flag with no value is an error instead of the default.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…the online WeSpeaker embedder [skip ci]

Diarizes one file with today's pipeline, re-embeds the same segments with
FluidWeSpeakerSegmentEmbedder (what Nemotron falls back to), and reports
per-segment, within-model, cross-model and per-cluster cosine stats against the
WeSpeaker match floor, plus a looksInterchangeable heuristic. Answers whether
Nemotron voiceprints could share speakers.sqlite instead of a separate DB.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Add DiarizationService.resolvedNemotronPresetName(environment:), a public
wrapper over the runner's internal resolver, so the harness records 'fast128'
instead of 'default' when TRANSCRIPTED_NEMOTRON_PRESET is unset, and refuses a
name Core would silently replace. The lab scorer treats unset, 'default' (older
dumps) and 'fast128' as one variant so existing caches stay reusable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
scripts/hillclimb/benches/speaker_lab.py speaks the hill-climb request/result
protocol (#1791) and drives run_speaker_lab.sh --single once per trial over
the requested AMI series, then splits scores.json + recognition-events.json
back into per-series items: recognition_rate, recognized, asked_again,
pipeline/raw DER, speaker-count error, objective, plus int gates
wrong_person and new_person_false_match. Missing/incomplete series and
series without returning speakers are item errors, never zeros. Driver env
twins are scrubbed so shell exports can't leak into a trial, and scores.json
must echo every knob that was set. app_revision hashes the harness binary,
the requested RTTM/audio, and the lab scripts.

Imports hc_benches when the climber is on the branch, else uses a local copy
of the protocol validator, so it works before and after the merge.

Also: config/hillclimb/suites/speaker-lab-ami.json (the 16 download_ami.sh
lab series, 12 dev / 4 holdout stratified by site) and
speaker_lab.README.md with the benches/knobs/objectives JSON for #1791.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Changes to the adapter, its tests/README/suite, or the lab driver/scorer it runs now select py_compile + speaker_lab.py --self-test. Kept as its own rule above the SpeakerEvalHarness block so it doesn't collide with #1791's scripts/hillclimb/** rule.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…eports it [skip ci]

run_speaker_lab.sh --embedding-parity (EMBEDDING_PARITY=1) runs the harness's
embedding-parity per meeting (RTTM labels on corpora, pyannote clusters on own
calls), cached in data/eval/<corpus>/parity/. The scorer pools the reports via
their 200-bin histograms into an additive embeddingParity block in scores.json
and an 'Embedding parity' section in REPORT.md, with the same interchangeability
rule as the harness. schemaVersion stays 1; runs without the flag are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
A cache written by a pinned FluidAudio 0.17 build carries a
.fluidaudio-revision marker naming a commit. The app resolves the
diarizer repo at main, so a shipped marker naming anything else would make
it delete files inside its own signed bundle. build-beta drops the marker
after copying the cache.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…kip ci]

Rebased onto e3a559f. The dump now records the resolved Nemotron preset and the scorer treats unset/default/fast128 as one variant, so the knob echo check normalizes presets the same way. fast128 is still not passed because the driver names variants after the preset string (passing it would fork the dump cache). EMBEDDING_PARITY joins the scrubbed env twins; --embedding-parity is a diagnostic and never passed per trial.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…ip ci]

The 0.17 bump can't merge until WER is no worse and blank-audio dictation
stops aren't slower (PR #909 retries blank V3 decodes up to 5 times).

scripts/stt_fluidaudio_ab.sh builds transcripted-cli twice in git worktrees
under ~/stt-fluidaudio-ab: the baseline from origin/main (pre-bump source,
since 0.15.4 can't compile the new ASRConfig call) with FLUID_AUDIO_VERSION
0.15.4, the candidate from HEAD with 0.17.0. Builds are reused while their
inputs are unchanged. It refuses a baseline ref that already uses the 0.17
API and checks the version SwiftPM actually resolved.

scripts/stt_fluidaudio_ab.py then measures, reusing the STT shootout's
lecture download, caption parsing and WER scorer (read from git when the
shootout isn't in the checkout):
- WER: first 20 min of the lecture plus 10 caption-aligned ~45 s pieces
- stop time: silence 1/3/8 s, near-silence, low noise, room rumble, a
  0.5 s noise burst (gated), plus `say` clips (reported), ABBA rounds,
  median of CLI processingSeconds
Gate flags: --max-wer-delta-pp 0.5, --max-time-ratio 1.25,
--max-time-delta-ms 300. Writes result.json + report.md; last stdout line
is the result.json path.

scripts/test_stt_fluidaudio_ab.py covers the math, fixtures and a fake-CLI
end-to-end run, and drives the bash wrapper against a throwaway repo with
stubbed Mac tools. Runs on Linux.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…t groups [skip ci]

The hill-climb lab needs >= 10 dev and >= 8 holdout independent units before an
objective can confirm anything; 16 series (12/4) left speaker-lab-recognition BLOCKED.
Adds ES2004 ES2011 ES2012 ES2013 IS1006 IS1007 TS3007 TS3008 (all four sessions have
pyannote RTTMs, and no participant appears in two series). Holdout re-picked per site by
lowest unit_hash under a new salt (speaker-lab-ami-v2): 4 ES, 2 IS, 2 TS.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Resolves the Core and SpeakerEvalHarness CLAUDE.md file maps: keeps main's
per-file map and adds this branch's diarization backend, FluidAudio
compatibility and speaker lab entries.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
claude and others added 7 commits September 24, 2026 06:38
…[skip ci]

main's --models-dir guard (keeps FluidAudio from deleting a Parakeet Ultra
install that fails to load) used DownloadUtils.enforceOffline, which 0.16
renamed to ModelHub.offlineMode. Same semantics: loadModels rethrows
instead of purging and re-downloading. Also corrects the local-install
layout comment: v3 file names are unchanged through 0.17.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
build-deps.sh: keep FluidAudio 0.17.0 with traits: [] and take main's
removal of mlx-swift-lm.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
A lab for speaker separation and naming. It builds meetings with an answer key
from YODAS3 voices and runs them through the real meeting pipeline and naming
sheet headless.

- scripts/speaker_lab/voicebank.py: YODAS3 shards -> verified single-voice
  identities (NVIDIA TitaNet and 3D-Speaker CAM++ must both agree; long voices
  densely re-checked)
- scripts/speaker_lab/meeting_sim.py: mic + call channels (per-person device,
  real Opus, room reverb, backchannels, overlap, echo leak) plus a noisy
  calendar invite; families A-E and multi-week companies (F, sound-alike mode)
- speaker-eval-harness meeting-series: real TranscriptionTaskManager over
  throwaway paths with a stub stats store; answers the naming sheet the way
  SpeakerNamingSheet builds updates
- speaker-eval-harness dump-e2e / dump-set: NVIDIA Sortformer and LS-EEND from
  our FluidAudio build, and the raw production diarizer under a knobs file
- score.py, naming_replay.py: rows vs people, fragments, merged voices, word
  attribution, learning speed and wrong names, offline naming-policy replay

Two Core seams the app never sets: TranscriptionTaskManager
.pipelineResultObserver (the lab reads the exact Phase 1 result) and
DiarizationService.labSpeakerBounds (per-meeting speaker-count bounds, added
below the lines config/hillclimb/knobs.json pins).

YODAS3 is CC BY 3.0; all audio and derived data stay under data/eval/
(gitignored). Plan and results: Tools/SpeakerEvalHarness/YODAS_LAB_*.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both default off. Tuned in the YODAS3 speaker lab and checked on 45 fresh
meetings the tuning never saw (Tools/SpeakerEvalHarness/YODAS_LAB_RESULTS.md).

Separate voices on calls (beta):
- SpeakerSeparation.swift: after the diarizer, fold voices with under 5 s of
  talk into the voice they sound most like, merge fingerprints at 0.6 and up,
  and cap at the calendar invite size (+1 spare seat on calls of 3+)
- DiarizationService diarizes at a custom clustering threshold (0.70) with a
  cached second manager sharing one loaded copy of the models
- TranscriptionTaskManager.speakerSeparationProvider; the app sets it from
  SpeakerSeparationPreferences and the invite (MeetingSpeakerSeparation)
- Fresh 6-8 person calls: exactly right 0% -> 60%, missed people 27 -> 3,
  words under the right person 55% -> 78%

Recognize people sooner (beta):
- SpeakerNamingPolicy.InviteeBars: a voice whose best match is someone
  expected on the call is named silently after 2 confirmed meetings (0.80
  similarity, 0.10 margin) instead of 5 (0.92, 0.12); everyone else unchanged
- "Expected" is the calendar invite, or with no invite the 12 named people
  heard most recently, captured before the meeting's voices are matched
- TranscriptionTaskManager.lineupNamingProvider; the app sets it from
  CalendarNamingPreferences (MeetingCalendarNaming)
- Company series end to end: naming work -31% with invites, -37% with none,
  zero wrong names

Lab: merge_replay.py, cross-recording company (family X), holdout sets,
--separation / --calendar-naming / --no-invite in meeting-series.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@r3dbars r3dbars changed the title Speaker lab on YODAS3: simulated meetings through the real pipeline Speaker lab on YODAS3, plus better speaker separation and naming (beta toggles) Sep 28, 2026
r3dbars and others added 4 commits September 28, 2026 15:19
Isolated FluidAudio 0.17.4 probe (never linked into the app) plus a
fingerprint-free cleanup scorer. On the 45 holdout meetings, Nemotron 3 fast128
plus a 5 s fold credits 92% of words to the right person on 3-4 and 6-8 person
calls, vs 80/55% today and 80/78% for the PyAnnote-based separation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolved:
- DiarizationService: keeps both the Nemotron backend (#1789) and the lab's
  custom clustering threshold / speaker-bound seams. The threshold stays a
  cosine similarity and is converted to FluidAudio 0.17's cut distance.
- FluidAudioCompatibility.tunedOfflineDiarizerConfig() now reads the four
  hill-climb LabKnobOverrides knobs (cosine threshold converted), and
  config/hillclimb/knobs.json points the 15 diarizer knobs at their new lines.
- Harness command list, docs and .gitignore keep both sides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rom the saved model, lineup naming always on

- Nemotron 3 Diarization becomes the default backend. Its turns are embedded
  with the pyannote path's own offline WeSpeaker model in a real 10 s context
  window, so they match the people already in speakers.sqlite (lab: 0.99
  speaker-level cosine, 29/29 clusters top-1). If Nemotron can't load,
  pyannote stands in for that session.
- Separation cleanup always on, tuned per backend: Nemotron folds voices
  under 5 s and caps only one-on-one invites.
- Lineup naming always on (invite, or the 12 most recently heard people).
- The two beta toggles and their preferences are gone.
- Lab: harness --backend/--sep-* flags, stress families, AMI converter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@r3dbars r3dbars changed the title Speaker lab on YODAS3, plus better speaker separation and naming (beta toggles) Meetings: Nemotron 3 speaker separation + sooner naming, on by default (FluidAudio 0.17) Sep 28, 2026
@r3dbars
r3dbars marked this pull request as ready for review September 28, 2026 21:18
@r3dbars
r3dbars merged commit 10ee530 into main Sep 28, 2026
8 checks passed
@r3dbars
r3dbars deleted the claude/yodas-speaker-lab branch September 28, 2026 23:52
r3dbars added a commit that referenced this pull request Sep 29, 2026
diarization_engine was hardcoded to pyannote_offline, so every Nemotron
meeting since #1887 said pyannote. The pipeline now reads the diarizer's
activeRunDescriptor right before it diarizes (a Nemotron load fallback
reads pyannote) and the formatter writes nemotron_offline /
pyannote_offline / none, plus an optional voiceprint_model key. The
'+ PyAnnote pipeline' log line and the raw footer name the real backend.
CaptureKit and MCP surface voiceprint_model; older files read unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants