Skip to content

Speakers: release held people on confirm, floor the separation folds - #1928

Merged
r3dbars merged 1 commit into
mainfrom
claude/speaker-release-and-fold-floor
Sep 29, 2026
Merged

r3dbars merged 1 commit into
mainfrom
claude/speaker-release-and-fold-floor

Conversation

@r3dbars

@r3dbars r3dbars commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Four speaker-naming bugs from code review. Three of them showed up on real hardware.

1. Held people never released during a session

Bug: releaseHeldVoiceprintConfirmations() only ran at the start of SpeakerVoiceprintMigration.run(), once per launch. A review confirmation wrote speaker_profile_confirmations but never touched the ledger, so a held person needed a manual confirm in every meeting until the next launch. On hardware, a colleague from 216 calls needed one every time.
Fix: The release rules moved into releaseHeldVoiceprintConfirmationsImpl(holders:). recordUserConfirmationsImpl calls it after its inserts, in the same transaction, for the profiles it just confirmed. The rules themselves didn't change: follow merges to the holder, the holder must exist and have a confirmation, then carry the held confirmations over and mark the row released. Only the timing changed. The pass is skipped when the database has no migration ledger. The launch run still does its full pass as a safety net. A held row that fails to decode is now logged as a count only (unreadable_count), with no names.
Tests (SpeakerVoiceprintMigrationTests):

  • testAudioThatDisagreesKeepsTheNameButWaitsForOneConfirmation now expects the release right after the confirmation, and expects the next run to release nobody.
  • New: testConfirmingSomeoneElseReleasesNobody
  • New: testConfirmingThePersonAHeldProfileWasMergedIntoReleasesIt
  • New: testAnUnreadableHeldRowStaysHeldAndDoesNotBlockTheConfirmation

2. Tiny voices folded with no similarity floor

Bug: Step 1 of SpeakerSeparation.apply folded every voice under 5 s into a plain argmax, with no minimum similarity. A voice with no fingerprint went to whoever talked most. On hardware, a real 3.4 s line from a different person got someone else's label.
Fix: New SpeakerSeparationOptions.foldSimilarity. Step 1 now folds only when the cosine to the target is at least that bar, and it never folds a voice that has no fingerprint. Both tuned presets set the bar to the active voiceprint model's microAbsorb: 0.62 for WeSpeaker, 0.626 for ReDimNet2, 0.45 for ERes2Net. I picked microAbsorb because it's the clusterer's existing bar for "absorb a very short cluster (<10 s) into another". That's the same decision as this step, and it's calibrated per model.
Tests (SpeakerSeparationTests):

  • testANearSilentVoiceWithNoFingerprintStaysItsOwnSpeaker replaces the old test that folded it into the loudest voice.
  • New: testAShortLineFromSomeoneElseKeepsItsOwnSpeaker (the 3.4 s hardware case)
  • New: testAShortPieceOfSomeoneWhoTalkedMoreStillFolds
  • New: testAShortVoiceFoldsOnlyWhenItClearsTheFoldBar

3. A 1:1 invite capped the call side to one voice

Bug: nemotronTuned sets maxSpeakers: 1 for a one-person invite. The cap step then folded every other voice into the loudest one with no similarity check. So a recurring 1:1 slot that held a different call, or a third person joining, collapsed everyone into one named speaker.
Fix: New capFoldSimilarity, set to the model's separationMerge (0.6 WeSpeaker, 0.598 ReDimNet2). The cap step folds the quietest voice that clears the bar. A voice that clears nobody stays, even if that leaves more voices than the cap. With no bar set, the old behavior is unchanged.
Tests:

  • New: testAOneOnOneCapNeverCollapsesADifferentPersonIntoTheInvitee
  • New: testAOneOnOneWithOnlyTheInviteeStillEndsWithOneVoice (the case the cap exists for still works)
  • New: testACapWithABarKeepsAVoiceWithNoFingerprint
  • New: testFoldAndCapBarsComeFromTheActiveVoiceprintModel
  • Preset assertions extended.

4. Nemotron fallback kept Nemotron options

Bug: speakerSeparationProvider captured diarizationBackend when the app started. If Nemotron failed to load, pyannote ran with Nemotron-tuned separation.
Fix: A new MeetingSpeakerSeparationProvider.make builds the provider. It reads diarizer.activeBackend on every call, and the pipeline makes that call after the models load. The MeetingSessionController edit is minimal, and check-source-pins --changed-only passes.
Tests: new fast test Tests/MeetingSpeakerSeparationProviderTests.swift. It flips the active backend after the provider is built and checks that the next meeting gets the new backend and its recording date. The new file is added to the run-tests.sh source list.

Judgment call to review

  • The pyannote cap now has the same bar. labTuned also sets foldSimilarity = microAbsorb and capFoldSimilarity = separationMerge. This matters because after fix 4, pyannote is exactly what a Nemotron fallback runs, and its 1:1 cap would otherwise collapse people the same way. The cost: pyannote's step 2 already merges everything at separationMerge, so its invite cap is now close to a no-op. The YODAS lab credited that cap with part of pyannote's 1:1 gain. Pyannote is only the fallback now. An extra row takes one click to merge in review; two people under one name can't be split. If you'd rather keep pyannote's cap as it was, drop capFoldSimilarity from labTuned.

Checks

  • bash build-deps.sh --force, then bash build.sh --no-open: builds
  • swift test --filter SpeakerTests: 522 tests, 0 failures (14 skipped)
  • bash run-tests.sh --filter MeetingSpeakerSeparationProviderTests: 4/4
  • python3 scripts/dev/check-source-pins.py --changed-only: pass
  • bash check.sh: pass (15/15)

Still needed before merge: an independent review of the full diff (per AGENTS.md), and a hardware re-check of the 1:1 and 3.4 s cases.

🤖 Generated with Claude Code

- A review confirmation now releases a held voiceprint-migration person in
  the same transaction, instead of waiting for the next launch's run.
  Unreadable held ledger rows are logged as a count.
- Separation step 1 folds a short voice only when it clears the model's
  microAbsorb bar; a voice with no fingerprint is never folded.
- The invite cap folds a voice only when it clears separationMerge, so a
  1:1 invite can't collapse a third person into the invitee.
- The separation provider reads the diarizer's active backend per meeting,
  so a Nemotron load failure gets pyannote's settings.
@r3dbars

r3dbars commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Independent review (coordinator session, full Sources diff vs main): APPROVE.

  • Release-on-confirm reuses the launch release rules (merge-survivor walk, profile exists, confirmation count > 0), scoped to the confirmed holders, inside the confirmation transaction; launch pass stays as backstop. Ledger-absent DBs skip cleanly.
  • Fold floor = model microAbsorb, cap floor = separationMerge; unfingerprinted voices no longer join the loudest. Matches the work-Mac mislabel (3.4 s, 1 segment) and the 1:1-invite collapse.
  • Provider reads the active backend per meeting, so a Nemotron load failure gets pyannote options.
  • Owner call on labTuned cap floor: keep it (conservative: an extra row beats two people under one name).
    Merge on green CI; hardware recheck with a 1:1 replay on the release build.

r3dbars added a commit that referenced this pull request Sep 29, 2026
@r3dbars
r3dbars merged commit 821e93b into main Sep 29, 2026
4 of 8 checks passed
@r3dbars
r3dbars deleted the claude/speaker-release-and-fold-floor branch September 29, 2026 21:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant