Skip to content

Release integration: all open PRs + pre-release review fixes - #1904

Merged
r3dbars merged 64 commits into
mainfrom
claude/release-integration-1167
Sep 28, 2026
Merged

r3dbars merged 64 commits into
mainfrom
claude/release-integration-1167

Conversation

@r3dbars

@r3dbars r3dbars commented Sep 28, 2026

Copy link
Copy Markdown
Owner

Merges every open PR into one tested build for the next release, plus the fixes from a pre-merge review pass. Merging this with a merge commit marks each PR below as merged.

Included PRs

Fixes found in review (5 Fable 5.1 deep reviews + 3 area test plans + 1 independent review of the fixes)

  • Island never records mic-only silently. First time: macOS allow box. After a denial: records the mic and asks about call audio every meeting; nothing remembered as a choice. Ask lifecycle owned by the session (MeetingCallAudioAsk).
  • 1:1 speaker cap counts people, not readable names (email-only guests no longer collapse a group call into one voice).
  • No-invite lineup: a lone match with no runner-up needs the standard 0.92 bar before silent naming. Invite lineups keep the lab bars.
  • Imports never use a calendar invite (no cap, no lowered bars).
  • Island always shows who was on the call (island mode only), hover "Not X?" to correct, saved as a correction on the profile lifeline. Recognized-only reviews don't block retries/updates and close after ~2 min.
  • Island name box: Return/Tab saves what you typed or arrowed to (no more "Chris" → Christina merges); Done/Later keep a typed name; Later ring only runs while visible.
  • Speakers card: naming a queued voice after a saved person merges into them (one mutation-service op with confirmations + rollback) instead of making a duplicate.
  • Writing: llama-server gets a per-launch API key (env, not argv) and --no-webui; every caller sends the bearer header. Closes a localhost prompt-probing hole inherited from Tilde.
  • Dictation: a hardware-muted mic's zero-only takes don't demote its speed path; an unconfirmed paste reads "Maybe pasted" in the island.
  • Release: build-beta.sh bundles the Nemotron fast128 diarizer (~190 MB) so the first meeting after an update doesn't download it inside the 120 s model budget; OpenMDW-1.1 notice added.

Checks (local, this Mac)

  • build.sh --no-open exit 0; swift test all targets pass.
  • run-tests.sh: everything passes except MeetingPromptDetectorTests, which fails only under heavy machine load (load avg 60–150 from parallel sessions); it passed (19,873/19,873) on an earlier head of this branch at lower load, and no change here touches the detector. Letting CI's clean runner decide.
  • check-source-pins, check-test-shape, check-telemetry-keys, check-doc-paths: pass. check.sh quick: VM-script guard test is flaky (passes on rerun).

Known, deferred (tracked separately)

Test checklist: https://claude.ai/artifact/R57o1UEghf3arPmCVtCCAT

🤖 Generated with Claude Code

0.17.0 adds Nemotron 3 Diarization, which the speaker lab needs. The bump
itself should change nothing users see, so:

- build-deps: tools 6.2 manifest so FluidAudio can opt out of its default
  NemoTextProcessing trait (a prebuilt Rust static lib we never call and
  would not archive). swiftLanguageModes [.v5] keeps in-tree targets on
  the language mode they built under before.
- DiarizationService: clusteringThreshold is now a Euclidean cut distance,
  so the tuned 0.6 cosine becomes sqrt(0.8); constrainedAssignment pinned
  off to match 0.15.x assignment.
- Parakeet (app + CLI): keep the 0.15.x long-form chunking
  (melChunkContext on, no seam-gap repair).
- Integration fake FluidAudio grows the ASRConfig shape.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
DiarizationBackendPreferences reads TRANSCRIPTED_DIARIZATION_BACKEND, then
the diarization-backend-preference default, then falls back to pyannote.
No Settings UI. Not wired into meetings yet; that lands with the Core
Nemotron backend.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… ci]

DiarizationBackend (.pyannote default, .nemotron experimental) picks the
model. The pyannote path is unchanged. The Nemotron path loads a FluidAudio
Nemotron3 preset (fast128 by default, TRANSCRIPTED_NEMOTRON_PRESET for the
lab), runs it on a private serial queue, turns frame probabilities into
exclusive speaker turns with the pure NemotronTurnBuilder, and embeds each
turn with the injected embedder or a new FluidWeSpeakerSegmentEmbedder,
since Nemotron emits no voiceprints.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
NemotronTurnBuilder: exclusivity, argmax on overlap, ties, gap bridging,
min-duration drop and blip rejoin, first-appearance remapping, empty and
malformed input. DiarizationService: pyannote default, backend round-trip,
not-loaded error, thresholds, shared reembed path. Nemotron preset override
and the WeSpeaker fallback embedder's pure helpers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
MeetingSessionController passes the chosen backend to DiarizationService.
With Nemotron and no ERes2Net, voiceprints come from FluidAudio's online
WeSpeaker model, so they go to their own speakers_wespeaker-fluid-online
database (meetings and Settings > People agree) until the lab proves they
match the offline pipeline's vectors. Default pyannote is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… thresholds and fingerprint-update knobs [skip ci]

dump takes --backend pyannote|nemotron and --embedder native|eres2net
(--eres2net-model for a path-based model) and records backend, embedder,
dimension, wall time, and audio length. Old dumps still decode.

replay takes --thresholds auto|weSpeaker|eRes2Net, --match adaptive,
--same-voice, --dedup, and the write-back EMA knobs. Defaults reproduce
the old behavior for WeSpeaker dumps. Per meeting it now reports raw
cluster count and whether each cluster matched an existing profile.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… [skip ci]

speaker_eval_common.py now holds RTTM parsing, overlap helpers, the
identity metrics (fragmentation, false merge, re-ID curve), and a pure
Python DER/JER that matches pyannote.metrics to float precision
(collar convention, UEM, overlap, Hungarian mapping). It is about 40x
faster, which matters once the lab scores many replays.

score_speaker_eval.py uses it and no longer needs pyannote installed.
Its markdown and JSON output are unchanged on a synthetic fixture.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…ng-speaker recognition [skip ci]

score_speaker_lab.py scores a lab run directory into scores.json and
REPORT.md: per variant, raw diarizer DER/JER and count error, pipeline
DER after clustering and DB matching, and the recognition scoreboard
(recognized / wrong person / asked again / undetected, plus new people
false-matched to a known profile). It picks each variant's best knob
setting and shows which knob values moved recognition. Own-calls mode
reports behavior without ground truth, agreement against a baseline
variant, and writes a self-contained timeline.html.

Also has helper subcommands the driver uses: grid, dump-ok, and
own-calls-list. Unit tests run on Linux with synthetic fixtures.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…ingle trial [skip ci]

run_speaker_lab.sh dumps each meeting once per variant (backend x
embedder x Nemotron preset) into its own cache, replays every knob
setting, and scores it all into reports/speaker-lab/<stamp>/. Every
knob is a flag with an env twin, the run is non-interactive, fails loud,
and prints the scores.json path as its last stdout line. --single runs
one variant at one setting for an outer optimizer. --own-calls runs the
same variants on saved meetings (call track read in place) and writes a
timeline page.

Tests drive the script end to end against a fake harness, so the
orchestration is covered on Linux too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…s [skip ci]

16 scenario series (8 Edinburgh, 4 Idiap, 4 TNO), 4 sessions each with
the same 4 people, for the speaker lab's cross-call recognition test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…[skip ci]

README covers what the lab measures, how to run it on AMI and on your
own saved calls, every knob, how an optimizer drives --single trials,
the scores.json schema, and how to add a diarizer or embedder. The
test matrix now runs the lab's syntax checks and unit tests when the
lab files change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
….17 [skip ci]

Review of #1789 found the bump was not as behavior-neutral as claimed:

- FluidAudio 0.17 pins speaker-diarization-coreml to one commit and
  deletes any cache without a matching revision marker. Every 0.15.x
  cache, and the offline-diarizer-models copy bundled in the app, has
  none, so the first meeting would delete inside the signed bundle (or
  fail offline). FluidAudioCompatibility.keepUnpinnedDiarizerCaches()
  resolves that repo at main, as 0.15.x did; DiarizationService,
  FluidWeSpeakerSegmentEmbedder and the CLI call it before loading.
- The tuned pyannote config moves to
  FluidAudioCompatibility.tunedOfflineDiarizerConfig() and a test pins
  the distance threshold, constrainedAssignment=false and every tuned
  value.
- CLI diarize/batch start from the 0.15.x default (cosine 0.6, no
  constrained assignment), and a config file's clusteringThreshold stays
  a cosine similarity, converted for 0.17.
- Parakeet comments now say only the two new defaults are pinned: 0.17
  also changed >15 s chunk merging with no switch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… [skip ci]

- THIRD_PARTY_LICENSES: FluidAudio 0.17's Japanese G2P (Misaki ports,
  UniDic) and Spanish/French lexicon notices; NeMo text processing is not
  linked.
- Parakeet comment: 0.17 retries blank v3 decodes up to five more times,
  so near-silent dictation can take longer and return text.
- build-deps: FluidAudio's resource bundle never ships, so LuxTts G2p must
  stay unused.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…and valueless flags [skip ci]

A dump with an unknown TRANSCRIPTED_NEMOTRON_PRESET recorded the typo while
Core's runner silently ran fast128, so the cached dump lied about its variant.
Validate with FluidAudio's Nemotron3Config.preset(named:) and record 'default'
for unset. --write-path-fixes now rejects anything but on|off (a typo used to
run the legacy path), and a flag with no value is an error instead of the default.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…the online WeSpeaker embedder [skip ci]

Diarizes one file with today's pipeline, re-embeds the same segments with
FluidWeSpeakerSegmentEmbedder (what Nemotron falls back to), and reports
per-segment, within-model, cross-model and per-cluster cosine stats against the
WeSpeaker match floor, plus a looksInterchangeable heuristic. Answers whether
Nemotron voiceprints could share speakers.sqlite instead of a separate DB.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Add DiarizationService.resolvedNemotronPresetName(environment:), a public
wrapper over the runner's internal resolver, so the harness records 'fast128'
instead of 'default' when TRANSCRIPTED_NEMOTRON_PRESET is unset, and refuses a
name Core would silently replace. The lab scorer treats unset, 'default' (older
dumps) and 'fast128' as one variant so existing caches stay reusable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
scripts/hillclimb/benches/speaker_lab.py speaks the hill-climb request/result
protocol (#1791) and drives run_speaker_lab.sh --single once per trial over
the requested AMI series, then splits scores.json + recognition-events.json
back into per-series items: recognition_rate, recognized, asked_again,
pipeline/raw DER, speaker-count error, objective, plus int gates
wrong_person and new_person_false_match. Missing/incomplete series and
series without returning speakers are item errors, never zeros. Driver env
twins are scrubbed so shell exports can't leak into a trial, and scores.json
must echo every knob that was set. app_revision hashes the harness binary,
the requested RTTM/audio, and the lab scripts.

Imports hc_benches when the climber is on the branch, else uses a local copy
of the protocol validator, so it works before and after the merge.

Also: config/hillclimb/suites/speaker-lab-ami.json (the 16 download_ami.sh
lab series, 12 dev / 4 holdout stratified by site) and
speaker_lab.README.md with the benches/knobs/objectives JSON for #1791.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Changes to the adapter, its tests/README/suite, or the lab driver/scorer it runs now select py_compile + speaker_lab.py --self-test. Kept as its own rule above the SpeakerEvalHarness block so it doesn't collide with #1791's scripts/hillclimb/** rule.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…eports it [skip ci]

run_speaker_lab.sh --embedding-parity (EMBEDDING_PARITY=1) runs the harness's
embedding-parity per meeting (RTTM labels on corpora, pyannote clusters on own
calls), cached in data/eval/<corpus>/parity/. The scorer pools the reports via
their 200-bin histograms into an additive embeddingParity block in scores.json
and an 'Embedding parity' section in REPORT.md, with the same interchangeability
rule as the harness. schemaVersion stays 1; runs without the flag are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
A cache written by a pinned FluidAudio 0.17 build carries a
.fluidaudio-revision marker naming a commit. The app resolves the
diarizer repo at main, so a shipped marker naming anything else would make
it delete files inside its own signed bundle. build-beta drops the marker
after copying the cache.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…kip ci]

Rebased onto e3a559f. The dump now records the resolved Nemotron preset and the scorer treats unset/default/fast128 as one variant, so the knob echo check normalizes presets the same way. fast128 is still not passed because the driver names variants after the preset string (passing it would fork the dump cache). EMBEDDING_PARITY joins the scrubbed env twins; --embedding-parity is a diagnostic and never passed per trial.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
… ci]

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…ip ci]

The 0.17 bump can't merge until WER is no worse and blank-audio dictation
stops aren't slower (PR #909 retries blank V3 decodes up to 5 times).

scripts/stt_fluidaudio_ab.sh builds transcripted-cli twice in git worktrees
under ~/stt-fluidaudio-ab: the baseline from origin/main (pre-bump source,
since 0.15.4 can't compile the new ASRConfig call) with FLUID_AUDIO_VERSION
0.15.4, the candidate from HEAD with 0.17.0. Builds are reused while their
inputs are unchanged. It refuses a baseline ref that already uses the 0.17
API and checks the version SwiftPM actually resolved.

scripts/stt_fluidaudio_ab.py then measures, reusing the STT shootout's
lecture download, caption parsing and WER scorer (read from git when the
shootout isn't in the checkout):
- WER: first 20 min of the lecture plus 10 caption-aligned ~45 s pieces
- stop time: silence 1/3/8 s, near-silence, low noise, room rumble, a
  0.5 s noise burst (gated), plus `say` clips (reported), ABBA rounds,
  median of CLI processingSeconds
Gate flags: --max-wer-delta-pp 0.5, --max-time-ratio 1.25,
--max-time-delta-ms 300. Writes result.json + report.md; last stdout line
is the result.json path.

scripts/test_stt_fluidaudio_ab.py covers the math, fixtures and a fake-CLI
end-to-end run, and drives the bash wrapper against a throwaway repo with
stubbed Mac tools. Runs on Linux.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
…t groups [skip ci]

The hill-climb lab needs >= 10 dev and >= 8 holdout independent units before an
objective can confirm anything; 16 series (12/4) left speaker-lab-recognition BLOCKED.
Adds ES2004 ES2011 ES2012 ES2013 IS1006 IS1007 TS3007 TS3008 (all four sessions have
pyannote RTTMs, and no participant appears in two series). Holdout re-picked per site by
lowest unit_hash under a new salt (speaker-lab-ami-v2): 4 ES, 2 IS, 2 TS.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
Resolves the Core and SpeakerEvalHarness CLAUDE.md file maps: keeps main's
per-file map and adds this branch's diarization backend, FluidAudio
compatibility and speaker lab entries.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SvYQfvK9JYWs3DB2fNkVNQ
r3dbars and others added 28 commits September 28, 2026 15:11
Both default off. Tuned in the YODAS3 speaker lab and checked on 45 fresh
meetings the tuning never saw (Tools/SpeakerEvalHarness/YODAS_LAB_RESULTS.md).

Separate voices on calls (beta):
- SpeakerSeparation.swift: after the diarizer, fold voices with under 5 s of
  talk into the voice they sound most like, merge fingerprints at 0.6 and up,
  and cap at the calendar invite size (+1 spare seat on calls of 3+)
- DiarizationService diarizes at a custom clustering threshold (0.70) with a
  cached second manager sharing one loaded copy of the models
- TranscriptionTaskManager.speakerSeparationProvider; the app sets it from
  SpeakerSeparationPreferences and the invite (MeetingSpeakerSeparation)
- Fresh 6-8 person calls: exactly right 0% -> 60%, missed people 27 -> 3,
  words under the right person 55% -> 78%

Recognize people sooner (beta):
- SpeakerNamingPolicy.InviteeBars: a voice whose best match is someone
  expected on the call is named silently after 2 confirmed meetings (0.80
  similarity, 0.10 margin) instead of 5 (0.92, 0.12); everyone else unchanged
- "Expected" is the calendar invite, or with no invite the 12 named people
  heard most recently, captured before the meeting's voices are matched
- TranscriptionTaskManager.lineupNamingProvider; the app sets it from
  CalendarNamingPreferences (MeetingCalendarNaming)
- Company series end to end: naming work -31% with invites, -37% with none,
  zero wrong names

Lab: merge_replay.py, cross-recording company (family X), holdout sets,
--separation / --calendar-naming / --no-invite in meeting-series.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on Not now, Later

Right-clicking the status item now opens the same popover as a left-click.
The separate right-click menu is gone (owner request); AGENTS.md's
keep-the-surface list is updated to match.

In the Notch island, the call-detected prompt drops the "Closes in Ns"
line and traces the countdown as a ring around Not now, the same ring the
dictation Dismiss button uses. Remind me soon reads Later. The prompt's
timeout pauses with the ring while the pointer is over the island, so it
never closes on someone reading it. An unanswered prompt still expires and
re-offers as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Isolated FluidAudio 0.17.4 probe (never linked into the app) plus a
fingerprint-free cleanup scorer. On the 45 holdout meetings, Nemotron 3 fast128
plus a 5 s fold credits 92% of words to the right person on 3-4 and 6-8 person
calls, vs 80/55% today and 80/78% for the PyAnnote-based separation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolved:
- DiarizationService: keeps both the Nemotron backend (#1789) and the lab's
  custom clustering threshold / speaker-bound seams. The threshold stays a
  cosine similarity and is converted to FluidAudio 0.17's cut distance.
- FluidAudioCompatibility.tunedOfflineDiarizerConfig() now reads the four
  hill-climb LabKnobOverrides knobs (cosine threshold converted), and
  config/hillclimb/knobs.json points the 15 diarizer knobs at their new lines.
- Harness command list, docs and .gitignore keep both sides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Phase 2 slice. DictationTrigger was nested inside DictationSessionController,
so the fast tests (which can't compile the controller) checked its raw values
by reading the controller as text. It's now top-level in
Sources/UI/Overlay/DictationTrigger.swift, unchanged. Uses inside the
controller already said plain DictationTrigger; the one outside use
(ContextCaptureEngine.routeDictationToggle) drops the qualifier.

DictationStartReadinessTests now checks hotkeyTriggerRawValues against
DictationTrigger(rawValue:) directly. Source-text baseline for that file 10 -> 8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ger-move

# Conflicts:
#	Sources/UI/CLAUDE.md
#	scripts/entrypoints/run-tests.sh
… call audio ask

When a saved meeting has voices Transcripted isn't sure about, the island
now asks "Who was on this call?" instead of opening the review window
(Notch island mode):

- voices it named on its own show as recognized; the pipeline now passes
  those names in SpeakerNamingRequest.recognizedSpeakerNames
- a likely match asks "Is this Maya?" with Yes / No
- No, or a voice with no guess, opens a name box with the calendar
  invitees as one-tap names (an arrow shows past three) and autocomplete
  from people already saved, invitees first
- the play button turns into pause while a clip plays, with a ring that
  fills until the clip ends
- Later carries the dictation Dismiss ring (20 s, pauses on hover, stops
  once you touch anything) and saves what was answered; Done saves and
  shows "Everyone's named" with Open transcript

Answers build the same SpeakerNameUpdates the window builds, and the
review reports the same analytics with surface speaker_review_island.
The island panel can take the keyboard only while a name box is in use.

Speakers: "Needs a name" becomes "Name these people", one card per call
(name, day, length) with that call's invitees as one-tap names and a saved
"Skip this call" (the per-voice Skip stays session-only).

Call audio off: in Notch island mode a meeting no longer waits on the
"can't hear the other side" alert. It starts with the mic, and the island
asks once while it records; turning it on applies from the next meeting,
since this recording never built the system-audio tap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a fixed 100 ms

The suite waited a fixed ~100 ms for MeetingPromptDetector's async evaluation
before asserting a prompt arrived, so a loaded machine could assert first:
seen as 'a real Meet tab right after a Not now ... should still prompt —
expected 2, got 0' in a full run, passing on rerun. The two waits that expect
a prompt now wait until it arrives (up to about 5 s); the final 'no new prompt'
check keeps the fixed wait, which a slow runner can't turn red.

Other suites in the file share the fixed helper and can flake the same way
under heavy load; fixing them for good needs a pending-evaluation hook in the
detector, noted in the testing plan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… is a card stack

The island's play/pause is now a fixed 30 pt round view (the NSButton
wasn't staying square, so it drew as a squashed blob with an off-center
ring). Answering a voice (Yes, or picking a name) stops its clip so the
next one is ready to play.

Speakers shows "Review and name these people" as a stack: only the top
call is open, the next ones peek out underneath, and naming everyone in
it brings the next call up. Later sends a call to the back for now; Skip
this call still removes it for good.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rom the saved model, lineup naming always on

- Nemotron 3 Diarization becomes the default backend. Its turns are embedded
  with the pyannote path's own offline WeSpeaker model in a real 10 s context
  window, so they match the people already in speakers.sqlite (lab: 0.99
  speaker-level cosine, 29/29 clusters top-1). If Nemotron can't load,
  pyannote stands in for that session.
- Separation cleanup always on, tuned per backend: Nemotron folds voices
  under 5 s and caps only one-on-one invites.
- Lineup naming always on (invite, or the 12 most recently heard people).
- The two beta toggles and their preferences are gone.
- Lab: harness --backend/--sep-* flags, stress families, AMI converter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cap's sleep/countdown loop moves out of installSessionTimeout into
DictationSessionCapTimer with an injected clock, sleep and countdown hook.
The controller still starts the deadline before the task runs and still
decides paste vs copy at the cap. Fake-clock tests cover the long sleep,
the 1 s countdown ticks, the single VoiceOver announcement, cancel, and a
late-starting task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Island: never remember mic only. First time shows the macOS allow box;
  after a denial each meeting records the mic and the island asks about
  call audio while it records. The ask no longer gets cleared by the
  previous meeting's status.
- 1:1 speaker cap counts people on the invite, not readable names, so
  email-only guests don't collapse a group call into one voice.
- No-invite lineup: a lone match (no runner-up) must clear the standard
  0.92 bar before naming silently. Invite lineups keep the lab bars.
- Imports never pick up a calendar invite (no cap, no lowered bars).
- Speakers card: naming a queued voice after one saved person (invitee
  chip or typed) merges into them instead of making a duplicate.
- A hardware-muted mic's zero-only takes don't demote its speed path.
- Island says "Maybe pasted" for an unconfirmed paste so the Paste
  button doesn't read as the fix and double the text.
The helper listened on 127.0.0.1:17891 with no key, CORS * and the web UI
on. Loopback TCP isn't per-user, and /completion returns timings.cache_n,
so with cache_prompt any local process or a web page could prefix-probe
the last prompt (typed text plus Screen Memory OCR text).

The host now makes a fresh 32-byte SecRandomCopyBytes key for every
launch, hands it over as LLAMA_API_KEY in the child's environment (not
argv, so ps can't show it) and adds --no-webui. Completions, scaffold
prewarm and both /health probes send Authorization: Bearer <key>. The key
lives only in memory. Checked against the pinned binary (2115b73):
LLAMA_API_KEY is honored, /completion answers 401 without it, /health
stays public, / is 404 with --no-webui.

Deliberate divergence from Tilde, recorded in the port ledger and plan.
…picks

- After every meeting with a remote voice, the island lists who was on the
  call, even when everyone was recognized ("On <meeting>", Done + ring, no
  questions). Hovering a recognized name offers "Not Taylor?", which opens
  the name box; the correction saves through the naming coordinator as a
  correction of the recognized person (match undone, dispute + lifeline
  outcome). The pipeline now extracts clips for silently recognized remote
  voices and passes them as SpeakerNamingRequest.recognizedSpeakers; only
  corrected ones join the save. The review window skips an all-recognized
  request.
- Return/Tab saves the arrowed-to row, else an exact match, else the typed
  name as a new person. Default highlight is the exact match or the new
  person row, and the arrows reach it. "Chris" no longer turns into
  "Christina".
- Done and Later keep a name typed but not submitted, read like Return.
- The Later ring only runs while the review is on screen and not hovered.
- Speakers card: naming a queued voice after a saved person now merges
  through one mutation-service op (mergeReviewedVoice) that names each
  queued row user_manual, records a merged confirmation per reviewed
  meeting for the kept person, and rolls transcripts back on failure.
- Recognized-only reviews: only queued (and clips only cut) when the
  Notch island lists them; they no longer count as a pending review for
  failed-meeting retry or update installs; the island closes one after
  ~2 min even while hidden.
- Cleanup of a pending review also removes recognized-voice clips.
- Island call-audio ask is owned by the session and tied to the start
  attempt: raised from the start outcome (so a first-time macOS Don't
  Allow asks too) and dropped when a start ends without recording.
- "Not X?" match telemetry now finds recognized voices.
The first meeting after an update no longer downloads ~190 MB inside the
120 s model-wait budget. build-beta.sh copies the fast128 preset and its
silence embedding into Resources/nemotron-diarizer-models, which
NemotronDiarizationRunner already loads first. Adds the OpenMDW-1.1
notice the weights require.
@r3dbars
r3dbars merged commit 00eb729 into main Sep 28, 2026
8 checks passed
@r3dbars
r3dbars deleted the claude/release-integration-1167 branch September 28, 2026 23:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants