Skip to content

Send exact dictation timings, model, and Mac class to PostHog - #1856

Merged
r3dbars merged 3 commits into
mainfrom
claude/transcription-speed-percentiles-yyabey
Sep 25, 2026
Merged

r3dbars merged 3 commits into
mainfrom
claude/transcription-speed-percentiles-yyabey

Conversation

@r3dbars

@r3dbars r3dbars commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Requested by Justin · project thread

Why

Before: PostHog only gets dictation speed as coarse ranges (for example "100 to 249 ms"). It doesn't know which speech model ran or what kind of Mac it was, and the start timing stops when the mic call returns, not when real sound arrives. So there's no way to get P50/P95/P99 per model or per machine.

After: each dictation sends exact timings (rounded to 10 ms) for key press to recording, key press to first real sound, model run time, and stop to pasted text, along with the model and a coarse Mac class (chip family and memory size).

Product Impact

  • Affects: dictation
  • Lane: dictation reliability
  • Why this matters: Justin wants dictation to start as fast as possible. This measures that on every Mac and model, so each speed fix can be checked against real users.

What changed

  • dictation_started: adds start_latency_ms, stt_model, mac_chip, memory_gb_bucket.
  • dictation_start_requested: adds stt_model, mac_chip, memory_gb_bucket, so failed and refused starts can be split by model and Mac too.
  • dictation_stop_latency_measured: adds first_sound_latency_ms / first_sound_latency_bucket (key press to first audio buffer), decode_latency_ms, stop_to_paste_latency_ms, stt_model, mac_chip, memory_gb_bucket. The local event also gets press_to_first_sound_ms.
  • ParakeetEngine stamps when the first buffer arrives, under the existing pendingSamplesLock, on both the engine tap and the pinned mic path. A fresh start clears the stamp, and a recovery restart keeps it. The stop event only uses it if it falls between this session's key press and its stop, so a stale stamp or one from the shared meeting mic is dropped.
  • New Sources/Observability/MachineClassTelemetry.swift turns the CPU brand string into m2_pro-style chip names (anything else becomes unknown) and memory into a bucket. It's cached once per process.
  • Allowlist and reviewed-properties updated, the new file added to fast-test APP_SOURCES, and Tests/MachineClassTelemetryTests.swift added.

How I checked it

  • bash scripts/dev/linux-checks.sh (48 passed, 0 failed)
  • python3 scripts/dev/check-source-pins.py --changed-only (353 pins hold)
  • python3 scripts/dev/check-telemetry-keys.py (every new key survives the sanitizer)
  • python3 scripts/ops/normalize-analytics-taxonomy.py --check
  • bash build.sh --no-open / bash run-tests.sh / swift test (CI green on 493983e: app-build, build-and-test, spm-tests, checks, repo-hygiene)

Checks I could not run, and why:

  • No Swift toolchain in the cloud session. CI was the Swift build.

Mac or hardware test still needed? A light one on the next local build:

  1. Dictate a few times. events.jsonl dictation_stop_latency_measured should show press_to_first_sound_ms, in the same range as request_to_recording_ms plus about 80 ms. On main the field doesn't exist.
  2. In PostHog (local build channel), dictation_started should carry mac_chip (for example m3_max), memory_gb_bucket, stt_model, start_latency_ms.

Risk Review

  • Privacy / local-first behavior reviewed: only a chip family, a memory bucket, the model id, and rounded timings. No device names, serials, or model identifiers.
  • New analytics properties avoid the sanitizer's drop fragments (checked by check-telemetry-keys.py)
  • Checked the text-pin tests for every file I edited
  • Storage path or migration impact reviewed (none)
  • Public-facing copy (none)
  • Release/update impact reviewed (none)
  • Agent PRs got an independent deep review of the full diff (READY at 493983e, notes in the project's next-release review folder)
  • No private transcripts, audio, tokens, personal paths, or customer data are included

Notes

The real-time rule holds: the new stamp is one nil check and one store inside a lock the tap block already takes. There's no allocation, no I/O, and no new lock. The pinned mic's admission runs on its capture queue, not the IOProc.

🤖 Generated with Claude Code

https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM


Generated by Claude Code

Dictation speed only reached PostHog as coarse buckets, with no model and
no machine, so P50/P95/P99 per model or per kind of Mac couldn't be
computed. The start timing also stopped when the mic call returned, not
when real audio arrived.

- dictation_started: start_latency_ms (10 ms rounding), stt_model,
  mac_chip, memory_gb_bucket.
- dictation_stop_latency_measured: first_sound_latency_ms/_bucket (key
  press to the first audio buffer), decode_latency_ms,
  stop_to_paste_latency_ms, plus the same model and Mac fields.
- ParakeetEngine stamps the first buffer's arrival under the existing
  pendingSamplesLock on both the engine tap and the pinned mic path; a
  recovery restart keeps the original stamp.
- MachineClassTelemetry turns the CPU brand string into a chip family
  ("m2_pro") and physical memory into a bucket. No identifiers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM
@r3dbars r3dbars self-assigned this Sep 25, 2026
Deep review follow-ups on the PostHog speed fields:
- privacy-first-observability.md names the 10 ms dictation timings plus
  stt_model, mac_chip and memory_gb_bucket as the reviewed raw-number case,
  and the learning plan's dictation rows list the new keys.
- AnalyticsEventPolicyTests checks every new key survives the
  dictation_stop_latency_measured allowlist and sanitizer.
- stt_model comes from the recording's model lease when there is one, so a
  mid-dictation setting change reports the model that actually ran.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM
AnalyticsEventPolicyTests pins that dictation_start_requested carries
everything dictation_started does except start latency, so a funnel can
compare attempts to successes. The new stt_model, mac_chip and
memory_gb_bucket now ride on the request event too, and the test's
latency-only difference names both start_latency_bucket and
start_latency_ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM
@r3dbars
r3dbars marked this pull request as ready for review September 25, 2026 16:20
@r3dbars
r3dbars merged commit 6ebfdfd into main Sep 25, 2026
8 checks passed
@r3dbars
r3dbars deleted the claude/transcription-speed-percentiles-yyabey branch September 25, 2026 16:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants