Send exact dictation timings, model, and Mac class to PostHog - #1856
Merged
Merged
Conversation
Dictation speed only reached PostHog as coarse buckets, with no model and
no machine, so P50/P95/P99 per model or per kind of Mac couldn't be
computed. The start timing also stopped when the mic call returned, not
when real audio arrived.
- dictation_started: start_latency_ms (10 ms rounding), stt_model,
mac_chip, memory_gb_bucket.
- dictation_stop_latency_measured: first_sound_latency_ms/_bucket (key
press to the first audio buffer), decode_latency_ms,
stop_to_paste_latency_ms, plus the same model and Mac fields.
- ParakeetEngine stamps the first buffer's arrival under the existing
pendingSamplesLock on both the engine tap and the pinned mic path; a
recovery restart keeps the original stamp.
- MachineClassTelemetry turns the CPU brand string into a chip family
("m2_pro") and physical memory into a bucket. No identifiers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM
Deep review follow-ups on the PostHog speed fields: - privacy-first-observability.md names the 10 ms dictation timings plus stt_model, mac_chip and memory_gb_bucket as the reviewed raw-number case, and the learning plan's dictation rows list the new keys. - AnalyticsEventPolicyTests checks every new key survives the dictation_stop_latency_measured allowlist and sanitizer. - stt_model comes from the recording's model lease when there is one, so a mid-dictation setting change reports the model that actually ran. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM
AnalyticsEventPolicyTests pins that dictation_start_requested carries everything dictation_started does except start latency, so a funnel can compare attempts to successes. The new stt_model, mac_chip and memory_gb_bucket now ride on the request event too, and the test's latency-only difference names both start_latency_bucket and start_latency_ms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Requested by Justin · project thread
Why
Before: PostHog only gets dictation speed as coarse ranges (for example "100 to 249 ms"). It doesn't know which speech model ran or what kind of Mac it was, and the start timing stops when the mic call returns, not when real sound arrives. So there's no way to get P50/P95/P99 per model or per machine.
After: each dictation sends exact timings (rounded to 10 ms) for key press to recording, key press to first real sound, model run time, and stop to pasted text, along with the model and a coarse Mac class (chip family and memory size).
Product Impact
dictationdictation reliabilityWhat changed
dictation_started: addsstart_latency_ms,stt_model,mac_chip,memory_gb_bucket.dictation_start_requested: addsstt_model,mac_chip,memory_gb_bucket, so failed and refused starts can be split by model and Mac too.dictation_stop_latency_measured: addsfirst_sound_latency_ms/first_sound_latency_bucket(key press to first audio buffer),decode_latency_ms,stop_to_paste_latency_ms,stt_model,mac_chip,memory_gb_bucket. The local event also getspress_to_first_sound_ms.ParakeetEnginestamps when the first buffer arrives, under the existingpendingSamplesLock, on both the engine tap and the pinned mic path. A fresh start clears the stamp, and a recovery restart keeps it. The stop event only uses it if it falls between this session's key press and its stop, so a stale stamp or one from the shared meeting mic is dropped.Sources/Observability/MachineClassTelemetry.swiftturns the CPU brand string intom2_pro-style chip names (anything else becomesunknown) and memory into a bucket. It's cached once per process.APP_SOURCES, andTests/MachineClassTelemetryTests.swiftadded.How I checked it
bash scripts/dev/linux-checks.sh(48 passed, 0 failed)python3 scripts/dev/check-source-pins.py --changed-only(353 pins hold)python3 scripts/dev/check-telemetry-keys.py(every new key survives the sanitizer)python3 scripts/ops/normalize-analytics-taxonomy.py --checkbash build.sh --no-open/bash run-tests.sh/swift test(CI green on 493983e: app-build, build-and-test, spm-tests, checks, repo-hygiene)Checks I could not run, and why:
Mac or hardware test still needed? A light one on the next local build:
events.jsonldictation_stop_latency_measuredshould showpress_to_first_sound_ms, in the same range asrequest_to_recording_msplus about 80 ms. Onmainthe field doesn't exist.dictation_startedshould carrymac_chip(for examplem3_max),memory_gb_bucket,stt_model,start_latency_ms.Risk Review
check-telemetry-keys.py)Notes
The real-time rule holds: the new stamp is one nil check and one store inside a lock the tap block already takes. There's no allocation, no I/O, and no new lock. The pinned mic's admission runs on its capture queue, not the IOProc.
🤖 Generated with Claude Code
https://claude.ai/code/session_016DYGa1i8HDv497ewpWgaCM
Generated by Claude Code