Skip to content

docs(eval): per-family accuracy (all 6 families, real gct) - #492

Merged
zzylol merged 1 commit into
mainfrom
docs/perfamily-accuracy
Jun 12, 2026
Merged

zzylol merged 1 commit into
mainfrom
docs/perfamily-accuracy

Conversation

@zzylol

@zzylol zzylol commented Jun 12, 2026

Copy link
Copy Markdown
Contributor

Folds the multi-sketch per-family accuracy into Fig 3(c), on real gct (1000 series): Sum exact, DDSketch/KLL quantiles in-envelope, CMS exact freq, HLL per-series 3e-5 — 5/6 families validated. DDSketch-vs-KLL head-to-head (KLL better median; DDSketch tighter tail, −26% wire). Two named gaps: CountSketch topk recall 0 (keys by item-not-host, ranks by occurrence-not-value → value-weighted topk needs a separate path) and HLL global rollup. Notes wall-clock anchoring fixed the warm read (timing, not reducer) + the warm-vs-archive range-routing finding. Data: datasets_eval/multisketch/ (feat/multisketch-eval + feat/multisketch-accuracy).

🤖 Generated with Claude Code

Adds Fig 3(c): Sum exact, DDSketch/KLL per-series quantiles (medians in-envelope,
p99 tail = small-N collapse variance), CMS exact freq, HLL per-series 3e-5 — 5/6
families validated. DDSketch-vs-KLL head-to-head (KLL better median, DDSketch tighter
tail + -26% wire). Two named gaps: CountSketch topk recall-0 (keys by item not host,
ranks by occurrence not value → value-weighted topk needs a separate path) and HLL
global rollup (per-series exact, global not served). Notes wall-clock anchoring was
the warm-read fix (timing not reducer) + the warm-vs-archive range-routing finding.
@zzylol
zzylol merged commit d66e882 into main Jun 12, 2026
@zzylol
zzylol deleted the docs/perfamily-accuracy branch June 12, 2026 21:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant