Skip to content

eval: first paper §6.2 sweep at N=1 - #179

Merged
zzylol merged 1 commit into
mainfrom
feat/sweep-eval-results
Apr 22, 2026
Merged

zzylol merged 1 commit into
mainfrom
feat/sweep-eval-results

Conversation

@zzylol

@zzylol zzylol commented Apr 22, 2026

Copy link
Copy Markdown
Contributor

First end-to-end run of all 5 baselines via PR #178's sweep script. 75s soak at (rate=1000/s, cardinality=1000) on log-normal gauge workload.

Baseline CPU RSS Out MiB/s × vs B0
B0 raw OTel 0.77c 214MB 16.34 1.00×
B1 Serf XOR 0.59c 271MB 3.44 1/ 4.8×
B5 Gorilla XOR 0.58c 297MB 3.84 1/ 4.3×
B2 full sketch 1.00c 506MB 163.77 10.03× (bigger!)
B3 delta (60s) 0.45c 225MB 0.053 1/305.8×

Two key findings for the paper:

  1. B2 full-sketch per-batch is 10× larger than raw at 1s cadence — motivates why the paper's §6.2 delta story matters. Per-batch sketch state at cardinality=1000 dwarfs the raw gauge points it summarizes.
  2. B3 delta at 60s window hits 305× reduction — the bandwidth claim the paper leads with, demonstrated end-to-end.

B1/B5 compression is modest on log-normal synthetic data (poor temporal correlation); Google-cluster-trace replay (TODO #4) would exercise them more fairly.

Reproducibility: BASELINES="b0-raw b1-serf b2-full b3-delta b5-gorilla" SCALE=N1 RATES=1000 CARDS=1000 SOAK_S=75 ./deploy/scripts/run-baseline-sweep.sh

Note: B3 needs ≥3min soak because rate() over 1m on a 60s-window emitter needs ≥2 samples; the archived number is from a 180s re-run.

🤖 Generated with Claude Code

First end-to-end run of all 5 baselines via the PR #178 sweep
script. Values are from the 75s soak at (rate=1000/s,
cardinality=1000) on a log-normal gauge workload:

  Baseline          CPU    RSS    Out MiB/s   × vs B0
  B0 raw OTel       0.77c  214MB  16.34       1.00×
  B1 Serf XOR       0.59c  271MB   3.44       1/ 4.8×
  B5 Gorilla XOR    0.58c  297MB   3.84       1/ 4.3×
  B2 full sketch    1.00c  506MB  163.77     10.03×    ⚠ 10× larger than raw
  B3 delta (60s)    0.45c  225MB   0.053      1/305.8×  ✓ paper's bandwidth claim

Two takeaways worth flagging:

  * B2 full-sketch is an order of magnitude WORSE than raw at
    1-second batch cadence — per-batch DDSketch+HLL state at
    cardinality=1000 dwarfs the raw gauge points it summarizes.
    This motivates why naive sketch-per-batch isn't a viable
    deployment, and why the paper needs delta + windowing.
  * B3 delta (60s) hits 305× reduction vs raw, exactly the
    bandwidth story §6.2 leads with. All four sketch families
    (DDSketch, HLL, CountSketch, CountMin) are delta-ready
    end-to-end (ASAPQuery-backend #60-#63 chain) so §6.4
    accuracy comparisons can run against reconstituted sketches.
  * B1 Serf + B5 Gorilla compression is modest (~5% and 2.4%)
    because log-normal synthetic data has poor temporal
    correlation; real trace data typically compresses much
    better. Google-cluster-trace replay (DC TODO #4) is the
    follow-up that would fix this.

Raw CSV archived for reproducibility; regenerate via:

  BASELINES="b0-raw b1-serf b2-full b3-delta b5-gorilla" \
    SCALE=N1 RATES=1000 CARDS=1000 SOAK_S=75 \
    ./deploy/scripts/run-baseline-sweep.sh > out.csv

B3's measured out-bytes takes a longer soak (≥3 min) because
of its 60s window — the sweep default of 75s yields one flush,
and the script's 1m rate() window returns NaN on a single
sample. The archived number (0.053 MiB/s) is from a separate
180s soak.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@zzylol
zzylol merged commit 44e008d into main Apr 22, 2026
@zzylol
zzylol deleted the feat/sweep-eval-results branch April 22, 2026 13:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant