eval: first paper §6.2 sweep at N=1 - #179
Merged
Merged
Conversation
First end-to-end run of all 5 baselines via the PR #178 sweep script. Values are from the 75s soak at (rate=1000/s, cardinality=1000) on a log-normal gauge workload: Baseline CPU RSS Out MiB/s × vs B0 B0 raw OTel 0.77c 214MB 16.34 1.00× B1 Serf XOR 0.59c 271MB 3.44 1/ 4.8× B5 Gorilla XOR 0.58c 297MB 3.84 1/ 4.3× B2 full sketch 1.00c 506MB 163.77 10.03× ⚠ 10× larger than raw B3 delta (60s) 0.45c 225MB 0.053 1/305.8× ✓ paper's bandwidth claim Two takeaways worth flagging: * B2 full-sketch is an order of magnitude WORSE than raw at 1-second batch cadence — per-batch DDSketch+HLL state at cardinality=1000 dwarfs the raw gauge points it summarizes. This motivates why naive sketch-per-batch isn't a viable deployment, and why the paper needs delta + windowing. * B3 delta (60s) hits 305× reduction vs raw, exactly the bandwidth story §6.2 leads with. All four sketch families (DDSketch, HLL, CountSketch, CountMin) are delta-ready end-to-end (ASAPQuery-backend #60-#63 chain) so §6.4 accuracy comparisons can run against reconstituted sketches. * B1 Serf + B5 Gorilla compression is modest (~5% and 2.4%) because log-normal synthetic data has poor temporal correlation; real trace data typically compresses much better. Google-cluster-trace replay (DC TODO #4) is the follow-up that would fix this. Raw CSV archived for reproducibility; regenerate via: BASELINES="b0-raw b1-serf b2-full b3-delta b5-gorilla" \ SCALE=N1 RATES=1000 CARDS=1000 SOAK_S=75 \ ./deploy/scripts/run-baseline-sweep.sh > out.csv B3's measured out-bytes takes a longer soak (≥3 min) because of its 60s window — the sweep default of 75s yields one flush, and the script's 1m rate() window returns NaN on a single sample. The archived number (0.053 MiB/s) is from a separate 180s soak. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First end-to-end run of all 5 baselines via PR #178's sweep script. 75s soak at (rate=1000/s, cardinality=1000) on log-normal gauge workload.
Two key findings for the paper:
B1/B5 compression is modest on log-normal synthetic data (poor temporal correlation); Google-cluster-trace replay (TODO #4) would exercise them more fairly.
Reproducibility:
BASELINES="b0-raw b1-serf b2-full b3-delta b5-gorilla" SCALE=N1 RATES=1000 CARDS=1000 SOAK_S=75 ./deploy/scripts/run-baseline-sweep.shNote: B3 needs ≥3min soak because rate() over 1m on a 60s-window emitter needs ≥2 samples; the archived number is from a 180s re-run.
🤖 Generated with Claude Code