Repository navigation
Decide: retire the launch-benchmark apparatus, or keep which tiers of it? #296
Description
Activity
breadcrumb: Claimed for grilling. The body's "Tier A is broken" half is overtaken: PR #373 (merged 2026-08-22) repaired the workflow, #292 is closed, bench.yml has 20 straight green runs, and the trend holds 82 points of which 36-82 measure the Rust binary. The alert-noise half is confirmed and current: 18 Performance Alerts fired in the 6 days since the repair — 15 of 19 alerting series-rows are
host-prep, 3devpod-up, 1attach, and still zero ever ontotalortools. Q1 therefore mutates: repair-vs-retire is moot (repair is sunk cost); the live question is retire the now-green apparatus, or keep it and fix the alert mis-scope.Resolution
Retire Tier A. Keep Tier B unchanged. Hold Tier C as-is. One gating e2e case inherits Tier A's only real catch. Decided with the owner, 2026-08-29.
Evidence at decision time
The ticket body's "broken" half was overtaken before this grilling ran: PR #373 (merged 2026-08-22) repaired the workflow, #292 closed, and
bench.ymlhad 20 straight green runs with 92 published points, 57 of them the Rust binary. The decision was therefore made against a working Tier A, on the noise half — which the Rust series confirmed rather than softened:- 31 alerts lifetime, 0 real regressions. 18 of them in the six days since the repair: of 19 alerting series rows, 15 were
host-prep, 3devpod-up, 1attach.totalandtools— the numbers the design exists to protect — have alerted zero times, ever. - The mis-scope is structural, not bad luck. The 150% point-to-previous threshold sits inside the noise band for small stages (warm
host-prepmedian is 42µs on the Rust series) and at the edge of it for the totals: warmtotalranged 1.15s–2.05s across 57 points with no code cause, max noise jump 1.46x — so a real +40% regression intotalis indistinguishable from runner jitter, and a real +20% can never fire. - Drift is invisible by design (each point compares only to the previous one, per Where does the trend live: github-action-benchmark vs committed JSONL, alert-don't-gate #197). The Rust series drifted — warm
totalmedian +14% first-third to last-third, coldhost-prep+75% — and the alert channel said nothing intelligible about either. - Every measurement that changed a decision came from Tier B (Collapse the three sequential devpod ssh round trips on the cold path #157/One cold-path setup pass: fold the hostname into the probe's round trip, with per-stage outcomes #168's fold, Cache the staged tools tar instead of re-staging ~342MB every transfer #158's stopped cache, Warm path: defer clone-manager construction past the fast-attach check #145's finding). The one real bug Tier A surfaced (Shared pixi cache mount makes ~/.cache root-owned in any image that lacks it, and hard-fails on a uid mismatch #240) it caught by failing to launch — a shape one e2e case can gate for a fraction of the code.
An alert channel at 31/31 false positives is negative value: it trains the owner to delete the notification unread. And the chart's passive value is served on demand by
pixi run bench --recordagainst any two commits, retroactively, on quiet hardware.The five decisions
- Tier A retired —
bench.yml,bench_points.py,bench_cold_reset.sh, the bench guards, anddev/bench/on gh-pages (a frozen chart that looks live is a trap; the points stay in that branch's history). Build ticket: Retire Tier A: delete the CI trend workflow, scripts, guards, and the published chart #503. - Tier B kept unchanged —
bench_launch.py+ fixture remain the way launch performance is measured, including Launch latency: dl <spec> -- <cmd> to a running command #139's pending Rust re-baseline. - Tier C held as-is — live wf contract (
DEVLAUNCH_HANDOFF_T0/DEVLAUNCH_PREWARM_FIRED_ATstamped by wf on every launch) and frozen public API per devlaunch-core module architecture and data model #251 §7. The "doesjsonmode earn its lines" shrink question is parked in the map's fog, weak while Tier B consumes that mode. - One e2e case inherits the foreign-uid real-repo launch that caught Shared pixi cache mount makes ~/.cache root-owned in any image that lacks it, and hard-fails on a uid mismatch #240 — build ticket e2e inherits the bench's foreign-uid real-repo launch coverage #502, wired to block Retire Tier A: delete the CI trend workflow, scripts, guards, and the published chart #503 so the coverage never lapses.
- devpod-up is still one opaque number: build #195's sub-lines and the unattributed remainder #293 closed unbuilt — pure Tier A; Can devpod up's output be decomposed into image / create / postCreate? #195's
devpod-updecomposition stands on the map as decided, deliberately not built (parser preserved onresearch/devpod-up-decomposition). wayfinder#159 gets a pointer comment to this resolution; the re-chart of that map is its own to do.
The map's Destination is rewritten accordingly (the trend-phrased original is preserved in this ticket's body and the map's history); the map now closes on #502 and #503.
- 31 alerts lifetime, 0 real regressions. 18 of them in the six days since the repair: of 19 alerting series rows, 15 were
Question
Is the launch-benchmark apparatus worth its code? Raised by the owner (2026-08-20): "i was considering deprecating the benchmarks as its a lot of code."
It is a lot of code — ~4,150 lines — but it is three separable things with three completely different track records, and the answer is different for each. Deciding this as one lump is the trap.
bench.yml(289),bench_points.py(311),bench_cold_reset.sh(43),test_bench_workflow.py(287),test_bench_points.py(428),test_bench_doc.py(356),test_bench_record_schema.py(149), plus thegh-pagesbranch and the published chartscripts/bench_launch.py(316),test/fixtures/bench_harness.py(107)DEVLAUNCH_TIMINGrust/devlaunch-core/src/timing.rs(1,866, of which 961 are tests)The recommendation up front: retire Tier A, keep B, keep C. That deletes ~1,860 lines — about 80% of the bench-specific code — and it deletes the tier with the worst record while keeping both tiers that have actually paid for themselves. The rest of this ticket is the evidence, because one part of it is genuinely not ours to decide alone.
Tier A has never once done its job
Charter decision 1 promised a trend that alerts on a real regression without gating. Six days of history, measured from the Actions API and the commit comments:
It is broken more than it works. 64 runs, 35 green and 29 failed. Green only in two bursts — roughly 15 hours on 08-14/15, then about 90 minutes on 08-20. Everything between 08-15T13:20 and 08-20T17:27 failed (~20 consecutive runs, five days), and everything since 08-20T19:19 has failed again — that is the
devpod-on-PATHbreak in #292, now six runs deep and still firing.Every alert it has ever raised was noise. 13 alerts fired. One — the first,
f5f13332— is the deliberate 1%-threshold dispatch recorded on #198, which named all nine series at once to prove the path worked. Of the remaining 12:host-prepin 10 of 12.devpod-upin 3 (devpod's own variance — explicitly out of scope for this map: "we time it; we don't fix it").totalin zero.toolsin zero.attachin zero.So the load-bearing number the whole design protects has never alerted, and the dominant cold cost (
tools, ~4.6–6s) has never alerted. Four of those alerts were warmhost-prep, and here are the actual values that paged you:Three tenths of a millisecond, four times. And this is not bad luck — it is a mis-scope provable from the repo's own tests. The charter's own words, excluding finer spans from the trend: "a 0.03s
gh auth tokenline swinging 50% must not page anyone." Butclients/gh.rsasserts, in a test:assert_eq!(span_labels(&document, "host-prep"), ["gh auth token"]). On the warm pathhost-prepis that line — the exact example the charter used for what must never alert was itself promoted to a first-class alerting stage. A relative threshold on a sub-millisecond stage cannot mean anything, ever.Net: 13 alerts against 35 published points — 37% of all published points paged the owner, and 0% of them were a regression. An alert channel with a 100% false-positive rate is worse than none: it trains you to ignore the one that matters.
And the numbers it published were the wrong binary's. All 35 points measure the Python build. Until
1654dafthese steps ranpixi run dl, which resolved to the editable install's console script — so after the 0.1.0 cutover the trend was charting a build that no longer shipped, and nothing said so. The commit that fixed that is the commit that broke the job. Not one point in the series describes the binary users run.What is not an argument against it: CI cost. Median green run 2m30s, 100 minutes total across all 64 runs. It is cheap. The cost is code and attention, not compute. (Caveat: the repaired version adds
cargo build --release, so a fixed Tier A would be slower than any run measured here.)Tier B is where all the value actually came from
Every measurement this map and #139 trade on came from the local harness, not the trend:
423 lines, no CI job, no gh-pages branch, no alert threshold, runs when asked. It is also exactly what #139's open re-baseline question needs — the Rust binary's numbers are still unknown, and
pixi run bench --recordis the method already built. Keep it.Tier C has two obligations that outlive the trend
timing.rsis not bench code, and this is the part to be careful with:wfnow stampsDEVLAUNCH_HANDOFF_T0andDEVLAUNCH_PREWARM_FIRED_ATon every launch.dl's timing is the only thing that reads them. Retire it and wf is writing to nobody.devlaunch-core's surface, andpublic-api.txtis enforced by thepublic-apiCI job. Removing them is a deliberate API break, not a cleanup.There is still a fair shrink question here — whether the
jsondocument mode earns its share of the 905 non-test lines once nothing consumes it on a schedule, or whether prose plus a leaner document would do. That is a separate decision and should not ride on this one.The one thing that is not ours to decide alone
blooop/wayfinder#159 "Launch benchmarking: wf's side of the seam" is open, and its destination reads "stamped across the seam so the devlaunch trend can carry the true end-to-end total." That map's two tickets are closed and its stamps have shipped; retiring Tier A pulls the consumer out from under an open map in another repo. Charter decision 2 made these two maps deliberately separate, so this needs a decision there, not a unilateral deletion here. Nothing about Tier A's record changes on that repo's side — but the re-chart is theirs to do.
The one thing Tier A caught, and whether it is covered
Honesty requires this: the bench did catch a real user-facing bug — #240, the shared-pixi-cache uid mismatch that hard-fails a launch in any image lacking the directory. But it caught it by failing to launch at all, not by charting anything: the trend line and the alert had no part in it. The gating
e2ejob already runs real devpod launches on every PR and merge — except that it uses a fixture image, while the bench launches a real GitHub repo (blooop/mcp-devtasks) whose devcontainer has a differentremoteUseruid and runspixi installin postCreate. That difference is precisely what exposed #240. So if Tier A goes, the question to settle in the same breath is whether one e2e case should inherit "launch one real repo with a foreign uid end-to-end" — which is cheap, and gates, and is the part that ever worked.What settling this decides downstream
devpod-upsub-lines) are both pure Tier A and both die with it. This ticket blocks them: repairing and then extending an apparatus that may be deleted is the one clearly wrong order.To grill
host-prepanddevpod-upgiven no threshold (which would leavetotalandtools— neither of which has ever alerted — as the only alerting series)?