Skip to content

Decide: retire the launch-benchmark apparatus, or keep which tiers of it? #296

Description

@blooop

Question

Is the launch-benchmark apparatus worth its code? Raised by the owner (2026-08-20): "i was considering deprecating the benchmarks as its a lot of code."

It is a lot of code — ~4,150 lines — but it is three separable things with three completely different track records, and the answer is different for each. Deciding this as one lump is the trap.

Tier What Lines
A — the CI trend bench.yml (289), bench_points.py (311), bench_cold_reset.sh (43), test_bench_workflow.py (287), test_bench_points.py (428), test_bench_doc.py (356), test_bench_record_schema.py (149), plus the gh-pages branch and the published chart 1,863
B — the local bench scripts/bench_launch.py (316), test/fixtures/bench_harness.py (107) 423
C — DEVLAUNCH_TIMING rust/devlaunch-core/src/timing.rs (1,866, of which 961 are tests) 1,866

The recommendation up front: retire Tier A, keep B, keep C. That deletes ~1,860 lines — about 80% of the bench-specific code — and it deletes the tier with the worst record while keeping both tiers that have actually paid for themselves. The rest of this ticket is the evidence, because one part of it is genuinely not ours to decide alone.

Tier A has never once done its job

Charter decision 1 promised a trend that alerts on a real regression without gating. Six days of history, measured from the Actions API and the commit comments:

It is broken more than it works. 64 runs, 35 green and 29 failed. Green only in two bursts — roughly 15 hours on 08-14/15, then about 90 minutes on 08-20. Everything between 08-15T13:20 and 08-20T17:27 failed (~20 consecutive runs, five days), and everything since 08-20T19:19 has failed again — that is the devpod-on-PATH break in #292, now six runs deep and still firing.

Every alert it has ever raised was noise. 13 alerts fired. One — the first, f5f13332 — is the deliberate 1%-threshold dispatch recorded on #198, which named all nine series at once to prove the path worked. Of the remaining 12:

  • host-prep in 10 of 12.
  • devpod-up in 3 (devpod's own variance — explicitly out of scope for this map: "we time it; we don't fix it").
  • total in zero. tools in zero. attach in zero.

So the load-bearing number the whole design protects has never alerted, and the dominant cold cost (tools, ~4.6–6s) has never alerted. Four of those alerts were warm host-prep, and here are the actual values that paged you:

0.000517 s  ->  0.000798 s        1.54x
0.000445 s  ->  0.000798 s        1.79x
0.000506 s  ->  0.000840 s        1.66x
0.000612 s  ->  0.000937 s        1.53x

Three tenths of a millisecond, four times. And this is not bad luck — it is a mis-scope provable from the repo's own tests. The charter's own words, excluding finer spans from the trend: "a 0.03s gh auth token line swinging 50% must not page anyone." But clients/gh.rs asserts, in a test: assert_eq!(span_labels(&document, "host-prep"), ["gh auth token"]). On the warm path host-prep is that line — the exact example the charter used for what must never alert was itself promoted to a first-class alerting stage. A relative threshold on a sub-millisecond stage cannot mean anything, ever.

Net: 13 alerts against 35 published points — 37% of all published points paged the owner, and 0% of them were a regression. An alert channel with a 100% false-positive rate is worse than none: it trains you to ignore the one that matters.

And the numbers it published were the wrong binary's. All 35 points measure the Python build. Until 1654daf these steps ran pixi run dl, which resolved to the editable install's console script — so after the 0.1.0 cutover the trend was charting a build that no longer shipped, and nothing said so. The commit that fixed that is the commit that broke the job. Not one point in the series describes the binary users run.

What is not an argument against it: CI cost. Median green run 2m30s, 100 minutes total across all 64 runs. It is cheap. The cost is code and attention, not compute. (Caveat: the repaired version adds cargo build --release, so a fixed Tier A would be slower than any run measured here.)

Tier B is where all the value actually came from

Every measurement this map and #139 trade on came from the local harness, not the trend:

423 lines, no CI job, no gh-pages branch, no alert threshold, runs when asked. It is also exactly what #139's open re-baseline question needs — the Rust binary's numbers are still unknown, and pixi run bench --record is the method already built. Keep it.

Tier C has two obligations that outlive the trend

timing.rs is not bench code, and this is the part to be careful with:

  1. It is a live cross-repo contract. wayfinder PR #162 merged 2026-08-14 and wf now stamps DEVLAUNCH_HANDOFF_T0 and DEVLAUNCH_PREWARM_FIRED_AT on every launch. dl's timing is the only thing that reads them. Retire it and wf is writing to nobody.
  2. The stage names are frozen public API. devlaunch-core module architecture and data model #251 §7 pinned them into devlaunch-core's surface, and public-api.txt is enforced by the public-api CI job. Removing them is a deliberate API break, not a cleanup.

There is still a fair shrink question here — whether the json document mode earns its share of the 905 non-test lines once nothing consumes it on a schedule, or whether prose plus a leaner document would do. That is a separate decision and should not ride on this one.

The one thing that is not ours to decide alone

blooop/wayfinder#159 "Launch benchmarking: wf's side of the seam" is open, and its destination reads "stamped across the seam so the devlaunch trend can carry the true end-to-end total." That map's two tickets are closed and its stamps have shipped; retiring Tier A pulls the consumer out from under an open map in another repo. Charter decision 2 made these two maps deliberately separate, so this needs a decision there, not a unilateral deletion here. Nothing about Tier A's record changes on that repo's side — but the re-chart is theirs to do.

The one thing Tier A caught, and whether it is covered

Honesty requires this: the bench did catch a real user-facing bug — #240, the shared-pixi-cache uid mismatch that hard-fails a launch in any image lacking the directory. But it caught it by failing to launch at all, not by charting anything: the trend line and the alert had no part in it. The gating e2e job already runs real devpod launches on every PR and merge — except that it uses a fixture image, while the bench launches a real GitHub repo (blooop/mcp-devtasks) whose devcontainer has a different remoteUser uid and runs pixi install in postCreate. That difference is precisely what exposed #240. So if Tier A goes, the question to settle in the same breath is whether one e2e case should inherit "launch one real repo with a foreign uid end-to-end" — which is cheap, and gates, and is the part that ever worked.

What settling this decides downstream

To grill

  1. Tier A: retire, or repair once more with host-prep and devpod-up given no threshold (which would leave total and tools — neither of which has ever alerted — as the only alerting series)?
  2. If retired: does an e2e case inherit the real-repo/foreign-uid launch coverage?
  3. Tier C: hold as-is on the wf contract and the frozen API, or open the shrink question separately?
  4. Who tells wayfinder#159?

Activity

  1. self-assigned this
    on Aug 29, 2026
  2. blooop commented on Aug 29, 2026

    @blooop
    OwnerAuthor

    breadcrumb: Claimed for grilling. The body's "Tier A is broken" half is overtaken: PR #373 (merged 2026-08-22) repaired the workflow, #292 is closed, bench.yml has 20 straight green runs, and the trend holds 82 points of which 36-82 measure the Rust binary. The alert-noise half is confirmed and current: 18 Performance Alerts fired in the 6 days since the repair — 15 of 19 alerting series-rows are host-prep, 3 devpod-up, 1 attach, and still zero ever on total or tools. Q1 therefore mutates: repair-vs-retire is moot (repair is sunk cost); the live question is retire the now-green apparatus, or keep it and fix the alert mis-scope.

  3. blooop commented on Aug 29, 2026

    @blooop
    OwnerAuthor

    Resolution

    Retire Tier A. Keep Tier B unchanged. Hold Tier C as-is. One gating e2e case inherits Tier A's only real catch. Decided with the owner, 2026-08-29.

    Evidence at decision time

    The ticket body's "broken" half was overtaken before this grilling ran: PR #373 (merged 2026-08-22) repaired the workflow, #292 closed, and bench.yml had 20 straight green runs with 92 published points, 57 of them the Rust binary. The decision was therefore made against a working Tier A, on the noise half — which the Rust series confirmed rather than softened:

    An alert channel at 31/31 false positives is negative value: it trains the owner to delete the notification unread. And the chart's passive value is served on demand by pixi run bench --record against any two commits, retroactively, on quiet hardware.

    The five decisions

    1. Tier A retired — bench.yml, bench_points.py, bench_cold_reset.sh, the bench guards, and dev/bench/ on gh-pages (a frozen chart that looks live is a trap; the points stay in that branch's history). Build ticket: Retire Tier A: delete the CI trend workflow, scripts, guards, and the published chart #503.
    2. Tier B kept unchanged — bench_launch.py + fixture remain the way launch performance is measured, including Launch latency: dl <spec> -- <cmd> to a running command #139's pending Rust re-baseline.
    3. Tier C held as-is — live wf contract (DEVLAUNCH_HANDOFF_T0 / DEVLAUNCH_PREWARM_FIRED_AT stamped by wf on every launch) and frozen public API per devlaunch-core module architecture and data model #251 §7. The "does json mode earn its lines" shrink question is parked in the map's fog, weak while Tier B consumes that mode.
    4. One e2e case inherits the foreign-uid real-repo launch that caught Shared pixi cache mount makes ~/.cache root-owned in any image that lacks it, and hard-fails on a uid mismatch #240 — build ticket e2e inherits the bench's foreign-uid real-repo launch coverage #502, wired to block Retire Tier A: delete the CI trend workflow, scripts, guards, and the published chart #503 so the coverage never lapses.
    5. devpod-up is still one opaque number: build #195's sub-lines and the unattributed remainder #293 closed unbuilt — pure Tier A; Can devpod up's output be decomposed into image / create / postCreate? #195's devpod-up decomposition stands on the map as decided, deliberately not built (parser preserved on research/devpod-up-decomposition). wayfinder#159 gets a pointer comment to this resolution; the re-chart of that map is its own to do.

    The map's Destination is rewritten accordingly (the trend-phrased original is preserved in this ticket's body and the map's history); the map now closes on #502 and #503.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions