Repository navigation
Conversation
f60b2bc to
a8282ce
Compare
c664c1a to
34600c6
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Moderate issues remain around unsupported parallel apply and inaccurate or overly broad option handling and resolved-value logging.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (2)
What changed in this PR
Adds --overlay-v2-optimized for Rust-overlay benchmark defaults, with opt-outs, configuration tuning, mesh recovery, tests, and documentation.
Changes:
- Adds CLI options, preset resolution, and resolved-value logging.
- Applies optimized resources, core settings, account sizing, and Soroban limits.
- Adds placement checks, PostgreSQL waits, bounded mesh recovery, tests, and documentation.
| File | Summary |
|---|---|
src/FSLibrary/StellarStatefulSets.fs |
Mesh progress probes |
src/FSLibrary/StellarMissionContext.fs |
Preset resolution and logging |
src/FSLibrary/StellarKubeSpecs.fs |
Resources, environment, labels, and database waits |
src/FSLibrary/StellarCoreHTTP.fs |
Non-blocking connectivity probes |
src/FSLibrary/StellarCoreCfg.fs |
Optimized core configuration |
src/FSLibrary/MissionMinBlockTimeMixed.fs |
Mixed mission defaults |
src/FSLibrary/MissionMinBlockTimeClassic.fs |
Classic mission defaults |
src/FSLibrary/MissionMaxTPSMixed.fs |
Mixed MaxTPS defaults |
src/FSLibrary/MissionMaxTPSClassic.fs |
Classic MaxTPS defaults |
src/FSLibrary/MinBlockTimeTest.fs |
Account checks, placement, and mesh recovery |
src/FSLibrary/MaxTPSTest.fs |
MaxTPS limits and mesh recovery |
src/FSLibrary.Tests/Tests.fs |
Unit coverage |
src/App/Program.fs |
CLI options and validation |
doc/measuring-minimum-block-time.md |
Preset usage documentation |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| `--overlay-v2-optimized` applies, in one flag, the settings that benchmarks of the experimental Rust-overlay stellar-core image use. Without it missions keep their standard configs, resources and limits, so runs against the stellar-core master image need nothing. Any setting passed explicitly still wins, and the run log lists the resolved values (`--overlay-v2-optimized: ...` lines). It sets: | ||
|
|
||
| * the tx-set buffer to 125% and the Soroban byte allowance to 9 MiB (see above); | ||
| * for the perf missions (`MinBlockTimeClassic`/`Mixed`, `MaxTPSClassic`/`Mixed`): in-memory BucketListDB (`--disk-backed-buckets` opts out), no test-only tx meta (`--keep-tx-meta` opts out, and is needed for images older than v27.0.0), and validators with an 8 vCPU request, no CPU limit and 16 GiB memory (`--validator-cpu-limit-mcpu` restores a limit) whose containers get `TOKIO_WORKER_THREADS=8` unless `--core-env` sets it; |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical issues remain in mesh restart handling, database support, and validator labeling.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 3
Open (6)
Route all v2 restarts through bounded mesh retry logic · New Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes · New Label only validator CoreSets as validator pods · New Restrict overlay-v2 no-key exemption to MinBlockTime missions · New Log effective Soroban byte allowance Correct resource documentation for MaxTPSClassic
| if not self.network.missionContext.enableParallelApply then | ||
| t.Add("EXPERIMENTAL_PARALLEL_LEDGER_APPLY", true) |> ignore |
| if self.missionContext.oneValidatorPerHost | ||
| || self.missionContext.overlayV2Optimized then | ||
| self.PodLabels() | ||
| |> Map.add CfgVal.validatorRoleLabelKey CfgVal.validatorRoleLabelValue |
a8282ce to
99c4f9d
Compare
34600c6 to
4bb4e5c
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical timeout and moderate configuration, metrics, and documentation issues remain.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 4
Open (7)
Add an explicit timeout to TryGetInfo HTTP probes · New Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Route all v2 restarts through bounded mesh retry logic Restrict overlay-v2 no-key exemption to MinBlockTime missions Log effective Soroban byte allowance Correct resource documentation for MaxTPSClassic
| member self.TryGetInfo() : Info.Info option = | ||
| try | ||
| let parsed = Info.Parse(self.fetch "info") | ||
|
|
99c4f9d to
abdd1da
Compare
4bb4e5c to
1ab522a
Compare
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Unresolved moderate issues affect database configuration, latency reporting, mesh timeout bounds, and CLI validation.
Review effort: Lite
Findings: 4
Open (7)
Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Route all v2 restarts through bounded mesh retry logic Restrict overlay-v2 no-key exemption to MinBlockTime missions Log effective Soroban byte allowance Correct resource documentation for MaxTPSClassic
abdd1da to
853f1fa
Compare
1ab522a to
23ce02e
Compare
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Three moderate findings remain unresolved.
Review effort: Lite
Findings: 4
Open (7)
Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Route all v2 restarts through bounded mesh retry logic Restrict overlay-v2 no-key exemption to MinBlockTime missions Log effective Soroban byte allowance Correct resource documentation for MaxTPSClassic
23ce02e to
c2889f4
Compare
853f1fa to
df242c3
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved issues affect validation, PostgreSQL compatibility, and mixed-mode latency behavior.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 4
Open (8)
Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Route all v2 restarts through bounded mesh retry logic Restrict overlay-v2 no-key exemption to MinBlockTime missions Log effective Soroban byte allowance Report resolved resources accurately for MaxTPSClassic · New Correct resource documentation for MaxTPSClassic
| (match ctx.validatorCpuLimitMcpu with | ||
| | Some l -> sprintf "perf validators: %dm CPU limit, 16Gi memory" l | ||
| | None -> "perf validators: 8 vCPU request, no CPU limit, 16Gi memory") |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical and moderate issues remain in bounded waits, timeout enforcement, and benchmark configuration.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 12
Open (18)
Optimized path can enter unbounded connection wait · New Boot deadline bypassed when final node answers late · New Remove unbounded mesh wait after bounded recovery check Collect e2e latency metrics from all MinBlockTime nodes Delay stall detection until all validators are ready Preserve classic allowance for classic-only workloads Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Prevent duplicate TOML entries for incompatible mission options Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Align warm-up documentation with actual measurement window Log the actual MaxTPS Soroban allowance Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
Resolved since last review (2)
| if context.overlayV2Optimized then formation.EnsureMeshedOrRedraw coreSets | ||
|
|
||
| // Setup overlay connections first before manually closing | ||
| // ledger, which kick off consensus | ||
| formation.WaitUntilConnected coreSets |
| elif s.meshStartSec.IsNone && not (List.isEmpty m.silent) then | ||
| (if nowSec >= b.bootTimeoutSec then BootTimedOut else KeepWaiting), s | ||
| else | ||
| let meshStart = defaultArg s.meshStartSec nowSec |
1a13191 to
8de7314
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved critical and moderate issues affect load sizing, PostgreSQL job setup, mesh-wait bounds, and run-report accuracy.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 15
Open (21)
Pre-generated transaction pools are too small for fixed-duration runs · New Job pod database configuration ignores usesPostgres · New Legacy connection wait reintroduces unbounded retries · New Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Collect e2e latency metrics from all MinBlockTime nodes Delay stall detection until all validators are ready Preserve classic allowance for classic-only workloads Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Prevent duplicate TOML entries for incompatible mission options Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Align warm-up documentation with actual measurement window Log the actual MaxTPS Soroban allowance Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic
And 1 more that still need to be addressed.
| // Measurement window at fixed TPS: ~5 min, enough for a | ||
| // stable read of the SLA metric without draining the tx | ||
| // source, or longer under --overlay-v2-optimized. | ||
| txs = fixedTxRate * loadDurationSec |
8de7314 to
8e24756
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
The redraw restart path can wait indefinitely for replica readiness, so the advertised recovery bound is not enforced.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 13
Open (18)
Unbounded replica readiness wait can hang mesh restart attempts · New Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Preserve classic allowance for classic-only workloads Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Prevent duplicate TOML entries for incompatible mission options Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Align warm-up documentation with actual measurement window Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
| |> List.map (fun cs -> async { self.Start cs.name }) | ||
| |> Async.Parallel |
8e24756 to
96d15a6
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Several moderate issues affect documented timing, mesh failure bounds, load-generator selection, and runtime logging.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 11
Open (17)
Unbounded replica readiness wait can hang mesh restart attempts Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Serial mesh probes can exceed advertised wall-clock bounds · New Align warm-up documentation with actual measurement window Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
Resolved since last review (2)
| let counts = | ||
| peers | ||
| |> List.map (fun p -> p, p.TryGetInfo() |> Option.map (fun i -> i.Peers.AuthenticatedCount)) |
96d15a6 to
f365b20
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Moderate issues remain in MinBlockTime sizing/scheduling and overlay mesh probing.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 11
Open (17)
Unbounded replica readiness wait can hang mesh restart attempts Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Metric retries can block scheduled ledger-age snapshots for 200 seconds · New Serial mesh probes can exceed advertised wall-clock bounds Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
Resolved since last review (1)
| let read () = | ||
| try | ||
| Ok(collectLedgerAgePercentiles formation coreSets) | ||
| with e -> Error e.Message |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Two unresolved findings remain, including a critical mesh-readiness issue and a moderate periodic-metrics timeout issue.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 12
Open (18)
Use live authenticated-peer gauge instead of cumulative connection counter · New Unbounded replica readiness wait can hang mesh restart attempts Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Metric retries can block scheduled ledger-age snapshots for 200 seconds Serial mesh probes can exceed advertised wall-clock bounds Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
| member self.TryGetAuthenticatedCount() : int option = | ||
| try | ||
| Some(ParseMetricCount(self.fetch "metrics") "overlay.connection.authenticated") | ||
| with _ -> None |
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Three unresolved moderate findings affect bounded metric sampling and overlay recovery behavior.
Review effort: Lite
Findings: 12
Open (18)
Use live authenticated-peer gauge instead of cumulative connection counter Unbounded replica readiness wait can hang mesh restart attempts Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Metric retries can block scheduled ledger-age snapshots for 200 seconds Serial mesh probes can exceed advertised wall-clock bounds Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Two moderate issues can invalidate mesh readiness or MinBlockTime measurement results.
Review effort: Lite
Findings: 12
Open (18)
Use live authenticated-peer gauge instead of cumulative connection counter Unbounded replica readiness wait can hang mesh restart attempts Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Metric retries can block scheduled ledger-age snapshots for 200 seconds Serial mesh probes can exceed advertised wall-clock bounds Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Four moderate review findings remain unresolved.
Review effort: Lite
Findings: 12
Open (18)
Use live authenticated-peer gauge instead of cumulative connection counter Unbounded replica readiness wait can hang mesh restart attempts Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Metric retries can block scheduled ledger-age snapshots for 200 seconds Serial mesh probes can exceed advertised wall-clock bounds Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic
One flag applies the settings that benchmarks of the experimental Rust-overlay (v2) stellar-core image use, so runs against stellar-core master need nothing and v2 runs need no long flag lists. It is the only knob: each setting it changes is fixed, and the run log lists them. Without it, configs, resources and limits are unchanged. This first step sets, for the perf missions (MinBlockTimeClassic/Mixed, MaxTPSClassic/Mixed): * in-memory BucketListDB; * DISABLE_TX_META_FOR_TESTING, so test builds skip the per-ledger tx meta copies on the apply path (images older than v27.0.0 reject the key); * one stellar-core pod per worker node, as --one-stellar-core-per-host: enforced scheduling, a placement check whenever pods start, and a fast failure when the cluster cannot provide the nodes; * validators that reserve 8 vCPU and 16 GiB but set no CPU limit. Bursts from core's worker threads plus the overlay's tokio runtime exceeded an 8-CPU CFS quota: up to 4,909 throttled periods per run (at 5k TPS), each freezing both processes for ~40-58 ms. Without the flag, perf missions keep their upstream resources (SimulatePubnetTier1PerfResources, or MaxTPSClassicResources); * TOKIO_WORKER_THREADS=8 in those validators' containers, since without a quota tokio sizes its runtime to every CPU on the node; --core-env can override it. It also sets 8 dependent-tx clusters in the Soroban limit upgrades: MaxTPS's (shared by SimulatePubnetMixedLoad and MinBlockTime's non-mixed Soroban modes) and MinBlockTimeMixed's per-candidate upgrade, each waiting until the network reports it. The network default of 1 serializes Soroban execution and holds a ledger to one cluster's instructions. Overlay-only MinBlockTimeMixed runs apply nothing, so there it shapes the tx sets: up to 8 clusters per stage, and up to 8x the instructions, leaving the tx count and byte limits to bind. BUCKETLIST_DB_INDEX_PAGE_SIZE_EXPONENT is now decided in one place, so it is emitted at most once. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…wances Under the flag: * MinBlockTime* sizes the per-candidate classic MaxTxSetSize and Soroban ledgerMaxTxCount at 125% of the offered txs per ledger instead of 2x. Every tx-set build on the Rust-overlay core pulls twice these limits from the mempool over IPC and validates them, so 2x meant pulling 4x the offered load per ledger; 125% still lets a slow ledger's backlog drain. * MinBlockTime* runs 960 s of load per candidate instead of 300 s. * MinBlockTime* splits core's 10 MiB tx-set byte budget between the classic and Soroban phases in proportion to the bytes it offers (classic payments ~200 B, Soroban at the mode's tx-size bound), with at least 1 MiB each, instead of core's 5 MiB each. Core otherwise caps the Soroban phase at 5 MiB, about 6900 SAC payments per tx set, whatever the network limits say. A Soroban-only run gets 9 MiB for Soroban and a classic-only run 9 MiB for classic; both phases grow alike with the close time, so the split holds for every candidate. MissionContext.txSetByteAllowances decides the allowances for the node configs and the run log, including the max-TPS modes' own splits; other missions keep core's defaults. Without the flag the sizing, load window and allowances are upstream's. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The Soroban tx-limit upgrade shared by MaxTPS, SimulatePubnetMixedLoad and MinBlockTime's non-mixed Soroban modes raises each per-transaction limit from the load's distributions but left the contract-events cap at the network default. An invoke emitting more events than that allows fails on apply, which reads as transactions vanishing rather than a limit. Raise it with the transaction size limit, as the mixed-pregen path already raises its own cap. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
MinBlockTime* now marks the core sets that generate load (generatesLoad), as --loadgen-keys does for pubnet topologies, so --measure-e2e-latency puts the metric on exactly the nodes that submit. The synthetic tier1 core sets were never marked, so the option measured nothing there, and it could not even be requested without --pubnet-data: it required --loadgen-keys, which requires --pubnet-data. It now needs --loadgen-keys only for missions other than MinBlockTime*. The marking changes nothing unless e2e latency is measured. Under --overlay-v2-optimized, MinBlockTime* measures e2e latency by default. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The max-TPS missions point DATABASE at postgres and add the postgres sidecar whatever the core set's dbType, but the pod's postgres setup (PGHOST/PGUSER, createdb and the pg_isready wait) only ran for dbType Postgres, and core sets default to Sqlite. new-db could then race the sidecar's startup and crash the container. The three were decided in separate places; MissionContext.usesPostgres now decides DATABASE, the sidecar and the setup steps together, for stateful and job pods alike, so they cannot disagree. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A libp2p simultaneous-dial collision can leave one peer edge missing after a boot or a mass restart. It never heals, and WaitUntilConnected then waits forever. Under the flag, formation.WaitUntilConnected, which every mission uses to wait for the overlay, runs EnsureMeshedOrRedraw instead of the unbounded per-node wait. It returns only once every node is fully connected, and bounds the wait: * every node must answer within 300 s of a (re)start (the core container's liveness probe already restarts a core that is silent for about 3 minutes), or the run fails naming the silent nodes; * once all nodes answer, a mesh that stops growing for 60 s (two ticks of the overlay's 30 s safety-net reconnect), or is still incomplete after 120 s, is redrawn by restarting all nodes; * after 3 attempts the run fails. Connections are read from each node's overlay.connection.authenticated metric, in parallel. With the Rust overlay, core refreshes the peer counts /info reports only when /metrics is requested, so /info can read 0 while the overlay is fully connected. A wedged mesh now fails a run in about 20 minutes at worst instead of hanging it. In every recorded run of the v2 image the mesh completed within a minute of the first check, without a redraw. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Core keeps ledger.age.closed-histogram over a sliding 5-minute window, and the SLA read it once after the load, so a load longer than 5 minutes was judged on its last 5 minutes only (#401). MinBlockTime* now reads every node's percentiles during the load, at consecutive 5-minute windows that end with the planned load after a warm-up of the remainder. It keeps reading every 5 minutes if the load runs past its planned end, and reads once more when the load ends unless the last read is under 30 s old, so every part of the load after the warm-up is judged. Every window must pass on every node; a read that fails fails the candidate instead of dropping its window. Reads taken during the load also avoid the idle ledgers core closes after it. A 300 s load is one read at its end, as before; the 960 s load of --overlay-v2-optimized is a 60 s warm-up and three windows, read at 360, 660 and 960 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Without the flag MinBlockTime* keeps main's 2000 ms ballot and nomination timeouts. Under it, the best-case measurement uses 500 ms. Reference measurements, 30 nodes at 500 SAC TPS (2026-07-27, image 3453), one run each, at T=3000: 2000 ms: externalize p75 2447 ms, ledger-age p75 4592 ms (+53%, fail) 500 ms: externalize p75 331 ms, ledger-age p75 3028 ms (+0.9%) A ballot round that stalls recovers only when its timer fires, so each stall costs a full timeout. At 2000 ms that stretches the ledger, the next tx set is bigger at the fixed TPS, and bigger sets stall more often, so the stalls snowball; an externalize p75 above the timeout means over a quarter of slots stalled. At 500 ms a stall costs little, ledgers stay near target and stalls stay rare. The reasoning lives with MissionContext.minBlockTimeScpTimeoutMs, and the run log lists the value. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A ledger-age window that could not be read failed the candidate. If the load completed, that is a harness or network problem, not a verdict: the window was never measured, and failing the candidate would let the search raise its lower bound on no evidence. The mission now aborts instead, as it does when a completed load's other metrics cannot be read. After a failed load the candidate has already failed, usually because a node dropped out, which also leaves its windows unreadable; those are still logged and judged as failed, and the search continues. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Three moderate review findings remain unresolved.
Review effort: Lite
Findings: 12
Open (18)
Use live authenticated-peer gauge instead of cumulative connection counter Unbounded replica readiness wait can hang mesh restart attempts Pre-generated transaction pools are too small for fixed-duration runs Boot deadline bypassed when final node answers late Optimized path can enter unbounded connection wait Remove unbounded mesh wait after bounded recovery check Delay stall detection until all validators are ready Keep zero-total mesh progress in bounded retry logic Treat zero-peer mesh progress as not ready Add an explicit timeout to TryGetInfo HTTP probes Label only validator CoreSets as validator pods Avoid enabling parallel apply on SQLite-backed MinBlockTime nodes Metric retries can block scheduled ledger-age snapshots for 200 seconds Serial mesh probes can exceed advertised wall-clock bounds Document MaxTPSClassic resource sizing exception Resource log reports incorrect MaxTPSClassic settings Report resolved resources accurately for MaxTPSClassic Correct resource documentation for MaxTPSClassic



Stack: #437 → #438 → #439
Third of three stacked PRs. This PR does a bunch of tuning of settings for overlay v2 so that we measure the best-case performance scenario. It also adds a flag so that we can easily turn on these settings, providing easier reproduction.
Commits are as follows:
--overlay-v2-optimizedflag. This enables a variety of performance enhancement flags, including--core-envcan override it)46e5ee8 Tune MinBlockTime sizing and the tx-set byte allowances:
5ab1e3d Raise the per-transaction contract-events cap, this doesn't matter in practice for core and we started hitting it.
39aa9fe MinBlockTime* marks its load-generating core sets, so
--measure-e2e-latencyputs the e2e latency metric on our artificial tier 1 core sets. Under the v2 flag, MinBlockTime* measures e2e latency by default.72d705c Pods wait for their postgres sidecar. Max-TPS always uses a postgres database, but the pod's postgres setup and wait only ran for a postgres dbType, so
new-dbcould race the sidecar and crash.b6ea066 Added restarts and monitoring during network bootstrap for overlay v2. We don't have a real peer discovery system in v2, so sometimes the network can get wedged while initializing peer connections. We now check that connection counts are increasing, and if not, restart the network and try again.
a08fb27 Judge close times over the whole load, not just its last 5 minutes (MinBlockTimeTest hardening #401). Core's ledger-age histogram only covers a sliding 5-minute window, so we now read it during the load at consecutive 5-minute windows (after a warm-up of the remainder) and every window must pass; the 960 s v2 load is a 60 s warm-up plus three windows.