diff --git a/doc/measuring-minimum-block-time.md b/doc/measuring-minimum-block-time.md index f57c90fc..2138e677 100644 --- a/doc/measuring-minimum-block-time.md +++ b/doc/measuring-minimum-block-time.md @@ -7,7 +7,7 @@ Supercluster provides two missions that measure the minimum ledger target close Other than the type of load generated, these two missions are identical. They both spin up a configurable network of stellar-core nodes, then search for the smallest `ledgerTargetCloseTimeMilliseconds` value that the network can sustain, using a binary search over the range `[--min-block-time-ms, --max-block-time-ms]`. For each candidate close time `T`, the missions upgrade the network's SCP timing settings to `T` (with proportionally scaled ballot and nomination timeouts), run ~5 minutes of load at the fixed transaction rate, then check the `ledger.age.closed-histogram` metric on every node against the SLA. If the SLA is met, the missions try again with a smaller `T`; otherwise, they try with a larger `T`. -The missions perform the binary search to find the minimum sustainable block time. Upon completion, the missions emit a log line of the form `Minimum sustainable block time: 4500 ms (fixed TPS 1000, image ...)`. +The missions perform the binary search to find the minimum sustainable block time. Upon completion, the missions emit a log line of the form `Minimum sustainable block time: 4000 ms (fixed TPS 1000, image ...)`. ## SLA: pass/fail criteria @@ -18,9 +18,11 @@ A candidate close time `T` is considered a **pass** if and only if, **on every n > **FIXME (P75 tolerance):** the intended P75 band is `[0.95·T, 1.05·T)` (±5%), but stellar-core currently has performance regressions that prevent the stricter band from being achievable under load. The tolerance has been temporarily widened to ±20% so the test can exercise the rest of the pipeline; tighten it back to ±5% (or narrower) once those regressions are fixed. -If any node violates any of these bounds, `T` is considered a **fail** and the binary search raises its lower bound. The same is true if the load run itself errors (e.g., stellar-core's internal `loadgen-run-failed` counter increments, nodes fall out of sync, or peers report inconsistent ledger hashes) — in that case the mission treats the iteration as a fail and the search continues upward. +If any node violates any of these bounds, `T` is considered a **fail** and the binary search raises its lower bound. The same is true if the load run itself errors (e.g., stellar-core's internal `loadgen-run-failed` counter increments, nodes fall out of sync, or peers report inconsistent ledger hashes) — in that case the mission treats the iteration as a fail and the search continues upward. If instead the load completes but its metrics cannot be read, the candidate was not measured, so the mission aborts rather than count it as a fail. -The search terminates when the upper and lower bounds are within 100 ms of each other. If no candidate in the range satisfied the SLA, the mission fails with `"No block time in [lo, hi] ms satisfied the SLA at TPS N"`. +Overlay-only candidates (`MinBlockTimeMixed`) never apply their transactions, but stellar-core's load generator still completes only once every transaction it submitted has been included in a closed ledger, so a load generator failure (such as load left out of ledgers) fails the candidate. That needs stellar-core master, or a Rust-overlay image built from 2026-07-31 on (earlier ones never count the transactions as included, so every overlay-only candidate fails). These candidates are judged while still in overlay-only mode: in addition to the close-time SLA above, the consistency and sync checks must pass and, as a cross-check after a completed load, every node must report the same `ledger.transaction.count` total, covering every offered transaction (a node still closing the last loaded ledger gets up to two close times to catch up). Otherwise, for example when a node fell further behind, the candidate fails. Apply is not re-enabled; the nodes are restarted, discarding what is left in their queues, before the next candidate. + +Candidate close times are always whole seconds: the search runs over the whole seconds in `[--min-block-time-ms, --max-block-time-ms]`, bounds included, so the default range evaluates `4000` ms and, only if that fails, `5000` ms. Equal bounds evaluate exactly that close time once, without rounding. If no candidate in the range satisfied the SLA, the mission fails with `"No block time in [lo, hi] ms satisfied the SLA at TPS N"`. ## Docker images with performance tests enabled @@ -38,9 +40,10 @@ These parameters affect both `MinBlockTimeClassic` and `MinBlockTimeMixed` missi * `--tx-rate`: The fixed transaction rate (TPS) used for every iteration of the search. For `MinBlockTimeMixed`, this is used only when neither `--classic-tx-rate` nor `--soroban-tx-rate` is set, in which case the mission splits it evenly between classic and Soroban streams. The mission answers the question "what is the smallest block time the network can sustain at this TPS?" so choosing a TPS the network clearly cannot sustain (e.g., above the network's max TPS at default block time) will result in the mission failing with no block time satisfying the SLA. * `--min-block-time-ms`: Binary search lower bound, in milliseconds. Defaults to `4000`. -* `--max-block-time-ms`: Binary search upper bound, in milliseconds. Defaults to `5000`, which is also the protocol's maximum allowed ledger target close time — setting this higher will cause the mission to fail at startup, since validators reject upgrades above the protocol cap. Must be strictly greater than `--min-block-time-ms`. +* `--max-block-time-ms`: Binary search upper bound, in milliseconds. Defaults to `5000`, which is also the protocol's maximum allowed ledger target close time — setting this higher will cause the mission to fail at startup, since validators reject upgrades above the protocol cap. Must not be less than `--min-block-time-ms`; when the two are equal the mission evaluates exactly that close time once, with no search. * `--num-pregenerated-txs`: Number of pre-generated signed classic transactions to create per loadgen node. `MinBlockTimeClassic` uses these on small networks (≤30 nodes) when it automatically switches classic payment load to `PayPregenerated`; `MinBlockTimeMixed` always uses them for the classic stream in its `MIXED_PREGEN_*` mode. Defaults to `2500000` * `--pubnet-data`: Network topology to use. Defaults to a topology of tier 1 validators. See [Specifying network topologies](#specifying-network-topologies) for details on how to specify a custom topology. +* `--tier-1-orgs-to-add`: Organizations (three validators each) to add to that default tier 1 topology of 10 organizations, from `0` (the default) to `30`. They are added in a fixed order (`x01`, `x02`, ...), spread over further cloud regions in North America, Europe, Asia, South America, Oceania, Africa and the Middle East; the simulated network delay between two validators grows with their distance, so larger counts also raise the network's latency floor. With `--pubnet-data`, the flag instead adds tier 1 organizations to that topology, as in `SimulatePubnet`. * `--netdelay-image`: Helper image providing simulated network delay for latency simulation. SDF provides a public image on dockerhub at `stellar/sdf-netdelay`. ### Additional options for mixed pre-generated classic and synthetic Soroban traffic @@ -51,6 +54,8 @@ In addition to the parameters in the previous section, `MinBlockTimeMixed` suppo * `--classic-tx-rate`: Classic payment TPS for the pre-generated classic stream. * `--soroban-tx-rate`: Soroban TPS for the selected synthetic Soroban stream. +The load runs on every validator of the load-generating organizations, each with its own slice of the genesis accounts, at an equal share of the offered rate for the whole load window. + If neither stream-specific TPS is set, `--tx-rate` is split evenly between classic and Soroban traffic. If either stream-specific TPS is set, any omitted stream defaults to `0`, and the fixed TPS for the mission is the sum of `--classic-tx-rate` and `--soroban-tx-rate`. Before enabling overlay-only mode, the mission upgrades classic max tx set size from `--classic-tx-rate * 15`, and Soroban network limits from the selected transaction type's per-transaction resources multiplied by `--soroban-tx-rate * 15` (~15 seconds of throughput as leeway). diff --git a/src/App/Program.fs b/src/App/Program.fs index 037ca63b..1d3138ba 100644 --- a/src/App/Program.fs +++ b/src/App/Program.fs @@ -501,7 +501,7 @@ type MissionOptions member self.SimulateApplyWeight = simulateApplyWeight [] member self.Tier1OrgsToAdd = tier1OrgsToAdd @@ -965,6 +965,7 @@ let main argv = minBlockTimeMixedMode = mission.MinBlockTimeMixedMode minBlockTimeMixedClassicTxRate = mission.MinBlockTimeMixedClassicTxRate minBlockTimeMixedSorobanTxRate = mission.MinBlockTimeMixedSorobanTxRate + pregenerateTxsPerValidator = false runForMinBlockTime = false forceOldStyleTriggerTimerPct = mission.ForceOldStyleTriggerTimerPct uniformDrift = List.ofSeq mission.UniformDrift diff --git a/src/FSLibrary.Tests/Tests.fs b/src/FSLibrary.Tests/Tests.fs index f19d06a3..41ee961c 100644 --- a/src/FSLibrary.Tests/Tests.fs +++ b/src/FSLibrary.Tests/Tests.fs @@ -152,6 +152,7 @@ let ctx : MissionContext = minBlockTimeMixedMode = "mixed_pregen_sac_payment" minBlockTimeMixedClassicTxRate = None minBlockTimeMixedSorobanTxRate = None + pregenerateTxsPerValidator = false runForMinBlockTime = false forceOldStyleTriggerTimerPct = 0 uniformDrift = [] @@ -764,7 +765,9 @@ type Tests(output: ITestOutputHelper) = let nCfgWithoutSimulateApply = MakeNetworkCfg { ctx with simulateApplyWeight = None; simulateApplyDuration = None } [ coreSet ] passOpt - let cmds = nCfgWithoutSimulateApply.getInitCommands PeerSpecificConfigFile coreSet.options + let cmds = + nCfgWithoutSimulateApply.getInitCommands PeerSpecificConfigFile coreSet.options None + let cmdStr = ShAnd(cmds).ToString() let exp = @@ -773,7 +776,7 @@ type Tests(output: ITestOutputHelper) = Assert.Equal(exp, cmdStr) - let cmds = nCfg.getInitCommands PeerSpecificConfigFile coreSet.options + let cmds = nCfg.getInitCommands PeerSpecificConfigFile coreSet.options None let cmdStr = ShAnd(cmds).ToString() Assert.Equal(exp, cmdStr) @@ -1235,3 +1238,409 @@ let ``core container env is unchanged when no extra variables are given`` () = [ "STELLAR_CORE_PEER_SHORT_NAME"; "ASAN_OPTIONS" ], c.Env |> Seq.map (fun e -> e.Name) |> List.ofSeq ) + +[] +let ``Min block time candidates are the whole seconds in the range, bounds included`` () = + Assert.Equal([ 4000; 5000 ], MinBlockTimeTest.wholeSecondCandidates 4000 5000) + Assert.Equal([ 2000; 3000; 4000 ], MinBlockTimeTest.wholeSecondCandidates 1500 4999) + Assert.Equal([ 1000 ], MinBlockTimeTest.wholeSecondCandidates 0 1000) + Assert.Empty(MinBlockTimeTest.wholeSecondCandidates 1100 1900) + +[] +let ``Equal min block time bounds evaluate exactly that close time`` () = + Assert.Equal([ 4000 ], MinBlockTimeTest.closeTimeCandidates 4000 4000) + // An explicit target is not rounded to a whole second. + Assert.Equal([ 1500 ], MinBlockTimeTest.closeTimeCandidates 1500 1500) + Assert.Equal([ 4000; 5000 ], MinBlockTimeTest.closeTimeCandidates 4000 5000) + Assert.Empty(MinBlockTimeTest.closeTimeCandidates 1100 1900) + +// The search result and the candidates it evaluated, in order. +let private searchMinPassingTrace (candidates: int list) (passes: int -> bool) : int option * int list = + let evaluated = ref [] + + let result = + MinBlockTimeTest.searchMinPassing + candidates + (fun t -> + evaluated.Value <- evaluated.Value @ [ t ] + passes t) + + result, evaluated.Value + +[] +let ``Overlay-only inclusion agrees only when every node holds the same count covering the offered load`` () = + let agrees = MinBlockTimeTest.inclusionAgrees 1000 + + Assert.True( + agrees [ "a", 1000.0 + "b", 1000.0 + "c", 1000.0 ] + ) + // A node that has not closed the last loaded ledger yet. + Assert.False( + agrees [ "a", 1000.0 + "b", 1000.0 + "c", 880.0 ] + ) + // A node counting transactions the others did not. + Assert.False( + agrees [ "a", 1000.0 + "b", 1000.0 + "c", 1003.0 ] + ) + // Every node agrees, but on fewer than were offered. + Assert.False( + agrees [ "a", 900.0 + "b", 900.0 + "c", 900.0 ] + ) + + Assert.False(agrees []) + +[] +let ``Min block time search finds the smallest passing candidate`` () = + let candidates = [ 1000 .. 1000 .. 5000 ] + + for threshold in candidates do + let result, evaluated = searchMinPassingTrace candidates (fun t -> t >= threshold) + Assert.Equal(Some threshold, result) + // Binary search over 5 candidates: at most ceil(log2 6) = 3 evaluations. + Assert.InRange(evaluated.Length, 1, 3) + + Assert.Equal(None, fst (searchMinPassingTrace candidates (fun _ -> false))) + Assert.Equal(None, fst (searchMinPassingTrace [] (fun _ -> true))) + + // The default [4000, 5000] range evaluates 4000 first, then 5000 only if 4000 fails. + Assert.Equal((Some 5000, [ 4000; 5000 ]), searchMinPassingTrace [ 4000; 5000 ] (fun t -> t >= 5000)) + Assert.Equal((Some 4000, [ 4000 ]), searchMinPassingTrace [ 4000; 5000 ] (fun t -> t >= 4000)) + // A single candidate is evaluated once. + Assert.Equal((None, [ 1500 ]), searchMinPassingTrace [ 1500 ] (fun _ -> false)) + +// The shape of a MinBlockTimeMixed milestone run: --tier-1-orgs-to-add 9 gives +// 19 organizations of 3 validators, so 57 load generators. 1,200,000 genesis +// accounts do not divide evenly among 57 (1,200,000 = 57 * 21,052 + 36), so +// the tests below see both floored and spread values. +let private milestoneOrgsToAdd = 9 + +let private milestoneOrgs = 10 + milestoneOrgsToAdd +let private milestoneValidators = milestoneOrgs * 3 +let private milestoneAccounts = 1200000 +let private milestoneAccountsPerValidator = milestoneAccounts / milestoneValidators + +[] +[] +[] +[] +[] +let ``All validator partitions preserve rate duration and disjoint account slices`` (rate: int) = + let durationSec = 960 + + let full = + { LoadGen.GetDefault() with + mode = MixedPregenSACPayment + accounts = milestoneAccounts + txrate = rate + txs = rate * durationSec + classicTxRate = Some 0 + sorobanTxRate = Some rate } + + let shares = StellarStatefulSets.PartitionValidatorLoad milestoneValidators true full + Assert.Equal(milestoneValidators, shares.Length) + // Rates are spread, at most one apart, so they add up to the offered rate; + // each transaction budget is its rate for the whole duration. + Assert.Equal(rate, shares |> List.sumBy (fun s -> s.txrate)) + Assert.Equal(rate * durationSec, shares |> List.sumBy (fun s -> s.txs)) + let rates = shares |> List.map (fun s -> s.txrate) + Assert.True(List.max rates - List.min rates <= 1) + // Accounts are floored: 57 * 21,052 = 1,199,964, leaving 36 unused. + Assert.Equal(milestoneValidators * milestoneAccountsPerValidator, shares |> List.sumBy (fun s -> s.accounts)) + + shares + |> List.iteri + (fun i s -> + Assert.Equal(milestoneAccountsPerValidator, s.accounts) + Assert.Equal(i * milestoneAccountsPerValidator, s.offset) + Assert.Equal(s.txrate * durationSec, s.txs) + Assert.Equal(Some s.txrate, s.sorobanTxRate) + Assert.Equal(Some 0, s.classicTxRate) + Assert.True(s.offset + s.accounts <= full.accounts)) + + for a, b in List.pairwise shares do + Assert.True(a.offset + a.accounts <= b.offset) + +[] +let ``Partition spreads rates but floors accounts and even-split txs`` () = + // 3 generators, and every value leaves a remainder. + let full = + { LoadGen.GetDefault() with + accounts = 1000 + txrate = 100 + txs = 30001 + spikesize = 10 } + + let shares = StellarStatefulSets.PartitionValidatorLoad 3 false full + let field f = shares |> List.map f + // Rates and spike sizes are spread: the first generators each take one of + // the remainder, so the shares still add up (100 = 34 + 33 + 33). + Assert.Equal([ 34; 33; 33 ], field (fun s -> s.txrate)) + Assert.Equal([ 4; 3; 3 ], field (fun s -> s.spikesize)) + // Accounts are floored to 1000 / 3 = 333 per generator, offset by slice, + // so the last account is unused. + Assert.Equal([ 333; 333; 333 ], field (fun s -> s.accounts)) + Assert.Equal([ 0; 333; 666 ], field (fun s -> s.offset)) + // Without a fixed duration each generator gets 30001 / 3 = 10000 txs, + // floored too. + Assert.Equal([ 10000; 10000; 10000 ], field (fun s -> s.txs)) + +[] +let ``Fixed duration partition spreads classic and Soroban rates separately`` () = + let full = + { LoadGen.GetDefault() with + accounts = 900 + txrate = 15 + txs = 15 * 300 + classicTxRate = Some 10 + sorobanTxRate = Some 5 } + + let shares = StellarStatefulSets.PartitionValidatorLoad 3 true full + let field f = shares |> List.map f + Assert.Equal([ Some 4; Some 3; Some 3 ], field (fun s -> s.classicTxRate)) + Assert.Equal([ Some 2; Some 2; Some 1 ], field (fun s -> s.sorobanTxRate)) + // Each generator's rate is the sum of its two streams, and its budget is + // that rate for the full 300 s, so the budgets add up exactly. + Assert.Equal([ 6; 5; 4 ], field (fun s -> s.txrate)) + Assert.Equal([ 1800; 1500; 1200 ], field (fun s -> s.txs)) + Assert.Equal(full.txs, shares |> List.sumBy (fun s -> s.txs)) + +[] +let ``Milestone topology pregenerates each validator's slice of its load partition`` () = + let sets = + StableApproximateTier1CoreSetsWithExtraOrgs "frozen-image" false milestoneOrgsToAdd + + Assert.Equal(milestoneOrgs, sets.Length) + Assert.Equal(milestoneValidators, sets |> List.sumBy (fun s -> s.options.nodeCount)) + + let full = + { LoadGen.GetDefault() with + mode = MixedPregenSACPayment + accounts = milestoneAccounts + txrate = 5000 + txs = 5000 * 960 + classicTxRate = Some 0 + sorobanTxRate = Some 5000 } + + let shares = StellarStatefulSets.PartitionValidatorLoad milestoneValidators true full + let keys = sets |> List.collect (fun s -> s.keys |> Array.toList) + Assert.Equal(milestoneValidators, keys |> List.map (fun k -> k.AccountId) |> Set.ofList |> Set.count) + + sets + |> List.iteri + (fun orgIndex cs -> + Assert.Equal(Some true, cs.options.tier1) + Assert.Equal(3, cs.options.nodeCount) + // Full mesh: every validator prefers the other 56. + Assert.Equal(milestoneValidators, cs.options.preferredPeersMap.Value.Count) + + for peers in cs.options.preferredPeersMap.Value.Values do + Assert.Equal(milestoneValidators - 1, peers.Length) + + match cs.options.quorumSet with + | ExplicitQuorum q -> + Assert.Equal(Some 67, q.thresholdPercent) + Assert.Equal(milestoneOrgs, q.innerQuorumSets.Length) + + for inner in q.innerQuorumSets do + Assert.Equal(Some 51, inner.thresholdPercent) + Assert.Equal(3, inner.validators.Count) + | _ -> failwith "Expected explicit organization quorum" + + // MinBlockTimeTest gives each organization the slice starting at + // its first validator's offset; PregenerationOptionsForPeer must + // then hand each validator the slice its load share uses. + let options = + { cs.options with + initialization = + { cs.options.initialization with + pregenerateTxs = + Some( + 10000, + milestoneAccountsPerValidator, + orgIndex * 3 * milestoneAccountsPerValidator + ) } } + + for i in 0 .. 2 do + let actual = PregenerationOptionsForPeer options i + let share = shares.[orgIndex * 3 + i] + Assert.Equal(Some(10000, share.accounts, share.offset), actual.initialization.pregenerateTxs)) + +[] +let ``Fixed duration partition rejects a partial second transaction budget`` () = + // Each validator's budget is its rate times the duration in whole seconds. + // 4,800,001 txs at 5,000 TPS is 960.0002 s, which no such budgets add up + // to, so the partition refuses it rather than dropping or adding a tx. + let full = + { LoadGen.GetDefault() with + accounts = milestoneAccounts + txrate = 5000 + txs = 4800001 } + + Assert.Throws + (fun () -> + StellarStatefulSets.PartitionValidatorLoad milestoneValidators true full + |> ignore) + |> ignore + +[] +let ``Load runs on every validator of each core set or on node 0 of each`` () = + let sets = + StableApproximateTier1CoreSetsWithExtraOrgs "frozen-image" false milestoneOrgsToAdd + + let selected = StellarStatefulSets.LoadgenPeerIndices true sets + Assert.Equal(milestoneValidators, selected.Length) + + Assert.Equal( + milestoneValidators, + selected + |> List.map (fun (cs, i) -> cs.keys.[i].AccountId) + |> Set.ofList + |> Set.count + ) + + Assert.Equal( + [ milestoneOrgs; milestoneOrgs; milestoneOrgs ], + [ for i in 0 .. 2 -> selected |> List.filter (fun (_, j) -> i = j) |> List.length ] + ) + + let nodeZero = StellarStatefulSets.LoadgenPeerIndices false sets + Assert.Equal(milestoneOrgs, nodeZero.Length) + Assert.All(nodeZero, (fun (_, i) -> Assert.Equal(0, i))) + +[] +let ``--tier-1-orgs-to-add extends the synthetic tier 1 topology with diverse organizations`` () = + let orgs (sets: CoreSet list) = sets |> List.map (fun cs -> cs.name.StringName) |> List.sort + let locsOf (sets: CoreSet list) name = (sets |> List.find (fun cs -> cs.name.StringName = name)).options.nodeLocs + let baseSets = StableApproximateTier1CoreSets "img" false + Assert.Equal(10, baseSets.Length) + Assert.Equal(orgs baseSets, orgs (StableApproximateTier1CoreSetsWithExtraOrgs "img" false 0)) + + // Organizations are added in tier1ExtraOrgs order with their own + // locations, and never change the base ones. + let twelve = StableApproximateTier1CoreSetsWithExtraOrgs "img" false 2 + Assert.Equal(List.sort (orgs baseSets @ [ "x01"; "x02" ]), orgs twelve) + Assert.Equal(Some [ Dublin; Amsterdam; Toronto ], locsOf twelve "x01") + + for org in orgs baseSets do + Assert.Equal(locsOf baseSets org, locsOf twelve org) + + // All 30: 40 organizations of 3 validators, no two extra ones at the same + // three locations. + let all = StableApproximateTier1CoreSetsWithExtraOrgs "img" false 30 + Assert.Equal(40, all.Length) + Assert.All(all, (fun cs -> Assert.Equal(3, cs.options.nodeCount))) + Assert.All(all, (fun cs -> Assert.Equal(3, cs.options.nodeLocs.Value.Length))) + + let extraLocSets = + all + |> List.filter (fun cs -> cs.name.StringName.StartsWith "x") + |> List.map (fun cs -> List.sort cs.options.nodeLocs.Value) + + Assert.Equal(30, extraLocSets |> List.distinct |> List.length) + + // The extra organizations reach regions with no base validator: Oceania, + // South America, Africa, the Middle East and more of Asia. + let locsIn (sets: CoreSet list) = sets |> List.collect (fun cs -> cs.options.nodeLocs.Value) + + for region in [ Sydney + Auckland + SaoPaulo + BuenosAires + CapeTown + Lagos + Dubai + TelAviv + Tokyo + Mumbai ] do + Assert.DoesNotContain(region, locsIn baseSets) + Assert.Contains(region, locsIn all) + + // The quorum grows with the organizations: one inner set per organization. + match all.Head.options.quorumSet with + | ExplicitQuorum q -> Assert.Equal(40, q.innerQuorumSets.Length) + | _ -> failwith "Expected explicit organization quorum" + + Assert.ThrowsAny(fun () -> StableApproximateTier1CoreSetsWithExtraOrgs "img" false -1 |> ignore) + |> ignore + + Assert.ThrowsAny(fun () -> StableApproximateTier1CoreSetsWithExtraOrgs "img" false 31 |> ignore) + |> ignore + +[] +let ``MIXED_PREGEN runs use at most one generator per requested TPS`` () = + let sets = StableApproximateTier1CoreSets "img" false + let count tps = (MinBlockTimeTest.activeLoadGenCoreSets tps sets).Length + // Load runs on every validator, so each organization brings 3 generators: + // 10 TPS fits 3 organizations (9 generators), not 4 (12). + Assert.Equal(10, count 5000) + Assert.Equal(10, count 30) + Assert.Equal(3, count 10) + // Always at least one organization, even below its 3 generators. + Assert.Equal(1, count 2) + Assert.Equal(1, count 0) + +[] +let ``MinBlockTime runs every MIXED_PREGEN mode, and only those, on every validator`` () = + // MinBlockTimeTest loads every validator exactly when isMixedPregenMode. + for mode in [ MixedPregenSACPayment; MixedPregenOZTokenTransfer; MixedPregenSoroswapSwap ] do + Assert.True(isMixedPregenMode mode) + + for mode in [ GeneratePaymentLoad; PayPregenerated ] do + Assert.False(isMixedPregenMode mode) + +[] +let ``Validators pregenerate their own account slices only when the MinBlockTime load runs on every validator`` () = + let pregenOpts = + { coreSetOptions with + nodeCount = 3 + initialization = { coreSetOptions.initialization with pregenerateTxs = Some(1000, 100, 0) } } + + let pregenSet = MakeLiveCoreSet "pregen" pregenOpts + + let nCfg (c: MissionContext) = + MakeNetworkCfg { c with installNetworkDelay = Some false } [ pregenSet ] passOpt + + let script (c: MissionContext) = + let pod = (nCfg c).ToPodTemplateSpec pregenSet + + let core = + pod.Spec.Containers + |> Seq.find (fun k -> k.Name = CfgVal.stellarCoreContainerName "run") + + String.concat " " core.Args + + // A MinBlockTime context whose mixed mode is only the option default (e.g. + // MinBlockTimeClassic) keeps one shared pregeneration command. + let shared = script { ctx with runForMinBlockTime = true } + Assert.Contains("'--offset 0'", shared) + Assert.DoesNotContain("'--offset 100'", shared) + let perValidatorCtx = { ctx with runForMinBlockTime = true; pregenerateTxsPerValidator = true } + let perValidator = script perValidatorCtx + + for offset in [ 0; 100; 200 ] do + Assert.Contains(sprintf "'--offset %d'" offset, perValidator) + + // One branch per validator inside the init chain, so a failed + // pregeneration still keeps stellar-core from starting, and a pod that + // matches no validator fails instead of starting without its txs. + let names = Array.init 3 (fun i -> ((nCfg perValidatorCtx).PodName pregenSet i).StringName) + + let cmds = + (nCfg perValidatorCtx).getInitCommands PeerSpecificConfigFile pregenSet.options (Some names) + + let branches = + cmds + |> Array.choose + (function + | ShIf (_, _, elifs, Some otherwise) -> Some(elifs.Length, otherwise.ToString()) + | _ -> None) + + Assert.Equal<(int * string) array>([| (2, "false") |], branches) diff --git a/src/FSLibrary/MaxTPSTest.fs b/src/FSLibrary/MaxTPSTest.fs index 5dbb101e..66d0ee73 100644 --- a/src/FSLibrary/MaxTPSTest.fs +++ b/src/FSLibrary/MaxTPSTest.fs @@ -129,9 +129,10 @@ let maxTPSTest (context: MissionContext) (baseLoadGen: LoadGen) (setupCfg: LoadG if context.pubnetData.IsSome then FullPubnetCoreSets context true false else - StableApproximateTier1CoreSets + StableApproximateTier1CoreSetsWithExtraOrgs context.image (if context.flatQuorum.IsSome then context.flatQuorum.Value else false) + context.tier1OrgsToAdd // PayPregenerated requires node restart between failed iterations to ensure validity of the pregenerated transactions // However, large-scale simulation restarts can be slow, so for now only use the new mode on small networks diff --git a/src/FSLibrary/MinBlockTimeTest.fs b/src/FSLibrary/MinBlockTimeTest.fs index 0ec8f8e2..14cc501b 100644 --- a/src/FSLibrary/MinBlockTimeTest.fs +++ b/src/FSLibrary/MinBlockTimeTest.fs @@ -22,7 +22,37 @@ open StellarSupercluster let private smallNetworkSize = 10 -let private searchThresholdMs = 100 +// Candidate close times are always whole seconds: sub-second values hit a +// known rounding issue in close-time handling, so the search must never +// propose one. The candidates are the whole seconds in [minMs, maxMs], bounds +// included, so the default [4000, 5000] range evaluates 4000 and, only if that +// fails, 5000. Exposed for unit tests. +let wholeSecondCandidates (minMs: int) (maxMs: int) : int list = + let first = max 1000 (((minMs + 999) / 1000) * 1000) + let last = (maxMs / 1000) * 1000 + [ first .. 1000 .. last ] + +// The close times a run evaluates. Equal bounds name one explicit target, such +// as a milestone latency, which is evaluated exactly once as given, in +// milliseconds and without whole-second rounding; otherwise the whole seconds +// in the range. Exposed for unit tests. +let closeTimeCandidates (minMs: int) (maxMs: int) : int list = + if minMs = maxMs then [ minMs ] else wholeSecondCandidates minMs maxMs + +// Binary search over ascending candidates for the smallest one that passes, +// assuming every candidate above a passing one passes too. Returns None when +// none passes. Exposed for unit tests. +let searchMinPassing (candidates: int list) (passes: int -> bool) : int option = + let arr = Array.ofList candidates + let mutable failIdx = -1 + let mutable passIdx = arr.Length + + while passIdx - failIdx > 1 do + let mid = (failIdx + passIdx) / 2 + + if passes arr.[mid] then passIdx <- mid else failIdx <- mid + + if passIdx < arr.Length then Some arr.[passIdx] else None // For the purposes of min block test, use high value to avoid noise from SCP timeouts let private timeout = 2000 @@ -223,28 +253,103 @@ let upgradeMixedPregenSorobanLimits waitForMixedPregenSorobanLimits peer limits -let private toggleOverlayOnlyMode (formation: StellarFormation) (coreSets: CoreSet list) = +// Exposed for reuse by MissionTriggerTimerMixConsensus. +let toggleOverlayOnlyMode (formation: StellarFormation) (coreSets: CoreSet list) = formation.NetworkCfg.EachPeerInSets (List.toArray coreSets) (fun peer -> let res = peer.ToggleOverlayOnlyMode() LogInfo "Toggled overlay-only mode on %s: %s" peer.ShortName.StringName res) -// Exposed for reuse by MissionTriggerTimerMixConsensus. -let withOverlayOnlyMode (formation: StellarFormation) (coreSets: CoreSet list) (f: unit -> unit) = - LogInfo "Enabling overlay-only mode" - toggleOverlayOnlyMode formation coreSets - - try - f () - finally - LogInfo "Disabling overlay-only mode" - toggleOverlayOnlyMode formation coreSets - let private readLedgerAgePercentiles (peer: Peer) : Peer * float * float = let h = peer.GetMetrics().LedgerAgeClosedHistogram peer, float h.``75``, float h.``99`` +// Transactions included in closed ledgers since the candidate cleared the +// metrics. ledger.transaction.count is Updated by LedgerManagerImpl even in +// overlay-only mode (it is marked before the apply-skip branch), so its sum +// counts every transaction that made it into a tx set with apply disabled. +let private readLedgerTxsIncluded (peer: Peer) : Peer * float = + peer, float (peer.GetMetrics().LedgerTransactionCount.Sum) + +// Whether every node reports the same count of transactions in its ledgers, +// covering the `offered` load. Exposed for unit tests. +let inclusionAgrees (offered: int) (included: ('node * float) list) : bool = + match included |> List.map snd |> List.distinct with + | [ txs ] -> txs >= float offered + | _ -> false + +// Every node's included-transaction count after a completed load, re-read for +// up to two close times until they agree (inclusionAgrees): a node that had +// not yet closed the last loaded ledger when first read catches up. After a +// completed load only empty ledgers close, so waiting cannot hide a mismatch. +let private collectLedgerTxsIncluded + (formation: StellarFormation) + (coreSets: CoreSet list) + (offered: int) + (targetMs: int) + : (Peer * float) list = + let readAll () = + formation.NetworkCfg.PeersInSets(List.toArray coreSets) + |> List.map (fun peer -> async { return readLedgerTxsIncluded peer }) + |> Async.Parallel + |> Async.RunSynchronously + |> Array.toList + + let clock = System.Diagnostics.Stopwatch.StartNew() + let mutable included = readAll () + + while not (inclusionAgrees offered included) + && clock.ElapsedMilliseconds < 2L * int64 targetMs do + System.Threading.Thread.Sleep 1000 + included <- readAll () + + included + +// Cross-check for overlay-only candidates after a completed load: every node +// must report the same count of transactions in its ledgers, covering the +// offered load. Loadgen completes only once every transaction its node +// submitted has been included in a closed ledger, and nodes that close the +// same ledgers count the same transactions, so a load the network carried +// agrees; this confirms it from the ledgers' own counts, which +// checkLedgerAgeSLA does not look at (closing near-empty ledgers on schedule +// passes it trivially). Comparing totals over the window, rather than txs per +// ledger against TPS x T, is not skewed by the idle ledgers around the load or +// by long ledgers carrying more. A node still behind after +// collectLedgerTxsIncluded's re-reads, or a shortfall on every node, fails the +// candidate. Returns why, or None. +let private inclusionFailure (included: (Peer * float) list) (targetMs: int) (offered: int) : string option = + if List.isEmpty included then + None + else + match included |> List.map snd |> List.distinct with + | [ txs ] when txs >= float offered -> + LogInfo + "Inclusion at T=%dms: every node's ledgers hold %.0f transactions for the %d offered -> PASS" + targetMs + txs + offered + + None + | [ txs ] -> + LogError + "Every node's ledgers hold %.0f of the %d offered transactions at T=%dms, although loadgen reported all of them included." + txs + offered + targetMs + + Some(sprintf "every node's ledgers hold %.0f of the %d offered transactions" txs offered) + | _ -> + LogError + "Nodes disagree on how many transactions their ledgers hold at T=%dms (%s; %d offered): a node fell behind, forked or did not close every ledger itself." + targetMs + (included + |> List.map (fun (peer, txs) -> sprintf "%s: %.0f" peer.ShortName.StringName txs) + |> String.concat ", ") + offered + + Some "nodes disagree on how many transactions their ledgers hold" + let private collectLedgerAgePercentiles (formation: StellarFormation) (coreSets: CoreSet list) @@ -299,6 +404,13 @@ let private checkLedgerAgeSLA (percentiles: (Peer * float * float) list) (target let p99Max = tf * 2.0 let mutable ok = true + LogInfo + "SLA thresholds at T=%dms: every peer's ledger.age.closed-histogram P75 in [%.0f, %.0f) ms and P99 <= %.0f ms (the P75 band is the temporarily widened +/-20%%; the target is a close-time target, not a P99 guarantee)" + targetMs + tLo + tHi + p99Max + let formatDeviation value = let deviation = (value - tf) / tf * 100.0 @@ -334,14 +446,29 @@ let private checkLedgerAgeSLA (percentiles: (Peer * float * float) list) (target ok +// The load-generating core sets a MIXED_PREGEN_* run uses. Its load runs on +// every validator of these sets, so they hold at most one validator per +// requested TPS (every generator gets a non-zero share), and always at least +// one set. Exposed for unit tests. +let activeLoadGenCoreSets (requestedTps: int) (loadGenNodes: CoreSet list) : CoreSet list = + let fitting = + loadGenNodes + |> List.scan (fun total (cs: CoreSet) -> total + cs.options.nodeCount) 0 + |> List.tail + |> List.takeWhile (fun total -> total <= max 1 requestedTps) + |> List.length + + List.truncate (max 1 fitting) loadGenNodes + let minBlockTimeTest (context: MissionContext) (baseLoadGen: LoadGen) (setupCfg: LoadGen option) = let allNodes = if context.pubnetData.IsSome then FullPubnetCoreSets context true false else - StableApproximateTier1CoreSets + StableApproximateTier1CoreSetsWithExtraOrgs context.image (if context.flatQuorum.IsSome then context.flatQuorum.Value else false) + context.tier1OrgsToAdd // Mirrors MaxTPSTest: on small networks, GeneratePaymentLoad runs out of // source accounts at high TPS, so switch to PayPregenerated which uses @@ -371,44 +498,47 @@ let minBlockTimeTest (context: MissionContext) (baseLoadGen: LoadGen) (setupCfg: else loadGenNodes - let isLoadGenNode cs = List.exists (fun (cs': CoreSet) -> cs' = cs) loadGenNodes + // MIXED_PREGEN_* load runs on every validator of the active core sets, + // each with its own account slice; other modes run it on node 0 of each + // load-generating core set. + let everyValidator = isMixedPregenMode baseLoadGen.mode let activeLoadGenNodes = - if isMixedPregenMode baseLoadGen.mode then + if everyValidator then let requestedCount = max (baseLoadGen.classicTxRate |> Option.defaultValue 0) (baseLoadGen.sorobanTxRate |> Option.defaultValue 0) - loadGenNodes - |> List.truncate (min (List.length loadGenNodes) (max 1 requestedCount)) + activeLoadGenCoreSets requestedCount loadGenNodes else loadGenNodes - // For pre-generated modes, partition genesis accounts evenly across - // loadgen nodes and assign offsets so each active node signs txs against - // its own slice. Mixed pregen keeps every tier1 core set initialized, but - // partitions accounts by the active loadgen count so low-TPS runs still - // have enough local accounts on the node that generates load. + let isActiveLoadGenNode cs = List.exists (fun (cs': CoreSet) -> cs' = cs) activeLoadGenNodes + + let context = { context with pregenerateTxsPerValidator = everyValidator } + + // For pre-generated modes, partition genesis accounts evenly across the + // active loadgen nodes and assign offsets so each signs txs against its own + // slice; the other core sets pregenerate nothing, so low-TPS mixed runs + // still have enough local accounts on the nodes that generate load. With + // load on every validator, each core set's slice covers all of its + // validators (StellarKubeSpecs.PregenerationOptionsForPeer splits it per + // pod). let allNodes = match context.numPregeneratedTxs, context.genesisTestAccountCount, baseLoadGen.mode with | Some txs, Some accounts, mode when usesPregeneratedTxs mode -> - let partitionCount = - if isMixedPregenMode mode then - List.length activeLoadGenNodes - else - List.length loadGenNodes - + let generators (cs: CoreSet) = if everyValidator then cs.options.nodeCount else 1 + let partitionCount = List.sumBy generators activeLoadGenNodes let accountsPerNode = accounts / partitionCount let mutable j = 0 List.map (fun (cs: CoreSet) -> let pregenerateTxs = - if isLoadGenNode cs then - let i = if isMixedPregenMode mode then j % partitionCount else j - - j <- j + 1 + if isActiveLoadGenNode cs then + let i = j + j <- j + generators cs Some(txs, accountsPerNode, accountsPerNode * i) else Some(0, 1, 0) @@ -509,42 +639,122 @@ let minBlockTimeTest (context: MissionContext) (baseLoadGen: LoadGen) (setupCfg: upgradeSorobanMaxTxSetSize targetMs formation.clearMetrics allNodes - // A loadgen failure is not an SLA signal — it usually means the - // requested TPS is too high for the network, or that loadgen - // itself lost a tx in the pipeline. Treating it as "SLA missed" - // would mislead the binary search, so fail the mission loudly. - try - if isMixedPregenMode baseLoadGen.mode then - withOverlayOnlyMode - formation - allNodes - (fun () -> formation.RunMultiLoadgen activeLoadGenNodes loadGen) - else + // MIXED_PREGEN_* candidates run in overlay-only mode. Core skips + // apply, but loadgen still completes only once every transaction + // its node submitted has been included in a closed ledger (core + // counts inclusion in this mode). Apply is never re-enabled: the + // restart before the next candidate discards whatever is left in + // the nodes' queues rather than making the network apply it at + // once, which previously pushed nodes out of sync minutes after + // the window. Everything below is therefore read in-mode. + let overlayOnly = isMixedPregenMode baseLoadGen.mode + + if overlayOnly then toggleOverlayOnlyMode formation allNodes + + // Per doc/measuring-minimum-block-time.md, each failure below (a + // loadgen failure, such as application lagging the offered load + // or a node dropping out mid-window, nodes disagreeing on the + // transactions in their ledgers, or nodes out of sync or + // inconsistent) fails the candidate, not the mission: the search + // raises its lower bound and continues. Killing the mission would + // let one bad window discard the whole search. + let loadgenFailure = + try formation.RunMultiLoadgen activeLoadGenNodes loadGen - with e -> failwithf "Loadgen failed at T=%dms; TPS might be too high (%s)" targetMs e.Message - - // Snapshot SLA metrics before consistency checks; those can take - // long enough to skew the ledger age percentiles. - let ledgerAgePercentiles = collectLedgerAgePercentiles formation allNodes + None + with e -> + LogWarn "Loadgen failed at T=%dms: %s" targetMs e.Message + Some(sprintf "load generation failed: %s" e.Message) + + // Snapshot the window's metrics before the health checks: core + // keeps closing ledgers after loadgen exits, and the checks can + // take long enough to skew the ledger age percentiles. + // + // Metrics that cannot be read after a completed load leave the + // candidate unmeasured, a harness or network problem rather than + // a verdict, so the mission aborts instead of letting the search + // move on as if it had failed. After a failed load the candidate + // has already failed, usually because a node dropped out, which + // also makes its metrics unreadable, so that is only logged. + let snapshot = + try + let percentiles = collectLedgerAgePercentiles formation allNodes + + // Only a completed load is cross-checked: a failed one has + // already failed the candidate. + let included = + if overlayOnly && loadgenFailure.IsNone then + collectLedgerTxsIncluded formation allNodes loadGen.txs targetMs + else + [] + + if context.measureE2eLatency then + logE2eLatencyMetrics formation activeLoadGenNodes + + Some(percentiles, included) + with + | e when loadgenFailure.IsNone -> + failwithf + "Could not read metrics at T=%dms after a completed load, so the candidate cannot be judged: %s" + targetMs + e.Message + | e -> + LogWarn "Could not read metrics at T=%dms after the failed load: %s" targetMs e.Message + None + + let healthFailure = + try + formation.CheckNoErrorsAndPairwiseConsistency() + formation.EnsureAllNodesInSync allNodes + None + with e -> + LogWarn "Health check failed at T=%dms: %s" targetMs e.Message + Some(sprintf "health check: %s" e.Message) + + // The close-time SLA is logged even when the candidate already + // failed, so every window reports where its close times landed. + let slaOk, missingInclusion = + match snapshot with + | Some (percentiles, included) -> + checkLedgerAgeSLA percentiles targetMs, inclusionFailure included targetMs loadGen.txs + | None -> false, None + + let failure = List.tryPick id [ loadgenFailure; missingInclusion; healthFailure ] + + let pass = failure.IsNone && slaOk - if context.measureE2eLatency then - logE2eLatencyMetrics formation activeLoadGenNodes + LogInfo + "Candidate T=%dms at %d TPS: %s%s" + targetMs + fixedTxRate + (if pass then "PASS" else "FAIL") + (match failure with + | Some reason -> sprintf " (%s)" reason + | None when not slaOk -> " (close-time SLA not met)" + | None -> "") - formation.CheckNoErrorsAndPairwiseConsistency() - formation.EnsureAllNodesInSync allNodes - checkLedgerAgeSLA ledgerAgePercentiles targetMs + pass - if context.minBlockTimeMs >= context.maxBlockTimeMs then + if context.minBlockTimeMs > context.maxBlockTimeMs then failwithf - "--min-block-time-ms=%d must be strictly less than --max-block-time-ms=%d" + "--min-block-time-ms=%d must not exceed --max-block-time-ms=%d" context.minBlockTimeMs context.maxBlockTimeMs - let mutable lo = context.minBlockTimeMs - let mutable hi = context.maxBlockTimeMs - let mutable bestPassing = None + let candidates = closeTimeCandidates context.minBlockTimeMs context.maxBlockTimeMs + + if List.isEmpty candidates then + failwithf + "No whole-second close time in [%d, %d] ms: --min-block-time-ms and --max-block-time-ms must include at least one whole second" + context.minBlockTimeMs + context.maxBlockTimeMs - LogInfo "Starting min block time search: T in [%d, %d] ms, fixed TPS = %d" lo hi fixedTxRate + LogInfo + "Starting min block time search: T in [%d, %d] ms (candidates: %s), fixed TPS = %d" + context.minBlockTimeMs + context.maxBlockTimeMs + (candidates |> List.map string |> String.concat ", ") + fixedTxRate // Restart-or-sleep between iterations. Pre-generated modes require // a full restart because the pregenerated txs have baked-in @@ -572,22 +782,30 @@ let minBlockTimeTest (context: MissionContext) (baseLoadGen: LoadGen) (setupCfg: System.Threading.Thread.Sleep(5 * 60 * 1000) formation.EnsureAllNodesInSync allNodes - let mutable needsRecovery = false - - while hi - lo > searchThresholdMs do - if needsRecovery then restartCoreSetsOrWait () - - let mid = lo + (hi - lo) / 2 - - if evaluateAt mid then - LogInfo "SLA met at T=%dms; lowering upper bound" mid - hi <- mid - bestPassing <- Some mid - needsRecovery <- false + // A ref cell rather than a mutable: evaluateCandidate is a closure, + // and F# closures cannot capture mutable locals. + let needsRecovery = ref false + + // One search step: recover from the previous candidate if needed, + // then evaluate T. + let evaluateCandidate (t: int) : bool = + if needsRecovery.Value then restartCoreSetsOrWait () + + if evaluateAt t then + LogInfo "SLA met at T=%dms; lowering upper bound" t + // Overlay-only runs never apply their transactions, so the + // nodes' state and queues are not fit for the next + // iteration: leftovers starve its upgrade-contract loadgen, + // which then fails hard. Restart after a pass as well, not + // just after a failure. + needsRecovery.Value <- isMixedPregenMode baseLoadGen.mode + true else - LogInfo "SLA not met at T=%dms; raising lower bound" mid - lo <- mid - needsRecovery <- true + LogInfo "SLA not met at T=%dms; raising lower bound" t + needsRecovery.Value <- true + false + + let bestPassing = searchMinPassing candidates evaluateCandidate match bestPassing with | Some t -> diff --git a/src/FSLibrary/MissionTriggerTimerMixConsensus.fs b/src/FSLibrary/MissionTriggerTimerMixConsensus.fs index dda69d29..366ad507 100644 --- a/src/FSLibrary/MissionTriggerTimerMixConsensus.fs +++ b/src/FSLibrary/MissionTriggerTimerMixConsensus.fs @@ -112,6 +112,16 @@ let private parseDrift (context: MissionContext) : ClockDriftDistribution = | u, [] -> failwith (sprintf "--uniform-drift requires exactly 2 values (lower,upper), got %d" u.Length) | [], b -> failwith (sprintf "--bimodal-drift requires exactly 4 values (min1,max1,min2,max2), got %d" b.Length) +let private withOverlayOnlyMode (formation: StellarFormation) (coreSets: CoreSet list) (f: unit -> unit) = + LogInfo "Enabling overlay-only mode" + toggleOverlayOnlyMode formation coreSets + + try + f () + finally + LogInfo "Disabling overlay-only mode" + toggleOverlayOnlyMode formation coreSets + let triggerTimerMixConsensus (baseContext: MissionContext) = // This mission assigns the trigger timer per node; a blanket setting for // all nodes would defeat its purpose. diff --git a/src/FSLibrary/StellarKubeSpecs.fs b/src/FSLibrary/StellarKubeSpecs.fs index 800ad348..81f82975 100644 --- a/src/FSLibrary/StellarKubeSpecs.fs +++ b/src/FSLibrary/StellarKubeSpecs.fs @@ -34,6 +34,21 @@ type ConfigOption = // container picks up a peer-specific config. | PeerSpecificConfigFile +// For MissionContext.pregenerateTxsPerValidator runs, whose load runs on every +// validator: the organization's options store its first account offset; each +// validator gets the following disjoint slice. +let PregenerationOptionsForPeer (opts: CoreSetOptions) (index: int) : CoreSetOptions = + if index < 0 || index >= opts.nodeCount then + invalidArg "index" "Invalid validator index" + + let init = opts.initialization + + let perNode = + init.pregenerateTxs + |> Option.map (fun (txs, accounts, offset) -> txs, accounts, offset + accounts * index) + + { opts with initialization = { init with pregenerateTxs = perNode } } + let CoreContainerVolumeMounts (peerOrJobNames: string array) (configOpt: ConfigOption) : V1VolumeMount array = let arr = [| V1VolumeMount(name = CfgVal.dataVolumeName, mountPath = CfgVal.dataVolumePath) |] @@ -656,7 +671,13 @@ type NetworkCfg with | None -> cfgs | Some (opts) -> Array.append cfgs [| self.JobConfigMap(opts) |] - member self.getInitCommands (configOpt: ConfigOption) (opts: CoreSetOptions) : ShCmd array = + // perValidatorPeerNames: the pods of a MissionContext.pregenerateTxsPerValidator + // core set, in validator order, when each pregenerates its own account slice. + member self.getInitCommands + (configOpt: ConfigOption) + (opts: CoreSetOptions) + (perValidatorPeerNames: string array option) + : ShCmd array = let cfgWords = cfgFileArgs configOpt InitCoreContainer let runCore args = @@ -731,17 +752,35 @@ type NetworkCfg with // we want. let newHistIgnoreError = ignoreError newHist - let pregenerate = - match init.pregenerateTxs with - | None -> None - | Some (txs, accounts, offset) -> - runCoreIf - true - [| "pregenerate-loadgen-txs" + let pregenerateTxs (txs: int, accounts: int, offset: int) = + runCore [| "pregenerate-loadgen-txs" "--count " + txs.ToString() "--accounts " + accounts.ToString() "--offset " + offset.ToString() |] + // Per validator, each pod pregenerates its own slice + // (PregenerationOptionsForPeer). The branch stays in the init chain, so + // a failure still keeps stellar-core from starting, and a pod that + // matches no validator fails instead of starting without its txs. + let pregenerate = + match init.pregenerateTxs, perValidatorPeerNames with + | None, _ -> None + | Some slice, None -> Some(pregenerateTxs slice) + | Some _, Some names -> + let branch i (name: string) = + let isPeer = + ShCmd [| ShWord.OfStr "test" + ShWord.Var CfgVal.peerNameEnvVarName + ShWord.OfStr "=" + ShWord.OfStr name |] + + isPeer, pregenerateTxs (PregenerationOptionsForPeer opts i).initialization.pregenerateTxs.Value + + match Array.mapi branch names |> List.ofArray with + | [] -> None + | (isFirst, first) :: rest -> + Some(ShCmd.ShIf(isFirst, first, Array.ofList rest, Some(ShCmd.OfStr "false"))) + let initialCatchup = runCoreIf init.initialCatchup [| "catchup"; "current/0" |] let cmds = @@ -828,7 +867,7 @@ type NetworkCfg with | None -> [| CoreContainerForCommand image cfgOpt asan self.missionContext.coreEnv res command [||] [| jobName |] |] | Some (opts) -> - let initCmds = self.getInitCommands cfgOpt opts + let initCmds = self.getInitCommands cfgOpt opts None let coreContainer = CoreContainerForCommand @@ -920,7 +959,11 @@ type NetworkCfg with let cfgOpt = PeerSpecificConfigFile let volumes = Array.append peerCfgVolumes [| dataVol; historyCfgVolume |] - let initCommands = self.getInitCommands cfgOpt coreSet.options + let initCommands = + self.getInitCommands + cfgOpt + coreSet.options + (if self.missionContext.pregenerateTxsPerValidator then Some peerNames else None) let runCmd = [| "run" |] diff --git a/src/FSLibrary/StellarMissionContext.fs b/src/FSLibrary/StellarMissionContext.fs index f5565d52..a994f078 100644 --- a/src/FSLibrary/StellarMissionContext.fs +++ b/src/FSLibrary/StellarMissionContext.fs @@ -165,6 +165,11 @@ type MissionContext = minBlockTimeMixedMode: string minBlockTimeMixedClassicTxRate: int option minBlockTimeMixedSorobanTxRate: int option + // Set by MinBlockTimeTest for MIXED_PREGEN_* load: each validator + // pregenerates transactions for its own account slice + // (StellarKubeSpecs.PregenerationOptionsForPeer), and RunMultiLoadgen + // runs the load on every validator rather than node 0 of each core set. + pregenerateTxsPerValidator: bool runForMinBlockTime: bool forceOldStyleTriggerTimerPct: int uniformDrift: int list diff --git a/src/FSLibrary/StellarNetworkData.fs b/src/FSLibrary/StellarNetworkData.fs index cffd4617..f3a47beb 100644 --- a/src/FSLibrary/StellarNetworkData.fs +++ b/src/FSLibrary/StellarNetworkData.fs @@ -147,6 +147,44 @@ let Taipei = { lat = 25.0329; lon = 121.5654 } // of nodes. let Tokyo = { lat = 35.6895; lon = 139.69171 } +// Further cloud and hosting regions, used only by the synthetic organizations +// StableApproximateTier1CoreSetsWithExtraOrgs adds beyond its base 10 (not by +// pubnet simulations, which draw from `locations` below). +let Amsterdam = { lat = 52.3676; lon = 4.9041 } // Azure westeurope, many DCs + +let Atlanta = { lat = 33.749; lon = -84.388 } +let Auckland = { lat = -36.8485; lon = 174.7633 } // AWS ap-southeast-6 +let Bahrain = { lat = 26.0667; lon = 50.5577 } // AWS me-south-1 +let Bogota = { lat = 4.711; lon = -74.0721 } +let BuenosAires = { lat = -34.6037; lon = -58.3816 } +let CapeTown = { lat = -33.9249; lon = 18.4241 } // AWS af-south-1 +let Chicago = { lat = 41.8781; lon = -87.6298 } +let Dallas = { lat = 32.7767; lon = -96.797 } // GCP us-south1 +let Dubai = { lat = 25.2048; lon = 55.2708 } // Azure uaenorth +let Dublin = { lat = 53.3498; lon = -6.2603 } // AWS eu-west-1 +let Jakarta = { lat = -6.2088; lon = 106.8456 } // AWS ap-southeast-3 +let Johannesburg = { lat = -26.2041; lon = 28.0473 } // Azure southafricanorth +let Lagos = { lat = 6.5244; lon = 3.3792 } +let LosAngeles = { lat = 34.0522; lon = -118.2437 } // GCP us-west2 +let Madrid = { lat = 40.4168; lon = -3.7038 } // GCP europe-southwest1 +let Melbourne = { lat = -37.8136; lon = 144.9631 } // GCP australia-southeast2 +let MexicoCity = { lat = 19.4326; lon = -99.1332 } +let Miami = { lat = 25.7617; lon = -80.1918 } +let Milan = { lat = 45.4642; lon = 9.19 } // AWS eu-south-1 +let Mumbai = { lat = 19.076; lon = 72.8777 } // AWS ap-south-1 +let Nuremberg = { lat = 49.4521; lon = 11.0767 } // Hetzner +let Osaka = { lat = 34.6937; lon = 135.5023 } // AWS ap-northeast-3 +let Paris = { lat = 48.8566; lon = 2.3522 } // AWS eu-west-3 +let Santiago = { lat = -33.4489; lon = -70.6693 } // GCP southamerica-west1 +let SanJose = { lat = 37.3382; lon = -121.8863 } +let Seoul = { lat = 37.5665; lon = 126.978 } // AWS ap-northeast-2 +let Stockholm = { lat = 59.3293; lon = 18.0686 } // AWS eu-north-1 +let Sydney = { lat = -33.8688; lon = 151.2093 } // AWS ap-southeast-2 +let TelAviv = { lat = 32.0853; lon = 34.7818 } // GCP me-west1 +let Toronto = { lat = 43.6532; lon = -79.3832 } // GCP northamerica-northeast2 +let Warsaw = { lat = 52.2297; lon = 21.0122 } // GCP europe-central2, OVH +let Zurich = { lat = 47.3769; lon = 8.5417 } // GCP europe-west6 + let locations = [ Ashburn Brussels @@ -955,9 +993,49 @@ let TestnetCoreSetOptions (image: string) = initialization = { CoreSetInitialization.Default with waitForConsensus = true } dumpDatabase = false } +// Synthetic organizations, in the order they are added, that extend +// StableApproximateTier1CoreSets past its 10 organizations (--tier-1-orgs-to-add). +// Most spread their three validators over two regions, like real operators. +// Every block of additions keeps the mix at roughly one third North America, +// one third Europe and one third elsewhere (Asia, South America, Oceania, +// Africa, Middle East). The base 10 use 13 distinct locations; with 19 +// organizations there are 37, with all 40 there are 51. +let tier1ExtraOrgs : (string * GeoLoc list) list = + [ ("x01", [ Dublin; Amsterdam; Toronto ]) + ("x02", [ SaoPaulo; Miami; Madrid ]) + ("x03", [ Tokyo; Seoul; SanJose ]) + ("x04", [ Mumbai; Bahrain; Frankfurt ]) + ("x05", [ Sydney; Singapore; LosAngeles ]) + ("x06", [ Paris; Stockholm; Chicago ]) + ("x07", [ Dallas; Atlanta; Warsaw ]) + ("x08", [ CapeTown; Johannesburg; Purfleet ]) + ("x09", [ Zurich; Milan; Ashburn ]) + ("x10", [ MexicoCity; Dallas; Bogota ]) + ("x11", [ Osaka; Taipei; Portland ]) + ("x12", [ Nuremberg; Falkenstein; Helsinki ]) + ("x13", [ Jakarta; Singapore; HongKong ]) + ("x14", [ Toronto; Beauharnois; Dublin ]) + ("x15", [ Santiago; BuenosAires; Miami ]) + ("x16", [ TelAviv; Dubai; Milan ]) + ("x17", [ Melbourne; Sydney; Auckland ]) + ("x18", [ Chicago; CouncilBluffs; Frankfurt ]) + ("x19", [ Lagos; Paris; Ashburn ]) + ("x20", [ SanJose; LosAngeles; Tokyo ]) + ("x21", [ Amsterdam; Warsaw; Stockholm ]) + ("x22", [ Atlanta; Clifton; Madrid ]) + ("x23", [ Chennai; Mumbai; Singapore ]) + ("x24", [ Columbus; Chicago; Brussels ]) + ("x25", [ Zurich; Frankfurt; Seoul ]) + ("x26", [ Portland; SanJose; Dublin ]) + ("x27", [ Purfleet; Paris; SaoPaulo ]) + ("x28", [ Dallas; MexicoCity; Toronto ]) + ("x29", [ Helsinki; Stockholm; Osaka ]) + ("x30", [ Miami; Ashburn; Nuremberg ]) ] + // This coreset is a synthetic approximation of the Tier1 group, intended -// to be used in stable benchmarks rather than experiments. -let StableApproximateTier1CoreSets (image: string) (flatQuorum: bool) : CoreSet list = +// to be used in stable benchmarks rather than experiments. Its 10 +// organizations are followed by the first extraOrgs of tier1ExtraOrgs. +let StableApproximateTier1CoreSetsWithExtraOrgs (image: string) (flatQuorum: bool) (extraOrgs: int) : CoreSet list = let allOrgs : Map = Map.ofList [ ("bd", [ Brussels; CouncilBluffs; Taipei ]) ("ct", [ Frankfurt; Purfleet; Brussels ]) @@ -970,6 +1048,14 @@ let StableApproximateTier1CoreSets (image: string) (flatQuorum: bool) : CoreSet ("rg", [ Brussels; Frankfurt; Brussels ]) ("sdf", [ Ashburn; Ashburn; Ashburn ]) ] + if extraOrgs < 0 || extraOrgs > List.length tier1ExtraOrgs then + failwithf + "--tier-1-orgs-to-add must be between 0 and %d for the synthetic tier 1 topology, got %d" + (List.length tier1ExtraOrgs) + extraOrgs + + let allOrgs = Map.ofList (Map.toList allOrgs @ List.truncate extraOrgs tier1ExtraOrgs) + let allOrgPairs = Map.toList allOrgs let orgKeys _ nodes = List.map (fun _ -> KeyPair.Random()) nodes let allKeys = Map.map orgKeys allOrgs @@ -1028,3 +1114,6 @@ let StableApproximateTier1CoreSets (image: string) (flatQuorum: bool) : CoreSet options = coreSetOpts } List.map orgCoreSet allOrgPairs + +let StableApproximateTier1CoreSets (image: string) (flatQuorum: bool) : CoreSet list = + StableApproximateTier1CoreSetsWithExtraOrgs image flatQuorum 0 diff --git a/src/FSLibrary/StellarStatefulSets.fs b/src/FSLibrary/StellarStatefulSets.fs index 4127b00d..eebed04f 100644 --- a/src/FSLibrary/StellarStatefulSets.fs +++ b/src/FSLibrary/StellarStatefulSets.fs @@ -70,6 +70,50 @@ let private getAveragePeerCount (topology: Map) : float = let nodeCount = Map.count topology if nodeCount > 0 then float total / float nodeCount else 0.0 +// Divide aggregate rate, account space and duration across actual generators. +// In fixed-duration mode (load on every validator), derive each transaction +// budget from its assigned rate; equal transaction budgets would make +// remainder-rate peers finish early. Otherwise every generator gets an equal +// share of the transactions. +let PartitionValidatorLoad (n: int) (fixedDuration: bool) (full: LoadGen) : LoadGen list = + if n <= 0 then invalidArg "n" "Need at least one generator" + + if fixedDuration && full.accounts < n then + invalidArg "n" "Need accounts for every generator" + + if fixedDuration && (full.txrate <= 0 || full.txs % full.txrate <> 0) then + invalidArg "full" "Fixed-duration load must have an integral duration" + + let fraction value i = value / n + (if i < value % n then 1 else 0) + + [ for i in 0 .. n - 1 do + let classic = full.classicTxRate |> Option.map (fun v -> fraction v i) + let soroban = full.sorobanTxRate |> Option.map (fun v -> fraction v i) + + let rate = + match classic, soroban with + | Some c, Some s -> c + s + | Some c, None -> c + | None, Some s -> s + | None, None -> fraction full.txrate i + + yield + { full with + accounts = full.accounts / n + offset = (full.accounts / n) * i + txs = if fixedDuration then rate * (full.txs / full.txrate) else full.txs / n + spikesize = fraction full.spikesize i + txrate = rate + classicTxRate = classic + sorobanTxRate = soroban } ] + +// A single generator selection is shared by launch and final submission +// accounting: every validator of each core set when everyValidator +// (MissionContext.pregenerateTxsPerValidator), else node 0 of each. +let LoadgenPeerIndices (everyValidator: bool) (coreSets: CoreSet list) : (CoreSet * int) list = + coreSets + |> List.collect (fun cs -> [ for i in 0 .. (if everyValidator then cs.options.nodeCount - 1 else 0) -> cs, i ]) + // Pure check of an observed stellar-core pod -> worker-node mapping for // --one-stellar-core-per-host: every pod in mustBeScheduled is listed exactly // once and on a node, and no two scheduled pods share a node. Pods not yet @@ -280,6 +324,10 @@ let autoscalerView type StellarFormation with + member self.LoadGenPeers(coreSets: CoreSet list) : Peer list = + LoadgenPeerIndices self.NetworkCfg.missionContext.pregenerateTxsPerValidator coreSets + |> List.map (fun (cs, i) -> self.NetworkCfg.GetPeer cs i) + member self.GetCoreSetForStatefulSet(ss: V1StatefulSet) = List.find (fun cs -> (self.NetworkCfg.StatefulSetName cs).StringName = ss.Name()) self.NetworkCfg.CoreSetList @@ -782,40 +830,13 @@ type StellarFormation with |> ignore // This is similar to RunLoadgen but runs a 1/N fractional portion of a - // given LoadGen on node 0 of each of N CoreSets. + // given LoadGen on each of N generators (LoadGenPeers: node 0 of each + // CoreSet, or every validator of each when each pregenerated its own + // account slice). member self.RunMultiLoadgen (coreSets: CoreSet list) (fullLoadGen: LoadGen) = - let n = List.length coreSets - - let fractionalLoadGen (i: int) : LoadGen = - // Spread remainder across the first r nodes instead of dumping it - // entirely on the last one, so per-node load stays even. - let getFraction attr = - let q = attr / n - let r = attr % n - if i < r then q + 1 else q - - let getOptionalFraction attr = Option.map getFraction attr - - let classicShare = getOptionalFraction fullLoadGen.classicTxRate - let sorobanShare = getOptionalFraction fullLoadGen.sorobanTxRate - - // In mixed mode derive txrate from the component rates so that - // txrate, classicTxRate, and sorobanTxRate stay consistent on - // every peer (independent splits would diverge by 1 per slice). - let txrateShare = - match classicShare, sorobanShare with - | Some c, Some s -> c + s - | Some c, None -> c - | None, Some s -> s - | None, None -> getFraction fullLoadGen.txrate - - { fullLoadGen with - accounts = fullLoadGen.accounts / n - txs = fullLoadGen.txs / n - spikesize = getFraction fullLoadGen.spikesize - txrate = txrateShare - classicTxRate = classicShare - sorobanTxRate = sorobanShare } + let everyValidator = self.NetworkCfg.missionContext.pregenerateTxsPerValidator + let loadGenPeers = self.LoadGenPeers coreSets + let shares = PartitionValidatorLoad loadGenPeers.Length everyValidator fullLoadGen let hasNonZeroRate (loadGen: LoadGen) = match loadGen.classicTxRate, loadGen.sorobanTxRate with @@ -824,22 +845,23 @@ type StellarFormation with | None, Some sorobanRate -> sorobanRate <> 0 | None, None -> true - let loadGenPeers = List.map (fun cs -> self.NetworkCfg.GetPeer cs 0) coreSets - let peerLoadGens = - loadGenPeers - |> List.indexed - |> List.map - (fun (i, peer) -> - let loadGen = fractionalLoadGen i - let offset = loadGen.accounts * i - (peer, { loadGen with offset = offset }, offset)) + List.zip loadGenPeers shares + |> List.map (fun (peer, loadGen) -> peer, loadGen, loadGen.offset) |> List.filter (fun (_, loadGen, _) -> hasNonZeroRate loadGen) if List.isEmpty peerLoadGens then failwith "Loadgen failed: no peer has a non-zero tx rate" for (peer, peerSpecificLoadgen, offset) in peerLoadGens do + LogInfo + "LOAD_SHARE peer=%s rate=%d txs=%d accounts=%d offset=%d" + peer.ShortName.StringName + peerSpecificLoadgen.txrate + peerSpecificLoadgen.txs + peerSpecificLoadgen.accounts + offset + LogInfo "Loadgen: %s with offset %d" (peer.GenerateLoad peerSpecificLoadgen) offset while List.exists