Deterministic, bare-metal C++20 market-making core coupling drift-aware stochastic inventory control with SIMD joint-portfolio optimization instrumented down to hardware counters and deep-tail latency.
Modern electronic market-making must post statistically sound quotes in hundreds of nanoseconds while respecting portfolio-level capacity constraints across dozens of assets. This repository implements and measures such a system end to end: (i) a Cartea–Jaimungal drift-aware reservation-price engine with online Bayesian intensity estimation, (ii) a fixed-iteration AVX2/FMA and AVX-512 ADMM solver for the 8–64 asset joint-pricing QP, (iii) wait-free single-producer transport (SPSC ring, seqlock-distributed shadow prices, AF_XDP kernel-bypass ingest), and (iv) a measurement methodology capable of resolving nanosecond-scale tails across 10⁷ consecutive quotes. The central architectural result is the decoupling of optimization from quotation: ADMM executes on a dedicated core and publishes dual variables through a lock-free seqlock, so the critical path performs only reservation pricing, Bayesian observation, and an optimistic seqlock read and records zero ADMM-attributable stalls over 10⁷ ticks.
- Stochastic control plane (
stochastic_controller): closed-form CJ drift annuity with branchless κ→0 asymptotics, Poisson λ(δ)=A·e^(−kδ) execution model, Gamma-conjugate rate filter plus Laplace–Newton shape update, and EWMA order-flow drift all FMA-vectorized, heap-free,noexcepton the hot path. - Deterministic joint optimizer (
admm_solver): consensus ADMM with compile-time trip counts (32 outer × 10 bisection), branchless box+halfspace projection, stack-residentalignas(64)iterates, explicit instantiations at N ∈ {8, 16, 32, 64}. - Wait-free systems fabric: 64-byte-isolated SPSC ring with
MAP_HUGETLB/aligned backing options; single-publisher/multi-reader seqlock (distributed_state) with bounded-retry optimistic reads andshm_open/mmapmulti-process sharing; AF_XDP zero-copy ingest with a portablerecvmmsgfallback. - Measurement science: cycle-fenced (
__rdtscp+_mm_lfence) 10M-event harnesses on pre-faulted 2 MB hugepages, TSC-vs-steady_clockcalibration, PMU capture (perf stat/c2c), vector-codegen audit (objdump), and log-binned ASCII CDFs with full outlier provenance.
NIC queue → AF_XDP UMEM (2 MB hugepages) → zero-copy LE parser
→ SpscRingBuffer<MarketTick> (Vyukov, 64 B-isolated cursors)
→ quoting thread: quote_single + observe + seqlock try_read(λ) [critical path]
⇅ DecouplingFabric (seqlock release/acquire, relaxed atomics)
→ optimizer thread: ADMM-16 QP → publish_snapshot(λ, capacity) [off path]
| Component | File(s) | Design invariant |
|---|---|---|
| Hot-path types | include/types.hpp |
Every object alignas(64), trivially copyable, sizeof % 64 == 0 |
| Controller | include/stochastic_controller.hpp, src/stochastic_controller.cpp |
Scalar + auto-vectorized batch kernels; observe() is branchless EWMA/Newton |
| ADMM solver | include/admm_solver.hpp, src/admm_solver.cpp |
Fixed counts, no early exit (no jitter), _mm256/_mm512 kernels |
| SPSC transport | include/spsc_ring_buffer.hpp |
Release/acquire only; head/tail never share a line |
| Distributed state | include/distributed_state.hpp |
Odd epoch = writing; reader retries ≤ 2 then falls back to cached λ |
| Ingest | include/xdp_receiver.hpp, src/xdp_receiver.cpp |
XDP_DRV→SKB→UDP-fallback selected at runtime; hot poll is allocation-free |
| Storm bench | benchmarks/main_bench.cpp |
10M producer→ring→consumer TSC experiment |
| Tail bench | benchmarks/jitter_analyzer.cpp |
10M-quote decoupled jitter experiment (this paper's Figure 1–2) |
| PMU harness | benchmarks/profile_perf.sh |
Counters, coherence, codegen audit |
Build (deterministic release contract).
-O3 -march=native -mtune=native -fno-omit-frame-pointer -flto -ffast-math -fno-math-errno,
C++20, Threads::Threads (+ rt m for the tail bench on UNIX):
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -jIsolation protocol. Bare-metal prescription is isolcpus, nohz_full,
rcu_nocbs, thread pinning via pthread_setaffinity_np, and — as root —
SCHED_FIFO priority 99 (pthread_setschedparam, graceful perror
otherwise). Timing uses __rdtscp fenced by _mm_lfence on both sides;
cycles convert via a 100 ms TSC-vs-steady_clock calibration. Sample buffers
are mmap(MAP_PRIVATE|MAP_ANONYMOUS|MAP_HUGETLB) 2 MB pages (fallback:
posix_memalign(64) + stderr notice), memset-prefaulted. Warm-up is
10⁵ iterations. Hot loops contain no allocation, formatting, or syscalls;
asm volatile("" ::: "memory") fences the timed region.
Experiments. (E1) 10M-order storm (bench): throughput, queue-inclusive
tick-to-trade percentiles, sub-100 ns→>1000 ns histogram, ADMM-16 solve stats.
(E2) PMU capture (profile_perf.sh): cycles, instructions, cache-misses,
L1-dcache-load-misses, branch-misses; perf c2c HITM coherence analysis;
ADMM TU disassembly (vfmadd*, ymm/zmm, scalar-transition audit).
(E3) 10M-quote tail (jitter_analyzer [quote_core] [opt_core]): in-place
sorted percentiles, dispersion |tᵢ − tᵢ₋₁|, σ, >1000 ns outlier provenance
(ADMM-cadence vs. migration vs. hypervisor/cache jitter), log-binned CDF.
| Metric | ns | Note |
|---|---|---|
| p50 / p90 / p99 / p99.9 | 131 / 152 / 190 / 268 | Steady-state quoting cost |
| p99.99 | 10 929 [ABOVE 400ns] |
Hypervisor preemption; see Discussion |
| min / max | 105 / 9 463 387 | Max is a single WSL2 scheduling event |
| mean ± σ | 139.77 ± 3069.03 | σ dominated by the far tail |
| Dispersion (mean / max) | 23.27 / 9 463 274 | Instantaneous ` |
| Outliers >1000 ns | 2538 (0.0254%) | ADMM-cadence stalls: 0; migrated: 0 |
CDF: 88.7% ≤ 150 ns, 99.88% ≤ 250 ns, 99.975% ≤ 1000 ns. The seqlock read succeeds on the fast path (miss rate ≈ 0); every outlier is classified as unclassified cache-miss/context-switch except cadence ticks, which perform no solver work by construction.
The experiment confirms the hypothesis and bounds the claim honestly: removing
ADMM from the critical path eliminates optimization-induced stalls entirely,
yet p99.99 remains hypervisor-bound under WSL2 virtualization (multimicrosecond
preemptions visible as isolated >1 µs events on an otherwise ~130 ns
distribution). The predicted bare-metal outcome — p99.99 < 400 ns under
isolcpus + pinned SCHED_FIFO is therefore stated as a falsifiable
follow-up, not a reported result. Threats to validity include TSC frequency
excursions (mitigated by recalibration), THP/hugepage availability, and single-
machine NUMA scope; all artifacts (percentiles, CDF, outlier table) print to
stdout for independent replication.
sudo isolcpus=2,3 nohz_full=2,3 rcu_nocbs=2,3 ./build/bench 2 3
./benchmarks/profile_perf.sh 2 3
sudo ./build/jitter_analyzer 2 3Bare-metal tail validation; multi-socket NUMA placement of the optimizer; closing the loop from published λ into reservation-price skew; formal characterization of seqlock miss rate under optimizer oversubscription.