Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DynamicPricing: Microsecond-SLA Stochastic Pricing under Joint Portfolio Constraints

Deterministic, bare-metal C++20 market-making core coupling drift-aware stochastic inventory control with SIMD joint-portfolio optimization instrumented down to hardware counters and deep-tail latency.

Abstract

Modern electronic market-making must post statistically sound quotes in hundreds of nanoseconds while respecting portfolio-level capacity constraints across dozens of assets. This repository implements and measures such a system end to end: (i) a Cartea–Jaimungal drift-aware reservation-price engine with online Bayesian intensity estimation, (ii) a fixed-iteration AVX2/FMA and AVX-512 ADMM solver for the 8–64 asset joint-pricing QP, (iii) wait-free single-producer transport (SPSC ring, seqlock-distributed shadow prices, AF_XDP kernel-bypass ingest), and (iv) a measurement methodology capable of resolving nanosecond-scale tails across 10⁷ consecutive quotes. The central architectural result is the decoupling of optimization from quotation: ADMM executes on a dedicated core and publishes dual variables through a lock-free seqlock, so the critical path performs only reservation pricing, Bayesian observation, and an optimistic seqlock read and records zero ADMM-attributable stalls over 10⁷ ticks.

Contributions

  1. Stochastic control plane (stochastic_controller): closed-form CJ drift annuity with branchless κ→0 asymptotics, Poisson λ(δ)=A·e^(−kδ) execution model, Gamma-conjugate rate filter plus Laplace–Newton shape update, and EWMA order-flow drift all FMA-vectorized, heap-free, noexcept on the hot path.
  2. Deterministic joint optimizer (admm_solver): consensus ADMM with compile-time trip counts (32 outer × 10 bisection), branchless box+halfspace projection, stack-resident alignas(64) iterates, explicit instantiations at N ∈ {8, 16, 32, 64}.
  3. Wait-free systems fabric: 64-byte-isolated SPSC ring with MAP_HUGETLB/aligned backing options; single-publisher/multi-reader seqlock (distributed_state) with bounded-retry optimistic reads and shm_open/mmap multi-process sharing; AF_XDP zero-copy ingest with a portable recvmmsg fallback.
  4. Measurement science: cycle-fenced (__rdtscp + _mm_lfence) 10M-event harnesses on pre-faulted 2 MB hugepages, TSC-vs-steady_clock calibration, PMU capture (perf stat/c2c), vector-codegen audit (objdump), and log-binned ASCII CDFs with full outlier provenance.

System architecture

NIC queue → AF_XDP UMEM (2 MB hugepages) → zero-copy LE parser
  → SpscRingBuffer<MarketTick> (Vyukov, 64 B-isolated cursors)
  → quoting thread: quote_single + observe + seqlock try_read(λ)   [critical path]
  ⇅ DecouplingFabric (seqlock release/acquire, relaxed atomics)
  → optimizer thread: ADMM-16 QP → publish_snapshot(λ, capacity)   [off path]
Component File(s) Design invariant
Hot-path types include/types.hpp Every object alignas(64), trivially copyable, sizeof % 64 == 0
Controller include/stochastic_controller.hpp, src/stochastic_controller.cpp Scalar + auto-vectorized batch kernels; observe() is branchless EWMA/Newton
ADMM solver include/admm_solver.hpp, src/admm_solver.cpp Fixed counts, no early exit (no jitter), _mm256/_mm512 kernels
SPSC transport include/spsc_ring_buffer.hpp Release/acquire only; head/tail never share a line
Distributed state include/distributed_state.hpp Odd epoch = writing; reader retries ≤ 2 then falls back to cached λ
Ingest include/xdp_receiver.hpp, src/xdp_receiver.cpp XDP_DRV→SKB→UDP-fallback selected at runtime; hot poll is allocation-free
Storm bench benchmarks/main_bench.cpp 10M producer→ring→consumer TSC experiment
Tail bench benchmarks/jitter_analyzer.cpp 10M-quote decoupled jitter experiment (this paper's Figure 1–2)
PMU harness benchmarks/profile_perf.sh Counters, coherence, codegen audit

Method

Build (deterministic release contract). -O3 -march=native -mtune=native -fno-omit-frame-pointer -flto -ffast-math -fno-math-errno, C++20, Threads::Threads (+ rt m for the tail bench on UNIX):

cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Isolation protocol. Bare-metal prescription is isolcpus, nohz_full, rcu_nocbs, thread pinning via pthread_setaffinity_np, and — as root — SCHED_FIFO priority 99 (pthread_setschedparam, graceful perror otherwise). Timing uses __rdtscp fenced by _mm_lfence on both sides; cycles convert via a 100 ms TSC-vs-steady_clock calibration. Sample buffers are mmap(MAP_PRIVATE|MAP_ANONYMOUS|MAP_HUGETLB) 2 MB pages (fallback: posix_memalign(64) + stderr notice), memset-prefaulted. Warm-up is 10⁵ iterations. Hot loops contain no allocation, formatting, or syscalls; asm volatile("" ::: "memory") fences the timed region.

Experiments. (E1) 10M-order storm (bench): throughput, queue-inclusive tick-to-trade percentiles, sub-100 ns→>1000 ns histogram, ADMM-16 solve stats. (E2) PMU capture (profile_perf.sh): cycles, instructions, cache-misses, L1-dcache-load-misses, branch-misses; perf c2c HITM coherence analysis; ADMM TU disassembly (vfmadd*, ymm/zmm, scalar-transition audit). (E3) 10M-quote tail (jitter_analyzer [quote_core] [opt_core]): in-place sorted percentiles, dispersion |tᵢ − tᵢ₋₁|, σ, >1000 ns outlier provenance (ADMM-cadence vs. migration vs. hypervisor/cache jitter), log-binned CDF.

Results (E3, WSL2 root, decoupled, N = 10⁷)

Metric ns Note
p50 / p90 / p99 / p99.9 131 / 152 / 190 / 268 Steady-state quoting cost
p99.99 10 929 [ABOVE 400ns] Hypervisor preemption; see Discussion
min / max 105 / 9 463 387 Max is a single WSL2 scheduling event
mean ± σ 139.77 ± 3069.03 σ dominated by the far tail
Dispersion (mean / max) 23.27 / 9 463 274 Instantaneous `
Outliers >1000 ns 2538 (0.0254%) ADMM-cadence stalls: 0; migrated: 0

CDF: 88.7% ≤ 150 ns, 99.88% ≤ 250 ns, 99.975% ≤ 1000 ns. The seqlock read succeeds on the fast path (miss rate ≈ 0); every outlier is classified as unclassified cache-miss/context-switch except cadence ticks, which perform no solver work by construction.

Discussion

The experiment confirms the hypothesis and bounds the claim honestly: removing ADMM from the critical path eliminates optimization-induced stalls entirely, yet p99.99 remains hypervisor-bound under WSL2 virtualization (multimicrosecond preemptions visible as isolated >1 µs events on an otherwise ~130 ns distribution). The predicted bare-metal outcome — p99.99 < 400 ns under isolcpus + pinned SCHED_FIFO is therefore stated as a falsifiable follow-up, not a reported result. Threats to validity include TSC frequency excursions (mitigated by recalibration), THP/hugepage availability, and single- machine NUMA scope; all artifacts (percentiles, CDF, outlier table) print to stdout for independent replication.

Reproducing

sudo isolcpus=2,3 nohz_full=2,3 rcu_nocbs=2,3 ./build/bench 2 3
./benchmarks/profile_perf.sh 2 3
sudo ./build/jitter_analyzer 2 3

Future work

Bare-metal tail validation; multi-socket NUMA placement of the optimizer; closing the loop from published λ into reservation-price skew; formal characterization of seqlock miss rate under optimizer oversubscription.

About

Bare-metal, ultra-low-latency C++20 dynamic pricing engine. Cartea–Jaimungal adverse selection stochastic control, lock-free Seqlock DecouplingFabric, SIMD ADMM solver, and HugePage-backed ring buffers (p99.9 < 270ns).

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages