Skip to content
Dreamer-TobyPublic

About

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Topics

Resources

Stars

80 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

arXiv paper GitHub code Usage guide

Welcome to the official code repository for STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization.

Your star means a lot to us in developing this project! ⭐⭐⭐

📰 News

  • [2026/09/30] 📄 Our paper is available on arXiv.
  • [2026/09/29] 🚀 STEPQuant code is available in this repository.

👀 Overview

STEPQuant compresses the persistent recurrent states of Delta-rule models. It decides when errors will last and where they will affect the output, then spends precision accordingly.

State memory grows with concurrency, long-lived states accumulate error, and uniform low-bit quantization loses accuracy.

Fixed-size state per request still means growing memory under concurrency. Uniform low-bit quantization also feeds its errors into every subsequent state update.

🧩 Method

  • When — lifetime-aware bit allocation: Give more bits to state units with larger, longer-lived errors; protect a few high-risk units in FP16.
  • Where — key-row-aware dual-axis fitting: Scale key rows and value columns separately, prioritizing rows that matter most to the readout.

Allocation is calibrated once; packed-state kernels and asynchronous writeback keep inference efficient.

📊 Results

Seven long-generation tasks · BF16 weights with quantized states · Average acc (%)

Model FP32 INT8 INT6 STEPQuant@6 STEPQuant@4
Qwen3.8-27B 80.60 71.86 45.04 80.59 80.51
Kimi-Linear-48B-A3B-Instruct 61.52 56.02 45.70 61.47 58.52

STEPQuant lowers serving memory and state-update time for Qwen and Kimi.

In the paper's serving memory accounting, STEPQuant@6 compresses recurrent-state memory by 5.03× / 5.08× and reduces total memory by 68.7% / 53.7% on Qwen / Kimi, respectively.

Paper-reported results. Memory accounting uses W4/AWQ weights and five prefix-state slots per request (Figure 3(c)). Full-model decode throughput uses BF16 weights, batches 32–512, and 128 prompt + 1024 decode tokens (Appendix F.5). Task evaluations use NVIDIA A800 GPUs. The 4/6-bit budgets are nominal; FP16 pivots and scales add storage.

⚙️ Quick Start

Run from the repository root. The tested stack uses SGLang 0.5.12, PyTorch 2.11.0, and Triton 3.7.1. You need model checkpoints, a CUDA toolkit, and four GPUs for this Qwen example.

1. Reuse environments and set model_path in configs/models.json.

bash scripts/reuse_sglang_env.sh /path/to/sglang-env
bash scripts/reuse_eval_env.sh .venv-serving
bash scripts/reuse_calibration_env.sh /path/to/qwen-env qwen

If the serving helper prints an export CUDA_HOME=... command, run it.

2. Calibrate once, then serve.

.venv-eval/bin/python -m stepquant.reproduction \
  --models qwen --formats stepquant6 --devices 0,1,2,3

CUDA_VISIBLE_DEVICES=0,1,2,3 .venv-serving/bin/python -m stepquant.sglang \
  --model-path /path/to/Qwen3.8-27B --tp-size 4 \
  --state-format stepquant6 --plan artifacts/reproduction/plans/qwen-stepquant6.pt \
  --attention-backend triton --reasoning-parser qwen3 \
  --context-length 131072 --chunked-prefill-size 2048 \
  --mem-fraction-static 0.85 --served-model-name stepquant \
  --host 127.0.0.1 --port 31080

3. Evaluate. Stop the manually launched server first; the evaluation suite starts its own.

.venv-eval/bin/python -m stepquant.evaluation.suite \
  --models qwen --benchmarks long --formats fp32 stepquant6 \
  --server-python .venv-serving/bin/python --devices 0,1,2,3 --tp 4

For Kimi setup, the zero-shot short-task protocol (2048 output tokens, thinking disabled), and paired FP32 decode benchmarks, see the usage guide.

📂 Contact

If you have further questions, please open an issue or contact yaobingchen0515@gmail.com or xuhb2001@gmail.com.

Discussions and potential collaborations are also welcome.

🧠 Related Work

More of our work on model quantization:

About

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Topics

Resources

Stars

80 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages