MicroGen (microgen-llm on PyPI) is a modular, hardware-aware Large Language Model (LLM) inference research framework and experimental substrate built from scratch in PyTorch. Designed to dissect memory, latency, and throughput trade-offs under controlled hardware and workload conditions, MicroGen isolates state-of-the-art serving techniques including Physical Paged KV Allocation, Hash-Based Prefix Reuse, INT8 Weight Quantization, Multi-GPU Tensor Parallelism, Speculative Decoding, and Continuous Request Batching.
MicroGen: An empirical LLM inference research substrate isolating optimization overheads, non-monotonic composition dynamics, and hardware trade-offs behind modular execution protocols. Tested via an
$N=30$ repeated-trial evaluation protocol across NVIDIA Tesla T4 and P100 GPUs and open-weights model families (GPT-2, Qwen2.5, Llama-3.2).
Topics/Tags for GitHub: llm-inference, systems-research, pytorch, kv-cache, continuous-batching, paged-attention, tensor-parallelism, quantization, speculative-decoding, fastapi, cuda.
Note on Research Framework Positioning: MicroGen is explicitly designed as an experimental systems research substrate for isolating optimization overheads and measuring non-monotonic interaction dynamics under controlled conditions, rather than competing directly with production C++/CUDA serving engines (e.g., vLLM, TensorRT-LLM, SGLang). Beyond the empirical benchmarking substrate described in the paper, MicroGen includes a working OpenAI-compatible HTTP serving layer (
microgen/api/with FastAPI, SSE streaming, and 149 passing unit/integration tests). The paper's$N=30$ throughput/latency figures reflect direct in-process engine-level measurement (isolating model, memory, and kernel dynamics).
-
🚀 Non-Monotonic Optimization Composition: Combining individually useful optimizations (+INT8, +Paged KV, +Prefix Cache) yields a statistically significant throughput regression to
$0.96\times$ baseline ($493.6 \pm 8.6\text{ tok/s}$ ,$p_{\text{adj}} < 0.001$ ), while continuous batching scheduler overhead drops throughput to$0.76\times$ baseline ($391.8 \pm 6.2\text{ tok/s}$ ,$p_{\text{adj}} < 0.001$ ) due to cumulative Python event loop and pointer indirection overheads ($\eta_{\text{overhead}} = 24.8%$ ). -
🧠 Physical Block Paged KV Allocation: Dynamically assigns
$B_{\text{block}}=16$ token physical blocks, eliminating external contiguous-allocation memory fragmentation ($F_{\text{ext}} = 1 - \frac{\text{max contiguous block}}{\text{total free VRAM}} = 0.0%$ ) under dynamic memory pressure regimes. -
⚡ Hash-Based Prefix Cache Reuse: Implements exact longest-common-prefix (LCP) key lookup as an experimental baseline approximation of shared-prefix caching, delivering up to a
$3.91\times$ prefill TTFT speedup ($6.6 \pm 0.3\text{ ms}$ vs$25.8 \pm 1.2\text{ ms}$ ,$p_{\text{adj}} < 0.001$ ) under 100% prompt overlap at$L_{\text{prompt}}=1024$ tokens, crossing into positive speedup once prompt overlap exceeds 25%. -
🔮 Speculative Decoding Acceptance Boundaries: Characterizes acceptance rate break-even threshold
$\alpha_{\text{threshold}} = \frac{T_{\text{draft_step}}}{T_{\text{target_step}}}$ . On small target models (tiny-gpt2), draft step overhead ($4.4\text{ ms}$ ) exceeds verification savings, resulting in a throughput regression ($0.45\times$ baseline,$p_{\text{adj}} < 0.001$ ). -
🌐 Multi-GPU Tensor Parallelism: Shards linear projections across dual NVIDIA T4 GPUs (
$TP=2$ ), accelerating memory-bound decoding for GPT-2 ($124\text{M}$ ) from$8.2\text{ tok/s}$ to$14.0\text{ tok/s}$ ($1.71\times$ speedup,$p_{\text{adj}} < 0.001$ ).
🎯 Portfolio & Resume Framing (Research $\rightarrow$ Methodology $\rightarrow$ Discovery $\rightarrow$ Result)
-
LLM Inference Systems Architecture: Designed and built MicroGen, a modular PyTorch LLM inference research framework isolating memory, latency, and throughput trade-offs across CPU, CUDA, and multi-GPU (
$TP=2$ ) execution protocols. -
Non-Monotonic Composition Analysis: Discovered through an
$N=30$ repeated-trial ablation protocol that composing individually positive optimizations (+INT8, +Paged KV, +Prefix Cache) yields non-monotonic throughput degradation ($0.96\times$ baseline,$p_{\text{adj}} < 0.001$ ). -
Prefix Reuse & TTFT Acceleration: Implemented an experimental hash-based longest-common-prefix (LCP) KV cache manager achieving a
$3.91\times$ prefill TTFT speedup ($6.6\text{ ms}$ vs$25.8\text{ ms}$ ,$p_{\text{adj}} < 0.001$ ) under 100% prompt overlap. -
Continuous Batching Overhead Profiling: Built a micro-profiling harness isolating Python event loop overhead (
$\eta_{\text{overhead}} = 24.8%$ ), causally explaining continuous batching throughput regressions ($0.76\times$ baseline,$p_{\text{adj}} < 0.001$ ) in research substrates. -
Memory Modeling & Multi-GPU Acceleration: Formalized external VRAM allocation fragmentation (
$F_{\text{ext}} = 1 - \frac{\text{max contiguous block}}{\text{total free VRAM}}$ ) and sharded linear projections across dual NVIDIA T4 GPUs ($TP=2$ ), delivering a$1.71\times$ throughput speedup ($14.0\text{ tok/s}$ vs$8.2\text{ tok/s}$ ,$p_{\text{adj}} < 0.001$ ).
flowchart TD
Client[Client / HTTP Request / CLI] --> API[FastAPI OpenAI Router / CLI Entry]
API --> RateLimiter[Token Bucket Rate Limiter]
RateLimiter --> Scheduler[Continuous Batching Scheduler]
subgraph Engine Core
Scheduler --> RequestQueue[Priority Request Queue]
Scheduler --> PrefixCache[Prefix KV Cache Manager]
Scheduler --> KVCache[Paged & INT8 Quantized KV Cache]
Scheduler --> Backend[Inference Backend Interface]
end
subgraph Hardware Backends
Backend --> PyTorchBackend[PyTorch Standard Backend]
Backend --> QuantizedBackend[Quantized INT8 Backend]
Backend --> TPBackend[Tensor-Parallel Multi-GPU Backend]
end
subgraph Devices & Hardware
PyTorchBackend --> CPUDevice[CPU Hardware Device]
PyTorchBackend --> CUDADevice[NVIDIA CUDA GPU Device]
TPBackend --> MultiGPU[Multi-Rank CUDA GPUs]
end
pip install microgen-llmgit clone https://github.com/Omdeepb69/MicroGen.git
cd MicroGen
python -m venv venv
source venv/bin/activate # On Linux/macOS
# or: venv\Scripts\activate on Windows
pip install -e .import microgen
# Initialize engine with PyTorch FP32 backend
engine = microgen.LLMEngine.from_pretrained(
"sshleifer/tiny-gpt2",
backend_type="pytorch",
device="cuda"
)
# Generate completion text
output = engine.generate("MicroGen is a fast LLM inference engine", max_tokens=32)
print("Output:", output)import microgen
# Load quantized backend (INT8 weights + dynamic INT8 KV cache)
engine = microgen.LLMEngine.from_pretrained(
"sshleifer/tiny-gpt2",
backend_type="quantized",
device="cuda"
)
output = engine.generate("Quantized inference reduces VRAM footprint", max_tokens=32)
print("Quantized Output:", output)import microgen
# Partition linear layers across 2 GPU ranks
tp_engine = microgen.LLMEngine.from_pretrained(
"gpt2",
backend_type="tensor_parallel",
tp_world_size=2
)
output = tp_engine.generate("Distributed tensor parallelism scales decoding", max_tokens=32)
print("TP Output:", output)Turn any causal LM into a probabilistic decision engine without generating a single token. Constrained scoring evaluates all candidates in one pass — with zero position bias.
from transformers import AutoModelForCausalLM, AutoTokenizer
from microgen.decision.huggingface import TransformersDecisionModel
from microgen.decision.engine import DecisionEngine
from microgen.decision.schema import Choice, ChoiceSchema
# Wrap any HuggingFace causal model
tokenizer = AutoTokenizer.from_pretrained("HuggingFaceTB/SmolLM-135M")
model = AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM-135M")
engine = DecisionEngine(TransformersDecisionModel(model, tokenizer))
context = "The drone is 5 meters from a building."
# Binary yes/no question — no generation, just logit comparison
result = engine.yes_no(context=context, question="Is the drone in danger?")
print(result.choice) # "YES"
print(result.top_probability) # 0.9508
print(result.entropy) # 0.2829 bits
# Multi-token structured choice with Trie-based constrained decoding
schema = ChoiceSchema(name="action", options=[
Choice("TURN LEFT"), Choice("PULL UP"),
Choice("BRAKE"), Choice("CONTINUE"),
])
result = engine.choose(context=context, schema=schema)
print(result.choice) # "TURN LEFT"
print(result.probabilities) # {"TURN LEFT": 0.67, "CONTINUE": 0.19, ...}
# Temperature-scale the distribution to reduce overconfidence
from microgen.decision.calibration import TemperatureScaler
scaler = TemperatureScaler()
scaler.fit(validation_results, true_labels) # fit on a labelled set
calibrated = scaler.transform(result) # ECE: 39.8% → 28.4%Key properties:
- Zero-decode: No token generation; scores candidates directly from logits.
- Trie-based Constrained Decoding: Shared token prefixes computed once — eliminates redundant forward passes.
-
Option-Order Invariance: Proven
$D_{KL}(P_{original} | P_{permuted}) = 0.0$ across all permutations. -
Temperature Calibration: Fits a temperature parameter
$T$ via L-BFGS to reduce Expected Calibration Error. -
Entropy Diagnostics: Every
DecisionResultcarries Shannon entropy over the candidate distribution.
Start the HTTP API server:
microgen serve --host 0.0.0.0 --port 8000 --model sshleifer/tiny-gpt2Test completion with curl:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "sshleifer/tiny-gpt2",
"messages": [{"role": "user", "content": "Explain LLM inference"}],
"max_tokens": 50,
"temperature": 0.7
}'Test Server-Sent Events (SSE) streaming:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "sshleifer/tiny-gpt2",
"messages": [{"role": "user", "content": "Write a poem"}],
"stream": true
}'MicroGen provides a rich Click-based unified CLI (microgen):
# 1. Start Server
microgen serve --port 8000 --model sshleifer/tiny-gpt2
# 2. Terminal Interactive Chat
microgen chat --model sshleifer/tiny-gpt2
# 3. Standalone Text Generation
microgen generate --prompt "Artificial Intelligence is" --max-tokens 32
# 4. Run Benchmark Suite
microgen benchmark --model sshleifer/tiny-gpt2
# 5. Profile Execution Bottlenecks
microgen profile --prompt "Benchmark continuous batching" --backend quantizedMicroGen includes an automated statistical benchmarking suite (
# Run End-to-End Latency & Throughput Benchmark
python scripts/e2e_benchmark.py
# Export LaTeX Paper Tables (paper/tables/*.tex)
python scripts/export_paper_tables.py
# Generate Publication Vector Figures (paper/figures/*.pdf)
python scripts/generate_paper_figures.pyArtifacts Produced:
paper/main.pdf: Compiled 14-page research manuscript.arxiv_submission.zip: Self-contained arXiv submission bundle.
MicroGen is covered by a comprehensive 149-test Pytest suite:
# Run full isolated test suite (149 passing tests)
PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python -m pytestmicrogen/
├── microgen/
│ ├── api/ # FastAPI HTTP app & SSE streaming endpoints
│ ├── backends/ # PyTorch, Quantized INT8, and TensorParallel backends
│ ├── caching/ # Prefix KV Cache & Token-Bucket Rate Limiter
│ ├── cli/ # Unified Click CLI commands
│ ├── devices/ # Hardware device abstractions (CPU & CUDA)
│ ├── profiling/ # Execution Profiler & Diagnostic Engine
│ ├── runtime/ # KVCacheState, Paged KV Allocator & Sliding Window Eviction
│ ├── sdk/ # High-level LLMEngine wrapper API
│ └── scheduler/ # Priority RequestQueue, Batching & Continuous Batching Scheduler
├── paper/ # LaTeX research manuscript, tables, vector figures, and arXiv zip
├── tests/ # 149 Unit & Integration Pytest test cases
├── scripts/ # End-to-End Benchmarking & Table export scripts
├── pyproject.toml # PyPI packaging specification (`microgen-llm`)
└── README.md # Primary repository documentation
This project is licensed under the MIT License — see the LICENSE file for details.
The decision module provides a System-One-style API over any causal Language Model. It transforms LMs into constrained, probabilistic decision engines without generating any output tokens.
Instead of generating text autoregressively and parsing the output, the engine performs a single prefill pass and evaluates valid candidates directly from the logits.
flowchart TD
Prompt[Context + Question] --> Prefill[Model Prefill Pass]
Prefill --> Logits[Logits extracted at last position]
Candidates[Schema: 'YES', 'NO'] --> Trie[Batched BFS Trie Frontier]
Logits --> Trie
Trie -->|Multi-token candidates| Decode[Targeted Decode Passes]
Decode --> Trie
Trie --> Score[Candidate Scores & Length Normalization]
Score --> Calibrate[Optional Temperature Scaling]
Calibrate --> Prob[Probability Distribution & Entropy]
Prob --> Result[DecisionResult]
The engine relies on strictly typed, immutable dataclasses to ensure decision boundaries are respected.
Choice: A single candidate option.name: str— The label used in probabilities.value: Any— Optional caller-defined payload.
ChoiceSchema: The full candidate set definition.name: str— Identifier (e.g., "boolq").options: list[Choice]— Must contain at least 2 options.
DecisionResult: The immutable output of an inference call.choice: str— The winning option.probabilities: dict[str, float]— Full normalized probability distribution.top_probability: float— Probability of the winning choice. (Note: explicitly not namedconfidenceunless calibrated).entropy: float— Shannon entropy of the distribution in bits.calibrated: bool— True if aTemperatureScalerwas applied.
CandidateStats: Diagnostic record for sequence scoring.- Contains
token_count,token_ids,raw_logprob, andnormalized_score.
- Contains
The DecisionEngine is the primary entrypoint. It wraps an underlying DecisionModel (via adapters.py) and routes requests to the fast single-token scorer or the Batched Trie sequence scorer.
choose(context, schema, temperature=1.0, alpha=1.0)
The core method. Scores each option in the schema against the context.
-
alpha: Length normalization penalty ($S_\alpha = \frac{\sum \log P}{|Y|^\alpha}$ ). - Example Use Case: Selecting a strategic action in a robotics pipeline.
from microgen.decision.schema import Choice, ChoiceSchema
from microgen.decision.engine import DecisionEngine
schema = ChoiceSchema(name="action", options=[
Choice("TURN LEFT"), Choice("TURN RIGHT"), Choice("STOP")
])
result = engine.choose("Obstacle detected ahead.", schema=schema)
print(result.choice) # "STOP"
print(result.probabilities) # {"TURN LEFT": 0.1, "TURN RIGHT": 0.1, "STOP": 0.8}
print(f"Entropy: {result.entropy:.2f} bits")yes_no(context, question, temperature=1.0)
A high-level wrapper around choose() for boolean tasks.
result = engine.yes_no(context="User is asking for financial advice.", question="Is this safe?")
if result.choice == "NO" and result.top_probability > 0.9:
block_request()score(context, criteria, scale, temperature=1.0)
A high-level wrapper around choose() for discrete integer scaling (e.g., 1-5 rating).
result = engine.score("The code is missing tests.", criteria="Quality", scale=[1, 2, 3, 4, 5])
print(result.choice) # "2"For multi-token candidates, scoring sequentially ($O(N)$) causes latency to explode as vocabulary spaces grow (e.g., 77 classes). MicroGen solves this via a Batched BFS Trie Frontier.
- Candidates are tokenized and organized into a Prefix Trie.
- The engine evaluates all active branches at depth
$D$ in a single batched PyTorch forward pass. - Latency scales by the depth of the longest candidate, not the breadth of the candidate space.
Metrics: Evaluated on a Tesla T4, latency for 10 candidates vs. 77 candidates remains essentially flat (~320ms) because both share a maximum depth of 6 tokens.
MicroGen evaluates the reliability of a decision using standard information theory metrics rather than raw logits.
- Shannon Entropy: Calculated natively for every decision, representing the uncertainty in bits across the candidate distribution.
- Temperature Scaling (
TemperatureScaler): Fits a scalarTusing L-BFGS over a validation dataset to minimize Negative Log Likelihood (NLL). Can transform uncalibratedDecisionResultobjects into calibrated probabilities.
from microgen.decision.calibration import TemperatureScaler
scaler = TemperatureScaler()
scaler.fit(validation_results, true_labels) # Fits T to minimize NLL
calibrated_result = scaler.transform(raw_result)
print(calibrated_result.calibrated) # TrueThe core generation architecture separates the high-level orchestration API (LLMEngine) from the hardware-specific forward pass execution (InferenceBackend). This allows the same DecisionEngine and LLMEngine to run seamlessly over quantized, tensor-parallel, or standard PyTorch weights.
The fundamental abstraction boundary. All core logic interacts only with the backend's prefill and decode methods, ensuring hardware agnosticism.
prefill(input_ids, attention_mask, cache)Performs the initial un-cached forward pass over a prompt. Returns(logits, updated_cache).decode(token_ids, attention_mask, cache)Performs a single-token forward pass utilizing the KV cache. Returns(logits, updated_cache).
Available Backend Implementations:
PyTorchBackend(pytorch.py): Standard HuggingFaceAutoModelForCausalLMloading and inference (FP32/FP16/BF16).QuantizedPyTorchBackend(quantized.py): Leverages BitsAndBytes for 8-bit (int8) or 4-bit/fp8 inference, drastically reducing VRAM footprints.TensorParallelPyTorchBackend(parallel.py): Implements Megatron-1D tensor parallelism. Shards linear projections (Attention Q/K/V/O, MLP Gate/Up/Down) across multiple GPUs viatorch.distributed, accelerating memory-bound decoding.
The high-level developer-facing LLM engine interface. Provides automated backend dispatch and exposes standard text generation parameters.
LLMEngine.from_pretrained(...)
Factory method that automatically instantiates the correct InferenceBackend based on user flags.
-
model_name_or_path: HuggingFace ID (e.g., "Qwen/Qwen2.5-1.5B-Instruct"). -
quantize: Set to"int8"or"fp8". Dispatches toQuantizedPyTorchBackend. -
tensor_parallel_size: If$>1$ , dispatches toTensorParallelPyTorchBackendallocating ranks across available GPUs.
generate(prompt, max_new_tokens=50, stream=False, temperature=1.0)
Implements classic autoregressive text generation over the backend primitives. If stream=True, yields an Iterator[str] for SSE event streaming or CLI type-writer effects.
from microgen.sdk.engine import LLMEngine
engine = LLMEngine.from_pretrained(
"Qwen/Qwen2.5-1.5B",
quantize="int8",
tensor_parallel_size=1
)
for token in engine.generate("The capital of France is", stream=True):
print(token, end="", flush=True)These modules power the continuous request batching and high-performance KV cache allocation that allow MicroGen to sustain high throughput.
The ContinuousBatchingScheduler implements iteration-level scheduling. Instead of waiting for an entire batch of requests to finish, it injects new requests at the prefill stage as soon as batch slots become available.
RequestQueue: Thread-safe priority queue managing pending inference requests (Requestobjects trackingttft_ms,tpot_ms, andpriority).step(): The core scheduling loop.- Pops up to
max_batch_sizepending requests and executes a batchedprefill. - Executes a batched
decodefor all currently active running requests. - Evicts requests that hit
max_new_tokensor emit an EOS token.
- Pops up to
- Micro-Profiling: The scheduler tracks internal execution overhead natively via
profiling_stats, isolating Python event loop overheads from CUDA kernel time.
from microgen.scheduler.scheduler import ContinuousBatchingScheduler
from microgen.scheduler.queue import Request
scheduler = ContinuousBatchingScheduler(backend=backend, kv_cache_manager=manager, max_batch_size=8)
scheduler.add_request(Request(request_id="1", prompt="Hello", prompt_ids=[1,2,3]))
# Run until all requests finish
completed = scheduler.run_until_complete()Memory management is crucial for LLM serving. MicroGen provides two paradigms:
-
KVCacheState: A dynamic per-request cache state that subclasses HuggingFace'sCache.- Tracks
key_cacheandvalue_cacheacross all layers. -
Trie Routing Support: Implements
expand_batchandgather_batchto support the Batched BFS Trie Frontier, allowing cache branching without redundant prefill passes. -
Quantization: Natively supports per-vector INT8 KV cache quantization (
quantize_kv=True), effectively halving memory requirements.
- Tracks
-
PagedKVCacheAllocator: Implements physical block paging (inspired by vLLM's PagedAttention).- Manages a pool of fixed-size
PhysicalBlockobjects (e.g., 16 tokens per block). -
BlockTablemaps logical sequence tokens to physical blocks dynamically. - Eliminates external memory fragmentation (
$F_{ext} = 0$ ) by avoiding contiguous memory allocations for unpredictable sequence lengths.
- Manages a pool of fixed-size
These modules provide the external boundaries of the MicroGen framework, exposing internal engine optimizations to users and CI systems.
As detailed in Section 2, the LLMEngine is the primary entrypoint for programmatic usage. It allows developers to load models in 1 line with automated hardware-aware backend dispatch (PyTorch, INT8, TP=2).
MicroGen exposes a rich, unified Click CLI for interactive terminal usage and automated benchmarking.
microgen chat: Launches an interactive, streaming terminal session.microgen serve: Boots the FastAPI HTTP server using continuous batching.microgen benchmark: Runs an automated synthetic throughput/latency benchmark usingWorkloadGenerator.microgen generate: Executes a one-shot generation prompt.microgen profile: Runs a micro-profiling trace (Prefill/Decode ratio) and outputs bottleneck diagnostics.
Example CLI Usage:
microgen chat --model Qwen/Qwen2.5-1.5B-Instruct --device cuda --quantize int8
microgen benchmark --model sshleifer/tiny-gpt2 --num-requests 100 --max-tokens 32A production-ready HTTP server exposing OpenAI-compatible endpoints. It deeply integrates with the ContinuousBatchingScheduler to manage concurrent incoming requests asynchronously.
Endpoints:
GET /health&GET /v1/modelsPOST /v1/completions: Standard text generation.POST /v1/chat/completions: Chat generation format (roles + messages).
Both completions endpoints natively support stream=True, returning standard Server-Sent Events (SSE) text/event-stream chunks for seamless frontend integration.