Skip to content

Repository files navigation

⚡ MicroGen: LLM Inference Optimization Research Framework

PyPI Python 3.11+ PyTorch FastAPI Tests License

MicroGen (microgen-llm on PyPI) is a modular, hardware-aware Large Language Model (LLM) inference research framework and experimental substrate built from scratch in PyTorch. Designed to dissect memory, latency, and throughput trade-offs under controlled hardware and workload conditions, MicroGen isolates state-of-the-art serving techniques including Physical Paged KV Allocation, Hash-Based Prefix Reuse, INT8 Weight Quantization, Multi-GPU Tensor Parallelism, Speculative Decoding, and Continuous Request Batching.


📌 Research Framing & Evaluation Substrate

MicroGen: An empirical LLM inference research substrate isolating optimization overheads, non-monotonic composition dynamics, and hardware trade-offs behind modular execution protocols. Tested via an $N=30$ repeated-trial evaluation protocol across NVIDIA Tesla T4 and P100 GPUs and open-weights model families (GPT-2, Qwen2.5, Llama-3.2).

Topics/Tags for GitHub: llm-inference, systems-research, pytorch, kv-cache, continuous-batching, paged-attention, tensor-parallelism, quantization, speculative-decoding, fastapi, cuda.

Note on Research Framework Positioning: MicroGen is explicitly designed as an experimental systems research substrate for isolating optimization overheads and measuring non-monotonic interaction dynamics under controlled conditions, rather than competing directly with production C++/CUDA serving engines (e.g., vLLM, TensorRT-LLM, SGLang). Beyond the empirical benchmarking substrate described in the paper, MicroGen includes a working OpenAI-compatible HTTP serving layer (microgen/api/ with FastAPI, SSE streaming, and 149 passing unit/integration tests). The paper's $N=30$ throughput/latency figures reflect direct in-process engine-level measurement (isolating model, memory, and kernel dynamics).


🌟 Key Empirical Discoveries & System Principles

  • 🚀 Non-Monotonic Optimization Composition: Combining individually useful optimizations (+INT8, +Paged KV, +Prefix Cache) yields a statistically significant throughput regression to $0.96\times$ baseline ($493.6 \pm 8.6\text{ tok/s}$, $p_{\text{adj}} < 0.001$), while continuous batching scheduler overhead drops throughput to $0.76\times$ baseline ($391.8 \pm 6.2\text{ tok/s}$, $p_{\text{adj}} < 0.001$) due to cumulative Python event loop and pointer indirection overheads ($\eta_{\text{overhead}} = 24.8%$).
  • 🧠 Physical Block Paged KV Allocation: Dynamically assigns $B_{\text{block}}=16$ token physical blocks, eliminating external contiguous-allocation memory fragmentation ($F_{\text{ext}} = 1 - \frac{\text{max contiguous block}}{\text{total free VRAM}} = 0.0%$) under dynamic memory pressure regimes.
  • ⚡ Hash-Based Prefix Cache Reuse: Implements exact longest-common-prefix (LCP) key lookup as an experimental baseline approximation of shared-prefix caching, delivering up to a $3.91\times$ prefill TTFT speedup ($6.6 \pm 0.3\text{ ms}$ vs $25.8 \pm 1.2\text{ ms}$, $p_{\text{adj}} < 0.001$) under 100% prompt overlap at $L_{\text{prompt}}=1024$ tokens, crossing into positive speedup once prompt overlap exceeds 25%.
  • 🔮 Speculative Decoding Acceptance Boundaries: Characterizes acceptance rate break-even threshold $\alpha_{\text{threshold}} = \frac{T_{\text{draft_step}}}{T_{\text{target_step}}}$. On small target models (tiny-gpt2), draft step overhead ($4.4\text{ ms}$) exceeds verification savings, resulting in a throughput regression ($0.45\times$ baseline, $p_{\text{adj}} < 0.001$).
  • 🌐 Multi-GPU Tensor Parallelism: Shards linear projections across dual NVIDIA T4 GPUs ($TP=2$), accelerating memory-bound decoding for GPT-2 ($124\text{M}$) from $8.2\text{ tok/s}$ to $14.0\text{ tok/s}$ ($1.71\times$ speedup, $p_{\text{adj}} < 0.001$).

🎯 Portfolio & Resume Framing (Research $\rightarrow$ Methodology $\rightarrow$ Discovery $\rightarrow$ Result)

  • LLM Inference Systems Architecture: Designed and built MicroGen, a modular PyTorch LLM inference research framework isolating memory, latency, and throughput trade-offs across CPU, CUDA, and multi-GPU ($TP=2$) execution protocols.
  • Non-Monotonic Composition Analysis: Discovered through an $N=30$ repeated-trial ablation protocol that composing individually positive optimizations (+INT8, +Paged KV, +Prefix Cache) yields non-monotonic throughput degradation ($0.96\times$ baseline, $p_{\text{adj}} < 0.001$).
  • Prefix Reuse & TTFT Acceleration: Implemented an experimental hash-based longest-common-prefix (LCP) KV cache manager achieving a $3.91\times$ prefill TTFT speedup ($6.6\text{ ms}$ vs $25.8\text{ ms}$, $p_{\text{adj}} < 0.001$) under 100% prompt overlap.
  • Continuous Batching Overhead Profiling: Built a micro-profiling harness isolating Python event loop overhead ($\eta_{\text{overhead}} = 24.8%$), causally explaining continuous batching throughput regressions ($0.76\times$ baseline, $p_{\text{adj}} < 0.001$) in research substrates.
  • Memory Modeling & Multi-GPU Acceleration: Formalized external VRAM allocation fragmentation ($F_{\text{ext}} = 1 - \frac{\text{max contiguous block}}{\text{total free VRAM}}$) and sharded linear projections across dual NVIDIA T4 GPUs ($TP=2$), delivering a $1.71\times$ throughput speedup ($14.0\text{ tok/s}$ vs $8.2\text{ tok/s}$, $p_{\text{adj}} < 0.001$).

🏗️ Architecture Overview

flowchart TD
    Client[Client / HTTP Request / CLI] --> API[FastAPI OpenAI Router / CLI Entry]
    API --> RateLimiter[Token Bucket Rate Limiter]
    RateLimiter --> Scheduler[Continuous Batching Scheduler]
    
    subgraph Engine Core
        Scheduler --> RequestQueue[Priority Request Queue]
        Scheduler --> PrefixCache[Prefix KV Cache Manager]
        Scheduler --> KVCache[Paged & INT8 Quantized KV Cache]
        Scheduler --> Backend[Inference Backend Interface]
    end

    subgraph Hardware Backends
        Backend --> PyTorchBackend[PyTorch Standard Backend]
        Backend --> QuantizedBackend[Quantized INT8 Backend]
        Backend --> TPBackend[Tensor-Parallel Multi-GPU Backend]
    end

    subgraph Devices & Hardware
        PyTorchBackend --> CPUDevice[CPU Hardware Device]
        PyTorchBackend --> CUDADevice[NVIDIA CUDA GPU Device]
        TPBackend --> MultiGPU[Multi-Rank CUDA GPUs]
    end
Loading

🛠️ Installation & Setup

Install from PyPI

pip install microgen-llm

Install from Source

git clone https://github.com/Omdeepb69/MicroGen.git
cd MicroGen

python -m venv venv
source venv/bin/activate  # On Linux/macOS
# or: venv\Scripts\activate on Windows

pip install -e .

🚀 Quickstart & Usage Examples

1. Fluent SDK Wrapper API (microgen.LLMEngine)

import microgen

# Initialize engine with PyTorch FP32 backend
engine = microgen.LLMEngine.from_pretrained(
    "sshleifer/tiny-gpt2",
    backend_type="pytorch",
    device="cuda"
)

# Generate completion text
output = engine.generate("MicroGen is a fast LLM inference engine", max_tokens=32)
print("Output:", output)

2. INT8 Quantized Model Execution

import microgen

# Load quantized backend (INT8 weights + dynamic INT8 KV cache)
engine = microgen.LLMEngine.from_pretrained(
    "sshleifer/tiny-gpt2",
    backend_type="quantized",
    device="cuda"
)

output = engine.generate("Quantized inference reduces VRAM footprint", max_tokens=32)
print("Quantized Output:", output)

3. Multi-GPU Tensor Parallel Execution ($TP=2$)

import microgen

# Partition linear layers across 2 GPU ranks
tp_engine = microgen.LLMEngine.from_pretrained(
    "gpt2",
    backend_type="tensor_parallel",
    tp_world_size=2
)

output = tp_engine.generate("Distributed tensor parallelism scales decoding", max_tokens=32)
print("TP Output:", output)

4. 🧠 Decode-Free Decision Engine (v1.1.0+)

Turn any causal LM into a probabilistic decision engine without generating a single token. Constrained scoring evaluates all candidates in one pass — with zero position bias.

from transformers import AutoModelForCausalLM, AutoTokenizer
from microgen.decision.huggingface import TransformersDecisionModel
from microgen.decision.engine import DecisionEngine
from microgen.decision.schema import Choice, ChoiceSchema

# Wrap any HuggingFace causal model
tokenizer = AutoTokenizer.from_pretrained("HuggingFaceTB/SmolLM-135M")
model = AutoModelForCausalLM.from_pretrained("HuggingFaceTB/SmolLM-135M")
engine = DecisionEngine(TransformersDecisionModel(model, tokenizer))

context = "The drone is 5 meters from a building."

# Binary yes/no question — no generation, just logit comparison
result = engine.yes_no(context=context, question="Is the drone in danger?")
print(result.choice)           # "YES"
print(result.top_probability)  # 0.9508
print(result.entropy)          # 0.2829 bits

# Multi-token structured choice with Trie-based constrained decoding
schema = ChoiceSchema(name="action", options=[
    Choice("TURN LEFT"), Choice("PULL UP"),
    Choice("BRAKE"),     Choice("CONTINUE"),
])
result = engine.choose(context=context, schema=schema)
print(result.choice)         # "TURN LEFT"
print(result.probabilities)  # {"TURN LEFT": 0.67, "CONTINUE": 0.19, ...}

# Temperature-scale the distribution to reduce overconfidence
from microgen.decision.calibration import TemperatureScaler
scaler = TemperatureScaler()
scaler.fit(validation_results, true_labels)  # fit on a labelled set
calibrated = scaler.transform(result)        # ECE: 39.8% → 28.4%

Key properties:

  • Zero-decode: No token generation; scores candidates directly from logits.
  • Trie-based Constrained Decoding: Shared token prefixes computed once — eliminates redundant forward passes.
  • Option-Order Invariance: Proven $D_{KL}(P_{original} | P_{permuted}) = 0.0$ across all permutations.
  • Temperature Calibration: Fits a temperature parameter $T$ via L-BFGS to reduce Expected Calibration Error.
  • Entropy Diagnostics: Every DecisionResult carries Shannon entropy over the candidate distribution.

5. OpenAI-Compatible HTTP Serving & SSE Streaming

Start the HTTP API server:

microgen serve --host 0.0.0.0 --port 8000 --model sshleifer/tiny-gpt2

Test completion with curl:

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sshleifer/tiny-gpt2",
    "messages": [{"role": "user", "content": "Explain LLM inference"}],
    "max_tokens": 50,
    "temperature": 0.7
  }'

Test Server-Sent Events (SSE) streaming:

curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "sshleifer/tiny-gpt2",
    "messages": [{"role": "user", "content": "Write a poem"}],
    "stream": true
  }'

💻 Command Line Interface (CLI)

MicroGen provides a rich Click-based unified CLI (microgen):

# 1. Start Server
microgen serve --port 8000 --model sshleifer/tiny-gpt2

# 2. Terminal Interactive Chat
microgen chat --model sshleifer/tiny-gpt2

# 3. Standalone Text Generation
microgen generate --prompt "Artificial Intelligence is" --max-tokens 32

# 4. Run Benchmark Suite
microgen benchmark --model sshleifer/tiny-gpt2

# 5. Profile Execution Bottlenecks
microgen profile --prompt "Benchmark continuous batching" --backend quantized

📊 Benchmarking & Reproducibility Package

MicroGen includes an automated statistical benchmarking suite ($N=30$ repeated trials):

# Run End-to-End Latency & Throughput Benchmark
python scripts/e2e_benchmark.py

# Export LaTeX Paper Tables (paper/tables/*.tex)
python scripts/export_paper_tables.py

# Generate Publication Vector Figures (paper/figures/*.pdf)
python scripts/generate_paper_figures.py

Artifacts Produced:

  • paper/main.pdf: Compiled 14-page research manuscript.
  • arxiv_submission.zip: Self-contained arXiv submission bundle.

🧪 Testing & Verification

MicroGen is covered by a comprehensive 149-test Pytest suite:

# Run full isolated test suite (149 passing tests)
PYTEST_DISABLE_PLUGIN_AUTOLOAD=1 python -m pytest

📁 Repository Structure

microgen/
├── microgen/
│   ├── api/             # FastAPI HTTP app & SSE streaming endpoints
│   ├── backends/        # PyTorch, Quantized INT8, and TensorParallel backends
│   ├── caching/         # Prefix KV Cache & Token-Bucket Rate Limiter
│   ├── cli/             # Unified Click CLI commands
│   ├── devices/         # Hardware device abstractions (CPU & CUDA)
│   ├── profiling/       # Execution Profiler & Diagnostic Engine
│   ├── runtime/         # KVCacheState, Paged KV Allocator & Sliding Window Eviction
│   ├── sdk/             # High-level LLMEngine wrapper API
│   └── scheduler/       # Priority RequestQueue, Batching & Continuous Batching Scheduler
├── paper/               # LaTeX research manuscript, tables, vector figures, and arXiv zip
├── tests/               # 149 Unit & Integration Pytest test cases
├── scripts/             # End-to-End Benchmarking & Table export scripts
├── pyproject.toml       # PyPI packaging specification (`microgen-llm`)
└── README.md            # Primary repository documentation

📜 License

This project is licensed under the MIT License — see the LICENSE file for details.


📚 Exhaustive Module Documentation

1. The microgen.decision Module

The decision module provides a System-One-style API over any causal Language Model. It transforms LMs into constrained, probabilistic decision engines without generating any output tokens.

Architecture & Flow

Instead of generating text autoregressively and parsing the output, the engine performs a single prefill pass and evaluates valid candidates directly from the logits.

flowchart TD
    Prompt[Context + Question] --> Prefill[Model Prefill Pass]
    Prefill --> Logits[Logits extracted at last position]
    
    Candidates[Schema: 'YES', 'NO'] --> Trie[Batched BFS Trie Frontier]
    
    Logits --> Trie
    Trie -->|Multi-token candidates| Decode[Targeted Decode Passes]
    Decode --> Trie
    
    Trie --> Score[Candidate Scores & Length Normalization]
    Score --> Calibrate[Optional Temperature Scaling]
    Calibrate --> Prob[Probability Distribution & Entropy]
    Prob --> Result[DecisionResult]
Loading

Core Data Structures (schema.py)

The engine relies on strictly typed, immutable dataclasses to ensure decision boundaries are respected.

  • Choice: A single candidate option.
    • name: str — The label used in probabilities.
    • value: Any — Optional caller-defined payload.
  • ChoiceSchema: The full candidate set definition.
    • name: str — Identifier (e.g., "boolq").
    • options: list[Choice] — Must contain at least 2 options.
  • DecisionResult: The immutable output of an inference call.
    • choice: str — The winning option.
    • probabilities: dict[str, float] — Full normalized probability distribution.
    • top_probability: float — Probability of the winning choice. (Note: explicitly not named confidence unless calibrated).
    • entropy: float — Shannon entropy of the distribution in bits.
    • calibrated: bool — True if a TemperatureScaler was applied.
  • CandidateStats: Diagnostic record for sequence scoring.
    • Contains token_count, token_ids, raw_logprob, and normalized_score.

The DecisionEngine (engine.py)

The DecisionEngine is the primary entrypoint. It wraps an underlying DecisionModel (via adapters.py) and routes requests to the fast single-token scorer or the Batched Trie sequence scorer.

choose(context, schema, temperature=1.0, alpha=1.0) The core method. Scores each option in the schema against the context.

  • alpha: Length normalization penalty ($S_\alpha = \frac{\sum \log P}{|Y|^\alpha}$).
  • Example Use Case: Selecting a strategic action in a robotics pipeline.
from microgen.decision.schema import Choice, ChoiceSchema
from microgen.decision.engine import DecisionEngine

schema = ChoiceSchema(name="action", options=[
    Choice("TURN LEFT"), Choice("TURN RIGHT"), Choice("STOP")
])
result = engine.choose("Obstacle detected ahead.", schema=schema)

print(result.choice)             # "STOP"
print(result.probabilities)      # {"TURN LEFT": 0.1, "TURN RIGHT": 0.1, "STOP": 0.8}
print(f"Entropy: {result.entropy:.2f} bits")

yes_no(context, question, temperature=1.0) A high-level wrapper around choose() for boolean tasks.

result = engine.yes_no(context="User is asking for financial advice.", question="Is this safe?")
if result.choice == "NO" and result.top_probability > 0.9:
    block_request()

score(context, criteria, scale, temperature=1.0) A high-level wrapper around choose() for discrete integer scaling (e.g., 1-5 rating).

result = engine.score("The code is missing tests.", criteria="Quality", scale=[1, 2, 3, 4, 5])
print(result.choice) # "2"

Batched Trie Sequence Scoring (constrained.py)

For multi-token candidates, scoring sequentially ($O(N)$) causes latency to explode as vocabulary spaces grow (e.g., 77 classes). MicroGen solves this via a Batched BFS Trie Frontier.

  1. Candidates are tokenized and organized into a Prefix Trie.
  2. The engine evaluates all active branches at depth $D$ in a single batched PyTorch forward pass.
  3. Latency scales by the depth of the longest candidate, not the breadth of the candidate space.

Metrics: Evaluated on a Tesla T4, latency for 10 candidates vs. 77 candidates remains essentially flat (~320ms) because both share a maximum depth of 6 tokens.

Calibration & Entropy (calibration.py, entropy.py)

MicroGen evaluates the reliability of a decision using standard information theory metrics rather than raw logits.

  • Shannon Entropy: Calculated natively for every decision, representing the uncertainty in bits across the candidate distribution.
  • Temperature Scaling (TemperatureScaler): Fits a scalar T using L-BFGS over a validation dataset to minimize Negative Log Likelihood (NLL). Can transform uncalibrated DecisionResult objects into calibrated probabilities.
from microgen.decision.calibration import TemperatureScaler

scaler = TemperatureScaler()
scaler.fit(validation_results, true_labels) # Fits T to minimize NLL
calibrated_result = scaler.transform(raw_result)
print(calibrated_result.calibrated) # True

2. The backends & engine Modules

The core generation architecture separates the high-level orchestration API (LLMEngine) from the hardware-specific forward pass execution (InferenceBackend). This allows the same DecisionEngine and LLMEngine to run seamlessly over quantized, tensor-parallel, or standard PyTorch weights.

The InferenceBackend Protocol (backends/base.py)

The fundamental abstraction boundary. All core logic interacts only with the backend's prefill and decode methods, ensuring hardware agnosticism.

  • prefill(input_ids, attention_mask, cache) Performs the initial un-cached forward pass over a prompt. Returns (logits, updated_cache).
  • decode(token_ids, attention_mask, cache) Performs a single-token forward pass utilizing the KV cache. Returns (logits, updated_cache).

Available Backend Implementations:

  1. PyTorchBackend (pytorch.py): Standard HuggingFace AutoModelForCausalLM loading and inference (FP32/FP16/BF16).
  2. QuantizedPyTorchBackend (quantized.py): Leverages BitsAndBytes for 8-bit (int8) or 4-bit/fp8 inference, drastically reducing VRAM footprints.
  3. TensorParallelPyTorchBackend (parallel.py): Implements Megatron-1D tensor parallelism. Shards linear projections (Attention Q/K/V/O, MLP Gate/Up/Down) across multiple GPUs via torch.distributed, accelerating memory-bound decoding.

The LLMEngine SDK (sdk/engine.py)

The high-level developer-facing LLM engine interface. Provides automated backend dispatch and exposes standard text generation parameters.

LLMEngine.from_pretrained(...) Factory method that automatically instantiates the correct InferenceBackend based on user flags.

  • model_name_or_path: HuggingFace ID (e.g., "Qwen/Qwen2.5-1.5B-Instruct").
  • quantize: Set to "int8" or "fp8". Dispatches to QuantizedPyTorchBackend.
  • tensor_parallel_size: If $>1$, dispatches to TensorParallelPyTorchBackend allocating ranks across available GPUs.

generate(prompt, max_new_tokens=50, stream=False, temperature=1.0) Implements classic autoregressive text generation over the backend primitives. If stream=True, yields an Iterator[str] for SSE event streaming or CLI type-writer effects.

from microgen.sdk.engine import LLMEngine

engine = LLMEngine.from_pretrained(
    "Qwen/Qwen2.5-1.5B", 
    quantize="int8", 
    tensor_parallel_size=1
)

for token in engine.generate("The capital of France is", stream=True):
    print(token, end="", flush=True)

3. The scheduler & memory Modules

These modules power the continuous request batching and high-performance KV cache allocation that allow MicroGen to sustain high throughput.

Continuous Batching Scheduler (scheduler/scheduler.py, queue.py)

The ContinuousBatchingScheduler implements iteration-level scheduling. Instead of waiting for an entire batch of requests to finish, it injects new requests at the prefill stage as soon as batch slots become available.

  • RequestQueue: Thread-safe priority queue managing pending inference requests (Request objects tracking ttft_ms, tpot_ms, and priority).
  • step(): The core scheduling loop.
    1. Pops up to max_batch_size pending requests and executes a batched prefill.
    2. Executes a batched decode for all currently active running requests.
    3. Evicts requests that hit max_new_tokens or emit an EOS token.
  • Micro-Profiling: The scheduler tracks internal execution overhead natively via profiling_stats, isolating Python event loop overheads from CUDA kernel time.
from microgen.scheduler.scheduler import ContinuousBatchingScheduler
from microgen.scheduler.queue import Request

scheduler = ContinuousBatchingScheduler(backend=backend, kv_cache_manager=manager, max_batch_size=8)
scheduler.add_request(Request(request_id="1", prompt="Hello", prompt_ids=[1,2,3]))

# Run until all requests finish
completed = scheduler.run_until_complete()

KV Cache Management (runtime/kv_cache.py, runtime/paged_kv.py)

Memory management is crucial for LLM serving. MicroGen provides two paradigms:

  1. KVCacheState: A dynamic per-request cache state that subclasses HuggingFace's Cache.

    • Tracks key_cache and value_cache across all layers.
    • Trie Routing Support: Implements expand_batch and gather_batch to support the Batched BFS Trie Frontier, allowing cache branching without redundant prefill passes.
    • Quantization: Natively supports per-vector INT8 KV cache quantization (quantize_kv=True), effectively halving memory requirements.
  2. PagedKVCacheAllocator: Implements physical block paging (inspired by vLLM's PagedAttention).

    • Manages a pool of fixed-size PhysicalBlock objects (e.g., 16 tokens per block).
    • BlockTable maps logical sequence tokens to physical blocks dynamically.
    • Eliminates external memory fragmentation ($F_{ext} = 0$) by avoiding contiguous memory allocations for unpredictable sequence lengths.

4. The api, cli, & sdk Modules

These modules provide the external boundaries of the MicroGen framework, exposing internal engine optimizations to users and CI systems.

Python SDK (sdk/engine.py)

As detailed in Section 2, the LLMEngine is the primary entrypoint for programmatic usage. It allows developers to load models in 1 line with automated hardware-aware backend dispatch (PyTorch, INT8, TP=2).

Command-Line Interface (cli/main.py)

MicroGen exposes a rich, unified Click CLI for interactive terminal usage and automated benchmarking.

  • microgen chat: Launches an interactive, streaming terminal session.
  • microgen serve: Boots the FastAPI HTTP server using continuous batching.
  • microgen benchmark: Runs an automated synthetic throughput/latency benchmark using WorkloadGenerator.
  • microgen generate: Executes a one-shot generation prompt.
  • microgen profile: Runs a micro-profiling trace (Prefill/Decode ratio) and outputs bottleneck diagnostics.

Example CLI Usage:

microgen chat --model Qwen/Qwen2.5-1.5B-Instruct --device cuda --quantize int8
microgen benchmark --model sshleifer/tiny-gpt2 --num-requests 100 --max-tokens 32

FastAPI Server (api/app.py)

A production-ready HTTP server exposing OpenAI-compatible endpoints. It deeply integrates with the ContinuousBatchingScheduler to manage concurrent incoming requests asynchronously.

Endpoints:

  • GET /health & GET /v1/models
  • POST /v1/completions: Standard text generation.
  • POST /v1/chat/completions: Chat generation format (roles + messages).

Both completions endpoints natively support stream=True, returning standard Server-Sent Events (SSE) text/event-stream chunks for seamless frontend integration.

About

A modular, production-grade PyTorch LLM inference engine featuring Continuous Batching, Paged & Quantized INT8 KV Cache, Multi-GPU Tensor Parallelism, Speculative Decoding, and OpenAI-compatible SSE streaming API.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages