Training models with ternary quantized weights using PyTorch
-
Updated
Jun 12, 2019 - Python
Training models with ternary quantized weights using PyTorch
Pre-quantized ternary models for VLMs, multimodal, and audio — the models GGUF can't touch
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
1.58-bit ternary Mamba LLM for Indian languages. Weights are {-1,0,+1} — inference uses only add/sub. 3B model fits in 750MB, runs 20+ tok/s on mobile.
Colab-friendly BitNet distillation engine: collect KD traces from a teacher, train a ternary Mini-BitNet, and dry-run 7B memory. Multi-provider + Drive/S3
Research implementation of activation-aware ternary and mixed-bit quantization for Qwen3.8-27B, targeting 7-9 GB text inference on 16 GB GPUs.
Independent forensics of Bonsai 2 27B ternary quantization: format and basis recovery, the trained-weights residual, and the public PTQ calibration artifact.
PILON (Primitive-Induced Linear Operator Network) explores a compositional weight parameterization for transformer FFN layers. The goal is to replace dense FFN matrices with shared low-rank primitives plus learned composition weights.
Run PrismML Ternary Bonsai 2 27B (PTQ1_0 packed-ternary GGUF) natively on vLLM — tensor-core fused Triton kernel, 5.9GB resident, 890 tok/s @ bs=32, 2.5x official fork with matched slots
jade_mage · agentprivacy dual-agent harness lane for JadeZaher/ternary-memory-research (fork: main tracks upstream; lane on branch harness/jade_mage; contributions go upstream by PR)
To associate your repository with the ternary-quantization topic, visit your repo's landing page and select "manage topics."