RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
-
Updated
Aug 10, 2026 - Python
RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
Medical QA fine-tuning of Qwen2.5 using Unsloth and LoRA for medical question answering. Features optimized training, reduced GPU memory usage, accelerated inference, and robust clinical knowledge adaptation for healthcare-focused conversational AI.
Efficient fine-tuning of Phi-3.5-Mini-Instruct using Unsloth and LoRA for Medical Question Answering. Demonstrates memory-efficient training, faster inference, and domain-specific adaptation to build an accurate healthcare AI assistant.
Self-hosted Vietnamese text-to-speech web app powered by VieNeu-TTS 0.3B Q4 GGUF, CUDA-accelerated llama.cpp, ONNX Int8 codec, FastAPI and uv. Generates 24 kHz WAV locally, supports long-form Vietnamese text, and requires no cloud TTS API or API key.
BitNet-inspired 1-bit Quantized Transformer for efficient protein function prediction and biological sequence modeling on low-power devices.
Add a description, image, and links to the quantized-models topic page so that developers can more easily learn about it.
To associate your repository with the quantized-models topic, visit your repo's landing page and select "manage topics."