You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Full post-training pipeline for Qwen2.5-1.5B — SFT → SimPO → GRPO on free T4/P100 GPUs. GSM8K accuracy jumps from 23% (base) to 61% (GRPO) using Unsloth 4-bit LoRA, TRL, and HuggingFace Hub checkpointing.