🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
-
Updated
Sep 3, 2026 - Python
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
Run a documented subset of verl-style OPD on one consumer GPU—typed config, Parquet prompts, and PEFT scale-out artifacts.
RL study guide — foundations through RLHF, DPO, GRPO, RLVR, agentic RL, and offline RL. Hand-written CS294 notes, 19 lecture drafts, 5 tested exercises, citations that resolve.
SAFi is an open source reasoning engine for AI agents that enforces your policies, governs agent actions, and records every decision for audit.
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
Cognitive training practices for AI agents. Self-applied. Open source. Built by an independent Vancouver yoga studio.
A modular, production-grade framework for Reinforcement Learning from Human Feedback (RLHF) with Proximal Policy Optimization (PPO) and Direct Preference Optimization (DPO).
Learning When to Answer: Behavior-Oriented Reinforcement Learning for Hallucination Mitigation
Constitutional midtraining for LLM alignment: values-based document generation, 2×2 factorial training pipeline, and full evaluation suite (including benchmarks on value conflict resolution and sycophantic pressure)
📚 350+ loss functions across 25+ AI subdomains — classification, GANs, diffusion, LLM alignment, RL, contrastive learning, audio, video, time series, and more. Chronologically ordered with paper links, math formulas, and implementations.
A clean, production-ready PyTorch implementation of Target Policy Optimization (TPO) for stable RLHF alignment.
CS336 作业 5:基于 Qwen2.5 模型的 LLM 对齐与推理强化学习。完整实现了监督微调(SFT)与组相对策略优化(GRPO)算法,并在 GSM8K 数据集上完成零样本、在策与离策的训练与评估对比。
Hands-on tutorials on training-free alignment of language models at inference time.
Deterministic LLM framework combining strict context governance (veritas/sim frameworks) with a high-fidelity somatic simulation engine tracking real-time biometric and status vectors.
Adversarial AI system to test and improve reliability under real-world pressure
Official implementation of "DZ-TiDPO: Non-Destructive Temporal Alignment for Mutable State Tracking". SOTA on Multi-Session Chat with negligible alignment tax.
The paper list related to activation steering
EDT-Former, a brige for LLM and graph data. Entropy-guided Dynamic Token Transformer for Graph-LLM alignment. Accepted at ICLR 2026.
A hands-on, ground-up implementation sandbox dedicated to mastering the inner mechanics of Large Language Models and Agentic AI systems.
LLM Post-training(SFT, RLVR, RLHF) 파이프라인 구축 및 평가 실습 아카이브
To associate your repository with the llm-alignment topic, visit your repo's landing page and select "manage topics."