Companion code for Post-Training for LLMs: SFT, Preference Optimization, RLVR, Agentic RL, and Distillation.
-
Updated
Aug 6, 2026 - Python
Companion code for Post-Training for LLMs: SFT, Preference Optimization, RLVR, Agentic RL, and Distillation.
Critic-free Reward-Integrated Self-distillation Policy Optimization
LLM Post-Training, RLHF, PPO, DPO, etc
Coordinated AI systems evolving together.
To associate your repository with the post-training-learning topic, visit your repo's landing page and select "manage topics."