ReOPD On-policy distillation for agents without the environment: replay teacher prefixes instead of live rollouts. Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12 baohao/Math_Qwen3-4B-Instruct-2507_SFT-RL 4B • Updated Jul 23 • 13 baohao/Math_Qwen3-4B-Instruct-2507_SFT 4B • Updated Jul 23 • 37 baohao/Math_Qwen3-30B-A3B-Instruct-2507_SFT 31B • Updated Jul 23 • 10
Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12
SAGE Self-Hinting Language Models Enhance Reinforcement Learning Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31 baohao/aime24 Viewer • Updated Feb 7 • 30 • 553 baohao/aime25 Viewer • Updated Feb 7 • 30 • 524 baohao/amc23 Viewer • Updated Feb 7 • 40 • 451
Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31
ReOPD On-policy distillation for agents without the environment: replay teacher prefixes instead of live rollouts. Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12 baohao/Math_Qwen3-4B-Instruct-2507_SFT-RL 4B • Updated Jul 23 • 13 baohao/Math_Qwen3-4B-Instruct-2507_SFT 4B • Updated Jul 23 • 37 baohao/Math_Qwen3-30B-A3B-Instruct-2507_SFT 31B • Updated Jul 23 • 10
Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12
SAGE Self-Hinting Language Models Enhance Reinforcement Learning Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31 baohao/aime24 Viewer • Updated Feb 7 • 30 • 553 baohao/aime25 Viewer • Updated Feb 7 • 30 • 524 baohao/amc23 Viewer • Updated Feb 7 • 40 • 451
Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31