Understanding Reasoning from Pretraining to Post-Training Paper • 2607.16097 • Published 13 days ago • 28
Pre2Post-Chess Collection Open-sourced models and datasets for training the chess reasoning models. • 7 items • Updated 9 days ago • 5
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay Paper • 2506.05316 • Published Jun 5, 2025 • 1
rlvr-weak-supervision Collection Models from "When Can LLMs Learn to Reason with Weak Supervision?" — Llama-3.2-3B with continual pre-training and Thinking SFT. • 3 items • Updated Apr 20 • 2
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning Paper • 2505.24846 • Published May 30, 2025 • 15