RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 11 days ago • 283
Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory Paper • 2610.02521 • Published 11 days ago • 59
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 11 days ago • 61
EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling Paper • 2610.02298 • Published 11 days ago • 58
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 16 days ago • 326
When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections Paper • 2609.32488 • Published 16 days ago • 27
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks Paper • 2609.36965 • Published 13 days ago • 26
Marathoner: Ultra-Long-Horizon Autonomous Intelligence Paper • 2609.34378 • Published 14 days ago • 43
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 13 days ago • 104