Inference Providers
Active filters: rlvr
junshim/When2Think-ThinkOnly-1.5B
Text Generation
• 2B • Updated • 366
• 7
Text Generation
• 2B • Updated • 821
• 3
thuml/webarena-world-model-rlvr
2B • Updated • 22
• 2
ngqtrung/video-hopchain-8b
Video-Text-to-Text
• 9B • Updated • 66
• 1
SultanR/SmolTulu-1.7b-Reinforced-GGUF
Text Generation
• 2B • Updated • 25
• 1
thuml/rt1-world-model-multi-step-rlvr
0.1B • Updated • 24
thuml/rt1-world-model-single-step-rlvr
0.1B • Updated • 11
thuml/bytesized32-world-model-rlvr-binary-reward
2B • Updated • 19
thuml/bytesized32-world-model-rlvr-task-specific-reward
2B • Updated • 24
DebateLabKIT/Llama-3.1-Argunaut-1-8B-HIRPO
Text Generation
• 8B • Updated • 43
• 1
Question Answering
• 4B • Updated • 37
• 3
thinkwee/NOVER1-Qwen2.5-7B
Question Answering
• 8B • Updated • 28
• 2
mradermacher/NOVER1-Qwen3-4B-GGUF
4B • Updated • 219
• 1
mradermacher/NOVER1-Qwen2.5-7B-GGUF
8B • Updated • 590
• 1
mradermacher/NOVER1-Qwen3-4B-i1-GGUF
4B • Updated • 1.11k
• 1
mradermacher/NOVER1-Qwen2.5-7B-i1-GGUF
8B • Updated • 1.06k
• 1
DebateLabKIT/Phi-4-Argunaut-1-HIRPO
Text Generation
• 415k • Updated • 28
mradermacher/Llama-3.1-Argunaut-1-8B-HIRPO-GGUF
8B • Updated • 608
• 1
mradermacher/Llama-3.1-Argunaut-1-8B-HIRPO-i1-GGUF
8B • Updated • 1.07k
• 1
Text Generation
• 2B • Updated • 107
• 10
Text Generation
• 4B • Updated • 15
• 1
mradermacher/airesupdated-v2-GGUF
Reinforcement Learning
• 4B • Updated • 494
ABaroian/Apertus-8B-RLVR-GSM
Text Generation
• Updated • 2
Anonymouslolol/qwen3-8B-hanabi-step110
Reinforcement Learning
• Updated • 2
Aletheia-Bench/GRPO-Think-7B-16k
Text Generation
• 8B • Updated • 72
Aletheia-Bench/GRPO-Think-1.5B-16k
Text Generation
• 2B • Updated • 71
Aletheia-Bench/GRPO-Think-14B-16k
Text Generation
• 15B • Updated • 64
Text Generation
• 4B • Updated • 1
Aletheia-Bench/DPO-Think-1.5B
Text Generation
• 2B • Updated • 75
Aletheia-Bench/DPO-Think-14B
Text Generation
• 15B • Updated • 123
• 2