| Model | Total Params. | MMMU | MathVision | ZeroBench(sub) | DYNAMATH | SimpleVQA | HallusionBench | AIME25 | HMMT25 | CNMO24 | GPQA-Diamond | LiveCodeBench (24.8-25.5) |
|
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Open-Source VLM | Step3 | 321B | 74.2 | 64.8 | 23.0 | 50.1 | 62.2 | 64.2 | 82.9 | 70.0 | 83.7 | 73.0 | 67.1 |
| ERINE4.5 - thinking | 300B/424B | 70.0 | 47.6 | 22.5 | 46.9 | 59.8 | 60.0 | 35.1 | 40.5* | 75.5 | 76.8 | 38.8 | |
| GLM-4.1V-thinking | 9B | 68.0 | 49.4 | 22.8 | 41.9 | 48.1 | 60.8 | 13.3 | 6.7 | 25.0 | 47.4 | 24.2 | |
| MiMo-VL | 7B | 66.7 | 60.4 | 18.6 | 45.9 | 48.5 | 59.6 | 60.0 | 34.6 | 69.9 | 55.5 | 50.1 | |
| QvQ-72B-Preview | 72B | 70.3 | 35.9 | 15.9 | 30.7 | 40.3 | 50.8 | 22.7 | 49.5 | 47.3 | 10.9 | 24.1 | |
| LLaMA-Maverick | 400B | 73.4 | 47.2 | 22.8 | 47.1 | 45.4 | 57.1 | 19.2 | 8.91 | 41.6 | 69.8 | 33.9 | |
| Open-Source LLM | MiniMax-M1-80k | 456B | - | - | - | - | - | - | 76.9 | - | - | 70.0 | 65.0 |
| Qwen3-235B-A22B-Thinking | 235B | - | - | - | - | - | - | 81.5 | 62.5 | - | 71.1 | 65.9 | |
| DeepSeek R1-0528 | 671B | - | - | - | - | - | - | 87.5 | 79.4 | 86.9 | 81.0 | 73.3 | |
| Qwen3-235B-A22B-Thinking-2507 | 235B | - | - | - | - | - | - | 92.3 | 83.9 | - | 81.1 | - | |
| Proprietary VLM | O3 | - | 82.9 | 72.8 | 25.2 | 58.1 | 59.8 | 60.1 | 88.9 | 70.1 | 86.7 | 83.3 | 75.8 |
| Claude4 Sonnet (thinking) | - | 76.9 | 64.6 | 26.1 | 48.1 | 43.7 | 57.0 | 70.5 | - | - | 75.4 | 55.9 | |
| Claude4 opus (thinking) | - | 79.8 | 66.1 | 25.2 | 49.3 | 47.2 | 59.9 | 75.5 | - | - | 79.6 | 56.6 | |
| Gemini 2.5 Flash (thinking) | - | 73.2 | 57.3 | 20.1 | 57.1 | 61.1 | 65.2 | 72.0 | - | - | 82.8 | 61.9 | |
| Gemini 2.5 Pro | - | 81.7 | 73.3 | 30.8 | 56.3 | 66.8 | 66.8 | 88.0 | - | - | 86.4 | 71.8 | |
| Grok 4 | - | 80.9 | 70.3 | 22.5 | 40.7 | 55.9 | 64.8 | 98.8 | 93.9 | 85.5 | 87.5 | 79.3 |