| Rank | Model | Score |
|---|---|---|
| 1 | minicpm-sala | 0.951 |
| 2 | kimi-k2-0905 | 0.945 |
| 3 | gpt-5-1 | 0.94 |
| 4 | claude-3-5-sonnet | 0.937 |
| 5 | gpt-5-2 | 0.936 |
| 6 | gpt-5 | 0.934 |
| 7 | gpt-5 | 0.934 |
| 8 | kimi-k2-instruct | 0.933 |
| 9 | kimi-k2-instruct | 0.933 |
| 10 | qwen2-5-coder-32b-instruct | 0.927 |
| 11 | qwen2-5-coder-32b-instruct | 0.926 |
| 12 | sarvam-30b | 0.921 |
| 13 | claude-3-5-sonnet-v1 | 0.92 |
| 14 | mistral-large-2 | 0.92 |
| 15 | deepseek-v3-2 | 0.92 |
| 16 | qwen2-5-vl-32b-instruct | 0.915 |
| 17 | gpt-4o | 0.902 |
| 18 | gemini-2-5-pro | 0.9 |
| 19 | deepseek-v3-1 | 0.898 |
| 20 | granite-3-3-8b-base | 0.897 |
| 21 | granite-3-3-8b-instruct | 0.897 |
| 22 | nova-2-pro | 0.89 |
| 23 | llama-3-1-405b-instruct | 0.89 |
| 24 | deepseek-v2-5 | 0.89 |
| 25 | nova-pro | 0.89 |
1 phaseActive
No evaluations yet for this benchmark.