| Rank | Model | Score |
|---|---|---|
| 1 | deepseek-r1-0528 | 0.981 |
| 2 | qwen2-5-72b-instruct | 0.979 |
| 3 | gpt-4-1 | 0.971 |
| 4 | llama-3-1-405b-instruct | 0.969 |
| 5 | claude-3-5-sonnet | 0.967 |
| 6 | claude-3-opus | 0.964 |
| 7 | r1 | 0.964 |
| 8 | gpt-4-0613 | 0.963 |
| 9 | gpt-4o | 0.963 |
| 10 | phi-4 | 0.96 |
| 11 | llama-3-3-70b-instruct | 0.96 |
| 12 | o4-mini | 0.959 |
| 13 | llama-3-1-70b-instruct | 0.948 |
| 14 | llama-4-maverick | 0.938 |
| 15 | gemini-1-5-pro | 0.914 |
| 16 | llama-4-scout | 0.907 |
1 phaseActive
No evaluations yet for this benchmark.