| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 4.6 (no thinking) Anthropic | 72.1% | agent: claude-sonnet-4-6 (100 steps) | 17 Feb 2026 | — | Source ↗ | |
| 2 | Claude Opus 4.5 (no thinking) Anthropic | 66.3% | agent: Claude Opus 4.5 | 24 Nov 2025 | — | Source ↗ | |
| 3 | Kimi K2.5 Open source MoonshotAI | 63.3% | agent: Kimi-K2.5 | 27 Jan 2026 | — | Source ↗ | |
| 4 | Claude Sonnet 4.5 (no thinking) Anthropic | 62.9% | agent: claude-sonnet-4-5-20250929 (100 steps) | 29 Sep 2025 | — | Source ↗ | |
| 5 | Claude Sonnet 4 Anthropic | 43.9% | agent: claude-4-sonnet-20250514 (50 steps) | 22 May 2025 | — | Source ↗ | |
| 6 | Claude 3.7 Sonnet Anthropic | 35.8% | agent: claude-3-7-sonnet-20250219 (50 steps) | 24 Feb 2025 | — | Source ↗ | |
| 7 | computer-use-preview-2025-03-11 | 31.3% | agent: computer-use-preview (50 steps) | 11 Mar 2025 | — | Source ↗ | |
| 8 | o3 (medium) OpenAI | 23.0% | agent: o3 (100 steps) | 16 Apr 2025 | — | Source ↗ | |
| 9 | Qwen2.5-72B Open source Alibaba | 5.0% | agent: qwen2.5-vl-72b-instruct (100 steps) | 19 Sep 2024 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.