other benchmark · included in ECI
FrontierSWE
FrontierSWE benchmark (score column: Score).
Results
15
Random baseline
0.0%
Score ceiling
100%
Released
2 Sep 2026
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra OpenAI | 65.5% | max | 3 Sep 2026 | — | Source ↗ | |
| 2 | Claude Opus 5.5 Anthropic | 62.3% | max | 22 Sep 2026 | — | Source ↗ | |
| 3 | Claude Fable 5.1 Anthropic | 56.3% | max | 1 Sep 2026 | — | Source ↗ | |
| 4 | Claude Opus 5 Anthropic | 52.0% | max | 24 Jul 2026 | — | Source ↗ | |
| 5 | Claude Fable 5 Anthropic | 47.0% | max | 9 Jun 2026 | — | Source ↗ | |
| 6 | GPT-5.6 Sol OpenAI | 32.2% | max | 9 Jul 2026 | — | Source ↗ | |
| 7 | GLM 5.3 Open source Z.ai | 30.2% | max | 14 Aug 2026 | — | Source ↗ | |
| 8 | Grok 4.7 (xhigh) xAI | 29.5% | xhigh | 21 Sep 2026 | — | Source ↗ | |
| 9 | Kimi K3 Open source MoonshotAI | 25.9% | max | 16 Jul 2026 | — | Source ↗ | |
| 10 | Grok 4.6 (xhigh) xAI | 25.3% | xhigh | 12 Aug 2026 | — | Source ↗ | |
| 11 | Gemini 3.7 Flash Google | 20.3% | high | 13 Aug 2026 | — | Source ↗ | |
| 12 | Gemini 3.8 Flash Google | 19.6% | high | 2 Sep 2026 | — | Source ↗ | |
| 13 | Qwen3.8 Max (xhigh) Alibaba | 15.8% | xhigh | 2 Aug 2026 | — | Source ↗ | |
| 14 | Muse Spark 1.2 (xhigh) Meta AI | 12.0% | xhigh | 5 Aug 2026 | — | Source ↗ | |
| 15 | Inkling (xhigh) Open source Thinking Machines | 4.1% | xhigh | 15 Jul 2026 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.