multimodal benchmark · included in ECI
GeoBench
GeoBench benchmark (score column: ACW Country %).
Results
30
Random baseline
0.0%
Score ceiling
100%
Released
1 Mar 2025
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3 Flash Preview Google | 88.0% | — | 17 Dec 2025 | — | Source ↗ | |
| 2 | Gemini 2.5 Pro Preview (Jun 2025) Google DeepMind | 86.0% | — | 5 Jun 2025 | — | Source ↗ | |
| 3 | Gemini 3 Pro Preview Google DeepMind | 84.0% | — | 18 Nov 2025 | — | Source ↗ | |
| 4 | GPT-5 (medium) OpenAI | 81.0% | medium | 7 Aug 2025 | — | Source ↗ | |
| 5 | Gemini 2.5 Pro Exp (Mar 2025) Google DeepMind | 81.0% | — | 25 Mar 2025 | — | Source ↗ | |
| 6 | o1 (medium) OpenAI | 80.0% | medium | 17 Dec 2024 | — | Source ↗ | |
| 7 | Gemini 2.0 Flash (Feb 2025) Google DeepMind,Google | 77.0% | — | 5 Feb 2025 | — | Source ↗ | |
| 8 | Gemini 2.5 Flash (May 2025) Google DeepMind | 76.0% | — | 20 May 2025 | — | Source ↗ | |
| 9 | Gemini 1.5 Flash (Sep 2024) Google DeepMind | 76.0% | — | 24 Sep 2024 | — | Source ↗ | |
| 10 | Claude Opus 4.5 (no thinking) Anthropic | 75.0% | — | 24 Nov 2025 | — | Source ↗ | |
| 11 | o3 (medium) OpenAI | 74.0% | medium | 16 Apr 2025 | — | Source ↗ | |
| 12 | Gemini 2.5 Flash Preview (Apr 2025) Google DeepMind | 73.0% | — | 17 Apr 2025 | — | Source ↗ | |
| 13 | GPT-4.1 OpenAI | 72.0% | — | 14 Apr 2025 | — | Source ↗ | |
| 14 | GPT-4o (Nov 2024) OpenAI | 71.0% | — | 20 Nov 2024 | — | Source ↗ | |
| 15 | Claude 3.7 Sonnet Anthropic | 68.0% | — | 24 Feb 2025 | — | Source ↗ | |
| 16 | Claude 3.7 Sonnet (15k thinking) Anthropic | 65.0% | 15K | 24 Feb 2025 | — | Source ↗ | |
| 17 | o4 Mini OpenAI | 64.0% | high | 16 Apr 2025 | — | Source ↗ | |
| 18 | o4-mini (medium) OpenAI | 64.0% | medium | 16 Apr 2025 | — | Source ↗ | |
| 19 | GPT-4o-mini OpenAI | 64.0% | — | 18 Jul 2024 | — | Source ↗ | |
| 20 | Claude 3.5 Sonnet (Oct 2024) Anthropic | 62.0% | — | 22 Oct 2024 | — | Source ↗ | |
| 21 | Qwen2.5-72B Open source Alibaba | 62.0% | — | 19 Sep 2024 | — | Source ↗ | |
| 22 | o3 OpenAI | 60.0% | high | 16 Apr 2025 | — | Source ↗ | |
| 23 | Llama 4 Maverick Open source Meta | 52.0% | — | 6 Apr 2025 | — | Source ↗ | |
| 24 | Gemma 3 27B Open source Google | 52.0% | — | 12 Mar 2025 | — | Source ↗ | |
| 25 | Llama 3.2 90B Open source Meta AI | 52.0% | — | 24 Sep 2024 | — | Source ↗ | |
| 26 | Claude Opus 4 Anthropic | 49.0% | 32K | 22 May 2025 | — | Source ↗ | |
| 27 | Grok 4 xAI | 45.0% | — | 9 Jul 2025 | — | Source ↗ | |
| 28 | Claude Sonnet 4 Anthropic | 37.0% | — | 22 May 2025 | — | Source ↗ | |
| 29 | Pixtral 12B Open source Mistral AI | 37.0% | — | 17 Sep 2024 | — | Source ↗ | |
| 30 | Claude 3.5 Haiku (Oct 2024) Anthropic | 34.0% | — | 22 Oct 2024 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.