other benchmark · included in ECI
ANLI
ANLI benchmark (score column: Score).
Results
9
Random baseline
33.3%
Score ceiling
100%
Released
31 Oct 2019
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | phi-3-small 7.4B Open source Microsoft | 58.1% | — | 23 Apr 2024 | — | Source ↗ | |
| 2 | GPT-3.5 Turbo (Nov 2023) OpenAI | 58.1% | — | 6 Nov 2023 | — | Source ↗ | |
| 3 | Llama 3-8B Open source Meta AI | 57.3% | — | 18 Apr 2024 | — | Source ↗ | |
| 4 | phi-3-medium 14B Open source Microsoft | 55.8% | — | 23 Apr 2024 | — | Source ↗ | |
| 5 | Mixtral 8x7B Open source Mistral AI | 55.2% | — | 11 Dec 2023 | — | Source ↗ | |
| 6 | phi-3-mini 3.8B Open source Microsoft | 52.8% | — | 23 Apr 2024 | — | Source ↗ | |
| 7 | Gemma 7B Open source Google DeepMind | 48.7% | — | 21 Feb 2024 | — | Source ↗ | |
| 8 | Mistral 7B v0.1 Open source Mistral AI | 47.1% | — | 27 Sep 2023 | — | Source ↗ | |
| 9 | Phi-2 Open source Microsoft | 42.5% | — | 12 Dec 2023 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.