Skip to content
other benchmark · included in ECI

ANLI

ANLI benchmark (score column: Score).
Results
9
Random baseline
33.3%
Score ceiling
100%
Released
31 Oct 2019

Score by model release date

Best score by organisation

#ModelScoreRelativeEvidence
1phi-3-small 7.4B Open source
Microsoft
58.1%Source ↗
2GPT-3.5 Turbo (Nov 2023)
OpenAI
58.1%Source ↗
3Llama 3-8B Open source
Meta AI
57.3%Source ↗
4phi-3-medium 14B Open source
Microsoft
55.8%Source ↗
5Mixtral 8x7B Open source
Mistral AI
55.2%Source ↗
6phi-3-mini 3.8B Open source
Microsoft
52.8%Source ↗
7Gemma 7B Open source
Google DeepMind
48.7%Source ↗
8Mistral 7B v0.1 Open source
Mistral AI
47.1%Source ↗
9Phi-2 Open source
Microsoft
42.5%Source ↗

Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.