other benchmark · included in ECI
PIQA
PIQA benchmark (score column: Score).
Results
47
Random baseline
50.0%
Score ceiling
100%
Released
26 Nov 2019
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | GPT-4o-mini OpenAI | 88.7% | — | 18 Jul 2024 | — | Source ↗ | |
| 2 | Phi-3.5-MoE Open source Microsoft | 88.6% | — | 17 Aug 2024 | — | Source ↗ | |
| 3 | Gemini 1.5 Flash (Sep 2024) Google DeepMind | 87.5% | — | 24 Sep 2024 | — | Source ↗ | |
| 4 | Llama 3.1-405B Open source Meta AI | 85.9% | — | 23 Jul 2024 | — | Source ↗ | |
| 5 | Falcon-180B Open source Technology Innovation Institute | 84.9% | — | 6 Sep 2023 | — | Source ↗ | |
| 6 | DeepSeek V3 Open source DeepSeek | 84.7% | — | 26 Dec 2024 | — | Source ↗ | |
| 7 | Inflection-1 Inflection AI | 84.2% | — | 22 Jun 2023 | — | Source ↗ | |
| 8 | DeepSeek-V2 (MoE-236B, May 2024) Open source DeepSeek | 83.9% | — | 7 May 2024 | — | Source ↗ | |
| 9 | Gemma 2 9B Open source Google DeepMind | 83.7% | — | 24 Jun 2024 | — | Source ↗ | |
| 10 | Mixtral 8x7B Open source Mistral AI | 83.6% | — | 11 Dec 2023 | — | Source ↗ | |
| 11 | Mistral Nemo Open source Mistral | 83.5% | — | 18 Jul 2024 | — | Source ↗ | |
| 12 | Stable Beluga 2 Open source Stability AI | 83.3% | — | 20 Jul 2023 | — | Source ↗ | |
| 13 | Mistral 7B v0.1 Open source Mistral AI | 83.0% | — | 27 Sep 2023 | — | Source ↗ | |
| 14 | Falcon-40B Open source Technology Innovation Institute | 83.0% | — | 25 May 2023 | — | Source ↗ | |
| 15 | Llama 2-70B Open source Meta AI | 82.8% | — | 18 Jul 2023 | — | Source ↗ | |
| 16 | LLaMA-65B Open source Meta AI | 82.8% | — | 24 Feb 2023 | — | Source ↗ | |
| 17 | Qwen2.5-72B Open source Alibaba | 82.6% | — | 19 Sep 2024 | — | Source ↗ | |
| 18 | Nemotron-4 15B NVIDIA | 82.4% | — | 26 Feb 2024 | — | Source ↗ | |
| 19 | LLaMA-33B Open source Meta AI | 82.3% | — | 24 Feb 2023 | — | Source ↗ | |
| 20 | Mistral 7B v0.2 Open source Mistral AI | 82.2% | — | 11 Dec 2023 | — | Source ↗ | |
| 21 | Llama 2-34B Meta AI | 81.9% | — | 18 Jul 2023 | — | Source ↗ | |
| 22 | MPT-30B Open source MosaicML | 81.9% | — | 22 Jun 2023 | — | Source ↗ | |
| 23 | Llama 3.1-8B Open source Meta AI | 81.2% | — | 23 Jul 2024 | — | Source ↗ | |
| 24 | Gemma 7B Open source Google DeepMind | 81.2% | — | 21 Feb 2024 | — | Source ↗ | |
| 25 | Phi-3.5-mini Open source Microsoft | 81.0% | — | 16 Aug 2024 | — | Source ↗ | |
| 26 | Llama 2-13B Open source Meta AI | 80.8% | — | 18 Jul 2023 | — | Source ↗ | |
| 27 | MPT-7B Open source MosaicML | 80.6% | — | 5 May 2023 | — | Source ↗ | |
| 28 | internlm-20b | 80.3% | — | 18 Sep 2023 | — | Source ↗ | |
| 29 | Falcon-7B Open source Technology Innovation Institute | 80.3% | — | 24 Apr 2023 | — | Source ↗ | |
| 30 | LLaMA-13B Open source Meta AI | 80.1% | — | 24 Feb 2023 | — | Source ↗ | |
| 31 | Qwen-14B Open source Alibaba | 79.9% | — | 24 Sep 2023 | — | Source ↗ | |
| 32 | LLaMA-7B Open source Meta AI | 79.8% | — | 24 Feb 2023 | — | Source ↗ | |
| 33 | Llama 2-7B Open source Meta AI | 78.8% | — | 18 Jul 2023 | — | Source ↗ | |
| 34 | Baichuan2-13B Open source Baichuan | 78.1% | — | 6 Sep 2023 | — | Source ↗ | |
| 35 | Qwen-7B Open source Alibaba | 77.9% | — | 28 Sep 2023 | — | Source ↗ | |
| 36 | internlm-7b | 77.9% | — | 5 Jul 2023 | — | Source ↗ | |
| 37 | vicuna-13b-v1.1 | 77.4% | — | 12 Apr 2023 | — | Source ↗ | |
| 38 | Gemma 2B Open source Google DeepMind | 77.3% | — | 21 Feb 2024 | — | Source ↗ | |
| 39 | RedPajama-INCITE-7B-Base | 76.9% | — | 4 May 2023 | — | Source ↗ | |
| 40 | Baichuan1-7B Open source Baichuan | 76.2% | — | 1 Jun 2023 | — | Source ↗ | |
| 41 | open_llama_7b | 76.0% | 7b | 7 Jun 2023 | — | Source ↗ | |
| 42 | XGen-7B Open source Salesforce | 75.5% | — | 27 Jun 2023 | — | Source ↗ | |
| 43 | Dolly 2.0-12b Open source Databricks | 75.4% | — | 11 Apr 2023 | — | Source ↗ | |
| 44 | Cerebras-GPT-13B Open source Cerebras Systems | 73.5% | — | 20 Mar 2023 | — | Source ↗ | |
| 45 | Qwen-1_8B | 73.3% | 8B | 30 Nov 2023 | — | Source ↗ | |
| 46 | chatglm2-6b | 69.6% | — | 24 Jun 2023 | — | Source ↗ | |
| 47 | stablelm-tuned-alpha-7b | 65.8% | — | 19 Apr 2023 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.