other benchmark · included in ECI
OpenBookQA
OpenBookQA benchmark (score column: Accuracy).
Results
32
Random baseline
25.0%
Score ceiling
100%
Released
8 Sep 2018
Score by model release date
Best score by organisation
Leaderboard
| # | Model | Score | Relative | Setting | Released | Run | Evidence |
|---|---|---|---|---|---|---|---|
| 1 | phi-3-small 7.4B Open source Microsoft | 88.0% | — | 23 Apr 2024 | — | Source ↗ | |
| 2 | phi-3-mini 3.8B Open source Microsoft | 88.0% | — | 23 Apr 2024 | — | Source ↗ | |
| 3 | phi-3-medium 14B Open source Microsoft | 87.4% | — | 23 Apr 2024 | — | Source ↗ | |
| 4 | GPT-3.5 Turbo (Nov 2023) OpenAI | 86.0% | — | 6 Nov 2023 | — | Source ↗ | |
| 5 | Mixtral 8x7B Open source Mistral AI | 85.8% | — | 11 Dec 2023 | — | Source ↗ | |
| 6 | Llama 3-8B Open source Meta AI | 82.6% | — | 18 Apr 2024 | — | Source ↗ | |
| 7 | Mistral 7B v0.1 Open source Mistral AI | 79.8% | — | 27 Sep 2023 | — | Source ↗ | |
| 8 | Gemma 7B Open source Google DeepMind | 78.6% | — | 21 Feb 2024 | — | Source ↗ | |
| 9 | Phi-2 Open source Microsoft | 73.6% | — | 12 Dec 2023 | — | Source ↗ | |
| 10 | Falcon-180B Open source Technology Innovation Institute | 64.2% | — | 6 Sep 2023 | — | Source ↗ | |
| 11 | Llama 2-70B Open source Meta AI | 60.2% | — | 18 Jul 2023 | — | Source ↗ | |
| 12 | LLaMA-65B Open source Meta AI | 60.2% | — | 24 Feb 2023 | — | Source ↗ | |
| 13 | Llama 2-7B Open source Meta AI | 58.6% | — | 18 Jul 2023 | — | Source ↗ | |
| 14 | LLaMA-33B Open source Meta AI | 58.6% | — | 24 Feb 2023 | — | Source ↗ | |
| 15 | Llama 2-34B Meta AI | 58.2% | — | 18 Jul 2023 | — | Source ↗ | |
| 16 | PaLM 2-M | 57.4% | — | 17 May 2023 | — | Source ↗ | |
| 17 | LLaMA-7B Open source Meta AI | 57.2% | — | 24 Feb 2023 | — | Source ↗ | |
| 18 | Llama 2-13B Open source Meta AI | 57.0% | — | 18 Jul 2023 | — | Source ↗ | |
| 19 | Falcon-40B Open source Technology Innovation Institute | 56.6% | — | 25 May 2023 | — | Source ↗ | |
| 20 | LLaMA-13B Open source Meta AI | 56.4% | — | 24 Feb 2023 | — | Source ↗ | |
| 21 | PaLM 2-S | 56.2% | — | 17 May 2023 | — | Source ↗ | |
| 22 | MPT-30B Open source MosaicML | 52.0% | — | 22 Jun 2023 | — | Source ↗ | |
| 23 | Falcon-7B Open source Technology Innovation Institute | 51.6% | — | 24 Apr 2023 | — | Source ↗ | |
| 24 | MPT-7B Open source MosaicML | 51.4% | — | 5 May 2023 | — | Source ↗ | |
| 25 | XGen-7B Open source Salesforce | 40.2% | — | 27 Jun 2023 | — | Source ↗ | |
| 26 | RedPajama-INCITE-7B-Base | 40.0% | — | 4 May 2023 | — | Source ↗ | |
| 27 | Dolly 2.0-12b Open source Databricks | 39.2% | — | 11 Apr 2023 | — | Source ↗ | |
| 28 | open_llama_7b | 39.0% | 7b | 7 Jun 2023 | — | Source ↗ | |
| 29 | Phi-1.5 Open source Microsoft | 37.2% | 5 | 11 Sep 2023 | — | Source ↗ | |
| 30 | Cerebras-GPT-13B Open source Cerebras Systems | 35.8% | — | 20 Mar 2023 | — | Source ↗ | |
| 31 | vicuna-13b-v1.1 | 33.0% | — | 12 Apr 2023 | — | Source ↗ | |
| 32 | stablelm-tuned-alpha-7b | 32.4% | — | 19 Apr 2023 | — | Source ↗ |
Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.