[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| Chess Puzzles | games | 1.0% ±1.0 | 1% | — | 28 Aug 2026 | Eval log ↗ |
| Lech Mazur Writing | other | 62.6% | 73% | — | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 13.8% ±3.7 | 14% | — | 25 Feb 2025 | Eval log ↗ |
| Balrog | games | 11.6% | 17% | — | — | External ↗ |
| GPQA diamond | science | 56.1% ±2.6 | 59% | — | 31 Jan 2025 | Eval log ↗ |
| MATH level 5 | math | 64.9% ±1.1 | 66% | — | 31 Jan 2025 | Epoch ↗ |
| MMLU | knowledge | 84.8% | 96% | — | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
13 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.