Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| LMCA | agents | 15.9% | 24% | — | — | External ↗ |
| DTBench | reasoning | 61.9% | 63% | — | — | External ↗ |
| ARC-AGI-2 | reasoning | 0.0% | 0% | — | — | External ↗ |
| GeoBench | multimodal | 52.0% | 59% | — | — | External ↗ |
| Fiction.LiveBench | long-context | 46.2% | 48% | — | — | External ↗ |
| Lech Mazur Writing | other | 62.0% | 72% | — | — | External ↗ |
| HLE | knowledge | 5.7% | 10% | — | — | External ↗ |
| WeirdML | coding | 24.5% ±0.0 | 26% | — | — | External ↗ |
| Aider polyglot | coding | 15.6% | 18% | — | — | External ↗ |
| SimpleBench | reasoning | 27.7% | 34% | — | — | External ↗ |
| ARC-AGI | reasoning | 4.4% | 4% | — | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
13 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.