Claude Opus 4.1
ModelActiveby Anthropic · family “claude-opus”
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains... (description from the OpenRouter listing)
in: imagein: textin: fileout: textReasoningTool use
Epoch Capabilities Index
144.1 #86 of 268
90% CI 141.5 – 145.8
Released
5 Aug 2025
Context window
200K tokens
Max output
32K tokens
Input $ / 1M tokens
$15.0
Output $ / 1M tokens
$75.0
Cached input $ / 1M
$1.50
Each benchmark shown separately with its own source
Benchmark results (19)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| Mystery Game Puzzles | games | 21.0% ±4.1 | 25% | 24K | 25 Jul 2026 | Epoch ↗ |
| EBR-bench | reasoning | 7.9% | 10% | — | 25 Jun 2026 | Epoch ↗ |
| FrontierMath-Tiers-1-3-v2-Private | math | 12.6% ±2.0 | 13% | 32K | 11 Jun 2026 | Eval log ↗ |
| FrontierMath-Tier-4-v2-Private | math | 2.4% ±2.4 | 2% | 32K | 11 Jun 2026 | Eval log ↗ |
| Chess Puzzles | games | 7.0% ±2.6 | 10% | — | 20 Jul 2026 | Eval log ↗ |
| GDPval | agents | 43.6% | 88% | — | — | External ↗ |
| FrontierMath-Tier-4-2025-07-01-Privatesuperseded | math | 4.2% ±2.9 | 9% | 27K | 5 Aug 2025 | Epoch ↗ |
| DeepResearch Bench | agents | 48.3% | 87% | — | — | External ↗ |
| Terminal Bench | agents | 38.0% ±2.6 | 45% | agent: Terminus 2 | 31 Oct 2025 | External ↗ |
| METR Time Horizons | agents | 66.8% | 78% | 16K | — | External ↗ |
| FrontierMath-2025-02-28-Privatesuperseded | math | 7.2% ±1.5 | 14% | 27K | 5 Aug 2025 | Epoch ↗ |
| Lech Mazur Writing | other | 84.7% | 98% | — | — | External ↗ |
| VPCT | multimodal | 35.0% | 38% | — | — | External ↗ |
| WeirdML | coding | 45.9% ±0.0 | 49% | 16K | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 68.9% ±7.0 | 69% | 27K | 5 Aug 2025 | Eval log ↗ |
| SimpleBench | reasoning | 60.0% | 73% | — | — | External ↗ |
| Cybench | coding | 42.0% | 45% | — | — | External ↗ |
| SWE-Bench verified | coding | 73.3% ±2.0 | 88% | — | 11 Feb 2026 | Eval log ↗ |
| GPQA diamond | science | 77.3% ±3.0 | 81% | 16K | 5 Aug 2025 | Eval log ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
API list price over time
Price history
$ per 1M tokens
13 recorded prices on 1 Oct 2025 (OpenRouter listing + Internet Archive snapshots); re-read on every data refresh, most recently 6h ago. Steps show when the price changed.
Events
No events recorded yet.