Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on... (description from the OpenRouter listing)
| Benchmark | Domain | Score | vs best recorded | Setting | Run | Source |
|---|---|---|---|---|---|---|
| Furniture Assembly | multimodal | 33.3% ±6.1 | 40% | max | 10 Sep 2026 | Epoch ↗ |
| LMCA | agents | 52.2% | 80% | max | — | External ↗ |
| DTBench | reasoning | 94.7% | 96% | max | — | External ↗ |
| Mystery Game Puzzles | games | 28.0% ±4.5 | 33% | max | 25 Jul 2026 | Epoch ↗ |
| EBR-bench | reasoning | 19.0% | 25% | max | 30 Jun 2026 | Epoch ↗ |
| MirrorCode | coding | 31.1% ±9.0 | 42% | high | 12 Aug 2026 | Epoch ↗ |
| OSWorld 2.0 | agents | 18.2% | 58% | max | — | External ↗ |
| FrontierMath-Tiers-1-3-v2-Private | math | 70.2% ±2.7 | 75% | max | 10 Jun 2026 | Eval log ↗ |
| FrontierMath-Tier-4-v2-Private | math | 31.7% ±7.4 | 32% | max | 10 Jun 2026 | Eval log ↗ |
| ProofBench | math | 54.0% | 54% | max | — | External ↗ |
| APEX-Agents | agents | 49.2% | 72% | max | — | External ↗ |
| Chess Puzzles | games | 7.0% ±2.6 | 10% | max | 6 Aug 2026 | Eval log ↗ |
| GSO-Bench | coding | 44.1% | 94% | high | — | External ↗ |
| ARC-AGI-2 | reasoning | 75.8% | 80% | max | — | External ↗ |
| WeirdML | coding | 76.4% ±0.0 | 82% | high | — | External ↗ |
| OTIS Mock AIME 2024-2025 | math | 86.7% ±5.1 | 87% | max | 6 Aug 2026 | Eval log ↗ |
| SWE-Bench verified | coding | 83.5% ±1.7 | 100% | max | 20 Apr 2026 | Eval log ↗ |
| GPQA diamond | science | 86.4% ±2.4 | 90% | max | 6 Aug 2026 | Eval log ↗ |
| ARC-AGI | reasoning | 93.5% | 95% | high | — | External ↗ |
Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.
6 observations (OpenRouter listing + Internet Archive snapshots). History accumulates with every ingest run; a single point means no change has been observed yet.