Skip to content

o3 (medium)

ModelActive
by OpenAI · family “o3-medium”
Epoch Capabilities Index
Not scored by Epoch AI
Released
16 Apr 2025
Context window
Unknown
Max output
Unknown
Input $ / 1M tokens
Unknown
Output $ / 1M tokens
Unknown
Cached input $ / 1M
Unknown
Each benchmark shown separately with its own source

Benchmark results (20)

BenchmarkDomainScorevs best recordedSettingRunSource
Mystery Game Puzzlesgames23.0% ±4.2
27%
medium27 Aug 2026Epoch ↗
FrontierMath-Tiers-1-3-v2-Privatemath29.8% ±2.7
32%
medium27 Aug 2026Eval log ↗
Chess Puzzlesgames38.0% ±4.9
53%
medium7 Aug 2026Eval log ↗
GDPvalagents30.8%
62%
medium—External ↗
DeepResearch Benchagents45.2%
82%
medium—External ↗
CadEvalcoding74.0%
100%
medium—External ↗
ARC-AGI-2reasoning3.0%
3%
medium—External ↗
METR Time Horizonsagents65.4%
77%
medium—External ↗
GeoBenchmultimodal74.0%
84%
medium—External ↗
FrontierMath-2025-02-28-Privatesupersededmath16.9% ±2.2
32%
medium16 Nov 2025Epoch ↗
Fiction.LiveBenchlong-context88.9%
91%
medium—External ↗
Lech Mazur Writingother83.9%
98%
medium—External ↗
VPCTmultimodal52.0%
57%
medium—External ↗
HLEknowledge19.2%
35%
medium—External ↗
Aider polyglotcoding76.9%
87%
medium—External ↗
OTIS Mock AIME 2024-2025math84.4% ±5.5
84%
medium7 Aug 2026Eval log ↗
SWE-Bench verifiedcoding62.3% ±2.2
75%
medium12 Feb 2026Eval log ↗
OSWorldagents23.0%
32%
agent: o3 (100 steps)—External ↗
GPQA diamondscience80.8% ±2.8
84%
medium7 Aug 2026Eval log ↗
ARC-AGIreasoning53.8%
55%
medium—External ↗

Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.

API list price over time

Price history

$ per 1M tokens
This model has no public API price listing

0 recorded prices on — (OpenRouter listing); re-read on every data refresh, most recently 8h ago. Steps show when the price changed.

Events