Skip to content

Grok 4

ModelActive
by xAI · family “grok”
Epoch Capabilities Index
146.5 #69 of 268
90% CI 144.4 – 148.1
Released
9 Jul 2025
Context window
Unknown
Max output
Unknown
Input $ / 1M tokens
Unknown
Output $ / 1M tokens
Unknown
Cached input $ / 1M
Unknown
Each benchmark shown separately with its own source

Benchmark results (19)

BenchmarkDomainScorevs best recordedSettingRunSource
Chess Puzzlesgames28.0% ±4.5
39%
—30 Jan 2026Eval log ↗
GDPvalagents21.1%
42%
high—External ↗
FrontierMath-Tier-4-2025-07-01-Privatesupersededmath2.1% ±2.1
4%
—11 Aug 2025Epoch ↗
DeepResearch Benchagents47.3%
86%
agent: True—External ↗
Terminal Benchagents27.2% ±3.1
32%
agent: OpenHands2 Nov 2025External ↗
ARC-AGI-2reasoning16.0%
17%
——External ↗
METR Time Horizonsagents66.6%
78%
——External ↗
GeoBenchmultimodal45.0%
51%
——External ↗
FrontierMath-2025-02-28-Privatesupersededmath19.7% ±2.3
38%
—13 Nov 2025Epoch ↗
Fiction.LiveBenchlong-context94.4%
97%
——External ↗
Lech Mazur Writingother76.9%
89%
——External ↗
WeirdMLcoding45.7% ±0.0
49%
——External ↗
Aider polyglotcoding79.6%
90%
——External ↗
OTIS Mock AIME 2024-2025math84.0% ±5.0
84%
——Epoch ↗
Balroggames43.6%
64%
——External ↗
SimpleBenchreasoning60.5%
74%
——External ↗
Cybenchcoding43.0%
46%
——External ↗
GPQA diamondscience87.0% ±2.0
91%
——Epoch ↗
ARC-AGIreasoning66.7%
68%
——External ↗

Source: Epoch AI Benchmarking Hub (CC BY 4.0). “External” rows are leaderboard results Epoch collects from third parties. Best reported setting per benchmark is shown.

API list price over time

Price history

$ per 1M tokens
This model has no public API price listing

0 recorded prices on — (OpenRouter listing); re-read on every data refresh, most recently 8h ago. Steps show when the price changed.

Events

No events recorded yet.