Skip to content
Benchmark

Epoch AI evaluates Grok 4.6 (xhigh)

GPQA diamond 93.2% · FrontierMath-Tiers-1-3-v2-Private 66.0% · FrontierMath-Tier-4-v2-Private 31.7% · OTIS Mock AIME 2024-2025 99.2%

Type
BENCHMARK_RESULT
Detection
benchmark
Confidence
high
Status

Primary source

Epoch AI Benchmarking Hub · epoch.ai/benchmarks ↗

Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).

Related events