Skip to content
Benchmark

Epoch AI evaluates DeepSeek-R1-Distill-Qwen-14B

OTIS Mock AIME 2024-2025 50.6% · Chess Puzzles 1.0%

Type
BENCHMARK_RESULT
Detection
benchmark
Confidence
high
Status

Primary source

Epoch AI Benchmarking Hub · logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F ↗

Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).