Skip to content
Benchmark

Epoch AI evaluates Mistral Small 3.2

OTIS Mock AIME 2024-2025 30.3% · Chess Puzzles 1.0%

Type
BENCHMARK_RESULT
Detection
benchmark
Confidence
high
Status

Primary source

Epoch AI Benchmarking Hub · logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F ↗

Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).

Related events