Benchmark
Epoch AI evaluates GLM 5.3
GPQA diamond 90.9% · OTIS Mock AIME 2024-2025 91.1% · Chess Puzzles 21.0%
- Type
- BENCHMARK_RESULT
- Detection
- benchmark
- Confidence
- high
- Status
Related entities
GPQA diamond 90.9% · OTIS Mock AIME 2024-2025 91.1% · Chess Puzzles 21.0%
Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).