Skip to content
Benchmark

Epoch AI evaluates Llama 3-8B

Chess Puzzles 0.0%

Type
BENCHMARK_RESULT
Detection
benchmark
Confidence
high
Status

Primary source

Epoch AI Benchmarking Hub · logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F ↗

Independent benchmark runs (GPQA Diamond, SWE-bench Verified, FrontierMath, ...) plus externally reported leaderboards and the Epoch Capabilities Index (ECI).

Related events