AI news & updates
266 events
- BenchmarkOpen TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice CloningHugging Face✓ primaryHugging Face
- BenchmarkEpoch AI evaluates Claude Sonnet 5.5
GPQA diamond 95.6% · FrontierMath-Tiers-1-3-v2-Private 88.8% · FrontierMath-Tier-4-v2-Private 80.5% · OTIS Mock AIME 2024-2025 100.0%
- BenchmarkEpoch AI evaluates Claude Opus 5.5
Mystery Game Puzzles 71.0%
- BenchmarkEpoch AI evaluates GPT-6 Luna
Furniture Assembly 44.2%
- BenchmarkEpoch AI evaluates Grok 4.5
Furniture Assembly 22.5%
- BenchmarkEpoch AI evaluates Grok 4.6 (xhigh)
Furniture Assembly 40.0%
- BenchmarkEpoch AI evaluates Grok 4.7 (xhigh)
Furniture Assembly 20.8%
- BenchmarkEpoch AI evaluates GPT-6 Luna
Furniture Assembly 28.3%
- BenchmarkEpoch AI evaluates GPT-6 Sol
Furniture Assembly 58.3%
- BenchmarkEpoch AI evaluates Claude Opus 5.5
EBR-bench 71.4% · Furniture Assembly 83.3%
- BenchmarkEpoch AI evaluates GPT-6 Sol
EBR-bench 53.3%
- BenchmarkHow UK AISI and EvalEval Are Making Benchmark Results ReproducibleHugging Face✓ primaryHugging Face
- BenchmarkEpoch AI evaluates Muse Spark 1.3
FrontierMath-Tiers-1-3-v2-Private 74.0% · FrontierMath-Tier-4-v2-Private 46.3% · Chess Puzzles 38.0% · Mystery Game Puzzles 25.0%
- BenchmarkEpoch AI evaluates Muse Spark 1.3 (xhigh)
OTIS Mock AIME 2024-2025 99.2%
- BenchmarkEpoch AI evaluates Muse Spark 1.3 (xhigh)
FrontierMath-Tiers-1-3-v2-Private 74.4% · FrontierMath-Tier-4-v2-Private 41.5% · Chess Puzzles 35.0% · Mystery Game Puzzles 14.1%
- BenchmarkEpoch AI evaluates Kimi K3
Furniture Assembly 34.2%
- BenchmarkEpoch AI evaluates Kimi K2.6
Furniture Assembly 21.7%
- BenchmarkEpoch AI evaluates Claude Fable 5.1
MirrorCode 73.3% · Furniture Assembly 70.0%
- BenchmarkEpoch AI evaluates Qwen3.8 Max (0902) (xhigh)
Furniture Assembly 20.0%
- BenchmarkEpoch AI evaluates Gemini 3.7 Flash
Furniture Assembly 26.7%
- BenchmarkEpoch AI evaluates Gemini 3.6 Flash
Furniture Assembly 23.3%
- BenchmarkEpoch AI evaluates Gemini 3.8 Flash
Furniture Assembly 31.7%
- BenchmarkEpoch AI evaluates Gemini 3.1 Pro Preview
Furniture Assembly 26.7%
- BenchmarkEpoch AI evaluates Claude Fable 5
Furniture Assembly 35.8%
- BenchmarkEpoch AI evaluates Claude Opus 5
Furniture Assembly 60.8%
- BenchmarkEpoch AI evaluates Claude Opus 4.8
Furniture Assembly 42.5%
- BenchmarkEpoch AI evaluates Claude Opus 4.7
Furniture Assembly 33.3%
- BenchmarkEpoch AI evaluates Claude Opus 4.6
Furniture Assembly 28.3%
- BenchmarkEpoch AI evaluates Claude Opus 4.5 (64k thinking)
Furniture Assembly 28.3%
- BenchmarkEpoch AI evaluates GPT-5.6 Luna
Furniture Assembly 42.5%
- BenchmarkEpoch AI evaluates GPT-6 Astra
Furniture Assembly 80.0%
- BenchmarkEpoch AI evaluates GPT-5.6 Sol
Furniture Assembly 56.7%
- BenchmarkEpoch AI evaluates GPT-5.6 Terra
Furniture Assembly 54.2%
- BenchmarkEpoch AI evaluates GPT-5.5 (xhigh)
Furniture Assembly 44.2%
- BenchmarkEpoch AI evaluates GPT-5.2 (xhigh)
Furniture Assembly 38.3%
- BenchmarkEpoch AI evaluates GPT-5.4 (xhigh)
Furniture Assembly 37.5%
- BenchmarkEpoch AI evaluates Claude Fable 5.1
EBR-bench 57.1%
- BenchmarkEpoch AI evaluates Claude Opus 5
EBR-bench 45.7%
- BenchmarkEpoch AI evaluates GPT-5.6 Sol
EBR-bench 44.8%
- BenchmarkEpoch AI evaluates Qwen3.8 Max (0902) (xhigh)
GPQA diamond 92.3% · FrontierMath-Tiers-1-3-v2-Private 65.6% · FrontierMath-Tier-4-v2-Private 34.1% · OTIS Mock AIME 2024-2025 100.0%