AI news & updates
840 events
- BenchmarkEpoch AI evaluates Muse Spark 1.3 (xhigh)
OTIS Mock AIME 2024-2025 99.2%
- BenchmarkEpoch AI evaluates Muse Spark 1.3 (xhigh)
FrontierMath-Tiers-1-3-v2-Private 74.4% · FrontierMath-Tier-4-v2-Private 41.5% · Chess Puzzles 35.0% · Mystery Game Puzzles 14.1%
- Open releaseInference.net releases Schematron V2 Turbo
Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in re
- Open releaseInference.net releases Schematron V2 Small
Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema
- Model launchSakana releases Fugu Max
Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
- Model launchSakana releases Fugu Ultra v2
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...
- BenchmarkEpoch AI evaluates Kimi K3
Furniture Assembly 34.2%
- BenchmarkEpoch AI evaluates Kimi K2.6
Furniture Assembly 21.7%
- Open releaseinclusionAI releases Ling 3.0 Flash VL
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
- Open releaseMajorDeepSeek releases DeepSeek V4.1 Flash (batch)
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
- Model launchMajorCognition releases SWE-2 (max)
- BenchmarkEpoch AI evaluates Claude Fable 5.1
MirrorCode 73.3% · Furniture Assembly 70.0%
- BenchmarkEpoch AI evaluates Qwen3.8 Max (0902) (xhigh)
Furniture Assembly 20.0%
- BenchmarkEpoch AI evaluates Gemini 3.7 Flash
Furniture Assembly 26.7%
- BenchmarkEpoch AI evaluates Gemini 3.6 Flash
Furniture Assembly 23.3%
- BenchmarkEpoch AI evaluates Gemini 3.8 Flash
Furniture Assembly 31.7%
- BenchmarkEpoch AI evaluates Gemini 3.1 Pro Preview
Furniture Assembly 26.7%
- BenchmarkEpoch AI evaluates Claude Fable 5
Furniture Assembly 35.8%
- BenchmarkEpoch AI evaluates Claude Opus 5
Furniture Assembly 60.8%
- BenchmarkEpoch AI evaluates Claude Opus 4.8
Furniture Assembly 42.5%
- BenchmarkEpoch AI evaluates Claude Opus 4.7
Furniture Assembly 33.3%
- BenchmarkEpoch AI evaluates Claude Opus 4.6
Furniture Assembly 28.3%
- BenchmarkEpoch AI evaluates Claude Opus 4.5 (64k thinking)
Furniture Assembly 28.3%
- BenchmarkEpoch AI evaluates GPT-5.6 Luna
Furniture Assembly 42.5%
- BenchmarkEpoch AI evaluates GPT-6 Astra
Furniture Assembly 80.0%
- BenchmarkEpoch AI evaluates GPT-5.6 Sol
Furniture Assembly 56.7%
- BenchmarkEpoch AI evaluates GPT-5.6 Terra
Furniture Assembly 54.2%
- BenchmarkEpoch AI evaluates GPT-5.5 (xhigh)
Furniture Assembly 44.2%
- BenchmarkEpoch AI evaluates GPT-5.2 (xhigh)
Furniture Assembly 38.3%
- BenchmarkEpoch AI evaluates GPT-5.4 (xhigh)
Furniture Assembly 37.5%
- Open releaseMajorDeepSeek releases DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
- Open releaseMajorDeepSeek releases DeepSeek V4.1 Flash (unknown)
- Model launchInception releases Mercury 2.5
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
- Model launchInception Labs releases Mercury 2.5 (unknown)
- BenchmarkEpoch AI evaluates Claude Fable 5.1
EBR-bench 57.1%
- Open releaseNex AGI releases Nex-N2.5-Pro
Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
- Open releaseNex AGI releases Nex-N2.5-Mini
Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
- BenchmarkEpoch AI evaluates Claude Opus 5
EBR-bench 45.7%
- Model launchMajorOpenAI releases GPT-6 Astra Pro
GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's do
- Model launchMajorOpenAI releases GPT-6 Astra (batch)
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-hor