AI news & updates
824 events
- Model launchMajorOpenAI releases GPT-6 Sol Pro
GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: ht
- Model launchMajorOpenAI releases GPT-6 Luna Pro
GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs:
- Model launchMajorOpenAI releases GPT-6 Sol (batch)
GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional...
- Model launchMajorOpenAI releases GPT-6 Luna (batch)
GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic...
- Model launchMajorAnthropic releases Claude Opus 5.5
Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...
- Model launchMajorOpenAI releases GPT-6 Sol Pro (batch)
GPT-6 Sol Pro is the same underlying model as [GPT-6 Sol](https://openrouter.ai/openai/gpt-6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: ht
- Model launchMajorOpenAI releases GPT-6 Luna Pro (batch)
GPT-6 Luna Pro is the same underlying model as [GPT-6 Luna](https://openrouter.ai/openai/gpt-6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs:
- Model launchMajorAnthropic releases Claude Opus 5.5 (batch)
Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code...
- BenchmarkEpoch AI evaluates Claude Opus 5.5
EBR-bench 71.4% · Furniture Assembly 83.3%
- BenchmarkEpoch AI evaluates GPT-6 Sol
EBR-bench 53.3%
- Model launchMajorAnthropic releases Claude Opus 5.5 (xhigh)
- Model launchMajorOpenAI releases GPT-6 Sol (xhigh)
- Model launchMajorOpenAI releases GPT-6 Luna (xhigh)
- Model launchMajorSpaceXAI releases Grok 4.7
Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...
- Open releaseXiaomi releases MiMo-V2.6-Pro
MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...
- Open releaseXiaomi releases MiMo-V2.6-Flash
MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...
- Model launchMajorQwen releases Qwen3.8 Omni Flash
Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization,...
- Model launchXiaomi releases MiMo-V2.6-Pro-UltraSpeed
MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly
- Model launchMajorxAI releases Grok 4.7 (unknown)
- Model launchMajorxAI releases Grok 4.7 (xhigh)
- Model launchZ.ai releases GLM 5.3 FlashX
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
- Open releasePrismML releases Ternary Bonsai 2 27B
Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks.
- BenchmarkEpoch AI evaluates Muse Spark 1.3
FrontierMath-Tiers-1-3-v2-Private 74.0% · FrontierMath-Tier-4-v2-Private 46.3% · Chess Puzzles 38.0% · Mystery Game Puzzles 25.0%
- Model launchunbiased releases Pareto
Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.
- BenchmarkEpoch AI evaluates Muse Spark 1.3 (xhigh)
OTIS Mock AIME 2024-2025 99.2%
- BenchmarkEpoch AI evaluates Muse Spark 1.3 (xhigh)
FrontierMath-Tiers-1-3-v2-Private 74.4% · FrontierMath-Tier-4-v2-Private 41.5% · Chess Puzzles 35.0% · Mystery Game Puzzles 14.1%
- Open releaseInference.net releases Schematron V2 Turbo
Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in re
- Open releaseInference.net releases Schematron V2 Small
Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. Extraction instructions must be supplied through a JSON schema
- Model launchSakana releases Fugu Max
Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
- Model launchSakana releases Fugu Ultra v2
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...
- BenchmarkEpoch AI evaluates Kimi K3
Furniture Assembly 34.2%
- BenchmarkEpoch AI evaluates Kimi K2.6
Furniture Assembly 21.7%
- Open releaseinclusionAI releases Ling 3.0 Flash VL
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
- Open releaseMajorDeepSeek releases DeepSeek V4.1 Flash (batch)
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
- Model launchMajorCognition releases SWE-2 (max)
- BenchmarkEpoch AI evaluates Claude Fable 5.1
MirrorCode 73.3% · Furniture Assembly 70.0%
- BenchmarkEpoch AI evaluates Qwen3.8 Max (0902) (xhigh)
Furniture Assembly 20.0%
- BenchmarkEpoch AI evaluates Gemini 3.7 Flash
Furniture Assembly 26.7%
- BenchmarkEpoch AI evaluates Gemini 3.6 Flash
Furniture Assembly 23.3%
- BenchmarkEpoch AI evaluates Gemini 3.8 Flash
Furniture Assembly 31.7%