OpenAI announced that GPT-6 is being rolled out to all ChatGPT users, alongside a new Intelligent UI. The excerpt provides no further details on capabilities, availability timing, or how the interface works.
Anthropic released Claude Haiku 5.5, a small model priced at about a quarter of Haiku 4.5's running cost. Anthropic also halved Sonnet 5.5 cache read prices from 20 cents to 10 cents per million tokens.
Xiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.
AIWhy it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.
ICYMI we've cut the price of Sonnet 5.5 cache reads in half. In the Claude Platform, it's now $0.10 per million tokens (input is $2, output is $10). This means Sonnet 5.5 now runs ~20% cheaper on most agentic work. (API only, no change to Claude Code usage limits.)
Artificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now ranks first. The model is available only through OpenAI's Daybreak program and records no safety blocks across the Index. Its overall score is 32 points higher than the publicly available GPT-6 Sol (max), at a cost of $1.77 per task versus $11.67 for Grok 4.7 (xhigh).
Odyssey-3 is a foundation world model, enabling many applications in physical AI, human experiences, and even how we train intelligences. We're particularly excited by agents learning from experience inside Odyssey-3, working to accomplish objectives.
JetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.
Anthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.
Across StepFun's 8-benchmark eval, Step 5 Preview beats Kimi K3 and GLM-5.3 on 6, including DeepSWE (67.7), ProgramBench (80.5), StepCodeBench (49.0), and FrontierFinance (66.4).
Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview
StepFun's Step 5 Preview is now live on OpenRouter, with a week of free access rolling out across opencode, Cline, NousResearch, and Kilo Code. The company positions it as flagship-tier intelligence for agentic and professional work at substantially lower task cost, letting users keep their workflow while switching models.
StepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.
JetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.
AIWhy it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.
Anthropic released Claude Haiku 5.5, its fastest and most affordable small model, at prices up to 90 percent lower for most prompts under 100,000 tokens. Artificial Analysis ranks it first among small-class models on its Intelligence Index with a score of 43, but it consumes about three times the output tokens per task that GPT-6 Luna needs at maximum effort.
Anthropic released Claude Haiku 5.5, priced the same as OpenAI's GPT-6 Luna, with a 1M-token context window. Artificial Analysis scored it 43 on its Intelligence Index, slightly ahead of GPT-6 Luna at 38, but it uses about 3x more output tokens at max effort.
Anthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.
Anthropic has released Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. It keeps a 1M token context window, up to 128K output tokens, and is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic reports 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, and the article notes that non-default temperature, top_p or top_k values return a 400 error.
For reference, Haiku 4.5 came out on October 15th, 2025 so these two columns are less than a year apart. and oh yeah Haiku 5.5 is also much faster and 75% cheaper...
Anthropic's Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest and most efficient model in the Claude 5.5 family and costs around 75 percent less than Claude Haiku 4.5 for most tasks. The post also covers pairing it with Claude Opus 5.5 as a subagent layer and provides Boto3, Converse, and Anthropic SDK examples for calling the model.
Anthropic has released Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released. On average it costs around 75% less to run than Claude Haiku 4.5, and the author says it is 10x cheaper than Haiku 4.5 under 100k tokens. It can be tried with computer use, workflows, and the API.
Anthropic has released Claude Haiku 5.5, which the author describes as its fastest and cheapest model to date. The source says it costs about 75% less to run than Claude Haiku 4.5 and is the first Haiku model with an adjustable effort setting. The attached benchmark table reports Haiku 5.5 scores on tasks including computer use (OSWorld 2.1 offline subset, 72.4%) and Terminal-Bench 4.0 (39.2%), compared with Haiku 4.5 and other models.
Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles repetitive work like summaries and classification, and pairs well with Claude Opus 5.5 and Sonnet 5.5 as a subagent on coding work. It’s also fast enough for live customer support and browser use.
d1-3B ranks first among models under 10B on the Decision Index v0.2.1, a benchmark for structured decision-making. Built from LFM2.5-VL-3B, it makes decisions from text and images in one pass. Use it for reranking, agent guardrails, and visual inspection. 2/
Mistral AI's Mistral Large 4 is listed on OpenRouter as a frontier multimodal model accepting text and image input. The listing says it is built for reasoning, coding, and agentic workloads and offers a 1M-token context window. The feed excerpt is truncated, so further details such as pricing or availability are not confirmed here.
On important coding and agentic benchmarks such as DeepSWE, AutomationBench, and AA-Briefcase, ML4 matches the performance of the best open-weight models. It is SOTA on finance and legal workflows, as well as on complex multimodal grounding benchmarks. It can navigate complex terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks. 4/n
Reflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.
Reflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.
The author says Beam, a 501B-parameter open model from Reflection AI, comes close to GLM 5.2 in capability but trails GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash in several areas. The author attributes Beam's competitiveness mainly to inference efficiency, with inference compute at roughly one-third to one-quarter of GLM 5.2's.
Sophia Yang congratulated Reflection AI on Beam, a 501B-parameter open model with 23B active per token. She attributes its efficiency to an RL length penalty that discourages unnecessary tokens and a sparse MoE architecture. Reflection says full weights will be released this month, and the quoted post reports training over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.
Reflection AI introduced Beam, an agentic open model with 501B total parameters and 23B active parameters, trained end-to-end from scratch. The quoted announcement says it targets frontier reasoning efficiency and coding and agentic tasks, with full weights due this month. Clément Delangue, Hugging Face's CEO, reposted it with a welcome to the Reflection organization on Hugging Face.
AIWhy it matters: The quoted announcement names Beam's parameter scale, active-parameter count, and coding and agentic focus, which helps readers gauge where it fits among open models.
Congrats to Reflection folks -- it's hard to get the first model out, and we hope you can rapidly accelerate contributions to the ecosystem from here :)
Ai2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.
AIWhy it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.
Pareto 26.10 Preview is a multimodal composite model built for research, coding, and agentic workflows. It is described as delivering frontier-level performance across a broad range of general-purpose tasks, though the source excerpt is a preview and provides no benchmark scores, parameter counts, pricing, or availability details.
Google has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.
AIWhy it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.