Microsoft releases MAI-Transcribe-2 speech transcription model
AIMustafa Suleyman points readers to a Microsoft AI page about the MAI-Transcribe-2 model. The post itself gives no details on capabilities, benchmarks, or availability beyond the link.
Updated
Updated
AIMustafa Suleyman points readers to a Microsoft AI page about the MAI-Transcribe-2 model. The post itself gives no details on capabilities, benchmarks, or availability beyond the link.
AIMicrosoft's Mustafa Suleyman shared a link to try MAI-Transcribe-2 on OpenRouter. The post provides no further details on the model's capabilities, specifications, or pricing.
AIBAAI introduced AREX, a research agent built on a 122B-parameter mixture-of-experts model with 10B active parameters. It drafts candidate answers, checks each constraint, and revisits unresolved points rather than running one long search. The post says AREX performs on hard search benchmarks comparable to GPT-5.4.
AILM Studio has launched its Bionic version, now live for users. The announcement is tied to DeepSeek-V4.1-Flash, the smallest model in DeepSeek's new architecture family, which adds native visual understanding and targets faster inference and higher throughput.
AIJunyang Lin asks on X whether v4p is back and whether k3-0.2 is coming. The post gives no details on either model, so no further specifics can be confirmed.
AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.
Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.
AIOllama says DeepSeek-V4.1-Flash is now fully rolled out on its cloud, hosted in the US and Europe. Prompts and responses are not logged or trained on, and per-token pricing matches the DeepSeek API, including off-peak pricing. The post repeats DeepSeek's claim that the model is more capable, faster, and more cost effective than prior DeepSeek models, including DeepSeek-V4-Pro.
AISakana AI released Fugu Max and Fugu Ultra v2, multi-agent orchestration systems that route tasks across a pool of open-weights and specialized models. The source says Fugu Max delivers performance within striking distance of elite models at two to six times lower cost, while Fugu Ultra v2 outperforms Opus 5 and Fable 5 on Chartography and outperforms models costing three to five times more per token on DeepSWE.
AILaravel announced last week that it no longer accepts GitHub issues and requires contributors to submit pull requests instead. The article argues that PR-only policies filter out spam issues and help maintainers, since users can now generate PRs with AI.
AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.
AIModel page:
AIOllama is rolling out DeepSeek-V4.1-Flash on its cloud, starting with Max and Team accounts. The company says it is quickly adding capacity to extend access to all subscribers.
AIOpenAI announced that a swarm of 10,000 agents produced a solution to the Navier-Stokes Millennium Problem, a result that angered mathematicians. NYU mathematician Tristan Buckmaster and Anthropic-employed collaborator Levent Alpöge had been working on related problems and released three draft papers of about 245 pages. Buckmaster said OpenAI's offer to merge efforts required acknowledging an OpenAI model and excluded Alpöge as co-author.
Why it matters: The piece separates the mathematical result from the collaboration dispute, showing how AI labs' compute spending is straining academic norms around credit and openness.
AIOpenAI's GPT-6 Astra is reported as the best-performing frontier model for antibody developability prediction in one benchmark, outperforming other frontier models tested on properties such as aggregation and stability. The post also says Astra built an interactive antibody visualization in about one hour.
AINVIDIA's Bryan Catanzaro welcomed the idea of models that understand how humans interact and thanked builders using Nemotron Ultra. Humansand introduced Persimmon, which it describes as the first large-scale model designed to realistically simulate how people talk and interact.
AISebastian Raschka says DeepSeek V4.1 contains a major architecture overhaul using an encoder-decoder setup, and he argues it could have been named V5. The attached diagrams compare DeepSeek V4-Flash (284B) with DeepSeek V4.1-Flash (552B), which has 1M supported context and a 10-layer encoder. The attached charts report a global KV cache per token of 890 bytes for V4.1-Flash, versus 3,514 for V4-Flash and 48,068 for DeepSeek-V3.2.
AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.
Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.
AIMustafa Suleyman says Microsoft launched eight new models in two months, with five debuting at #1 on their respective leaderboards. MAI-Transcribe-2 is described as the fastest, most accurate, and cheapest transcription model, while MAI-Image 2.5, 2.6, and 2.6 Flash debuted at #1 on the Artificial Analysis leaderboard. MAI-Code-1.1-Flash, launched in GitHub Copilot four weeks ago, now accounts for a third of small-model traffic there.
AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.
AIDeepSeek says its more efficient V4.1-Flash architecture lets it serve more users at lower cost and pass the savings on through lower API prices. Peak/off-peak pricing continues, with off-peak rates at 50% of peak rates. The new pricing takes effect at 04:00 UTC on September 10, 2026.
AIDeepSeek says V4.1-Flash is now live on its API with native multimodal support, accessed through the model name deepseek-flash. The older V4-Flash and V4-Flash-Vision-Exp are retired, while deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting 04:00 UTC on Sept 14, 2026, until V4.1-Pro launches.
AIDeepSeek says its V4.1-Flash model needs only 1/4 the HBM and 1/8 the SSD storage for its KV cache compared with the previous generation. Because cache-hit charges often make up a large share of agent costs, the company says the compressed cache significantly reduces those costs.
AIDeepSeek has introduced a 552B-parameter MoE model built on a new Causal Encoder–Decoder architecture, activating 8B parameters for input and 16B for output. The company says new pre-training methods and larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.
Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.
AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.
Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.
AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.
AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.
Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.
AIMeta announced Muse, a personal AI agent designed to get things done for users across many parts of life. The product is powered by Muse Spark 1.3, and the post links to an app download and a page describing how Muse was built.
Why it matters: The announcement names Muse Spark 1.3 as the underlying model, giving readers a concrete product and model pairing to track.
AIMark Chen announced that a group of agents produced a solution to the Navier-Stokes Millennium Prize Problem, using an unnamed OpenAI next-generation model. The post says the problem concerns whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and that it had been open for roughly 90 years. The quoted OpenAI post and the attached illustration of inward spiral and axial stretching are cited as context, but the source provides no proof details.
Why it matters: The post claims an AI-produced proof of a famous open problem, but the source gives no proof details or independent verification, so the claim itself is the main point.
AIGoogle AI Studio is encouraging developers to start building apps with Gemini 3.8 Flash, linking to a guide on ai.dev. The post offers no benchmarks, pricing, or feature details beyond that invitation.
AICohere Labs has released Tiny Aya Base 32K, an open-weights pretrained model with 3.35 billion parameters and a 32K context window. The model covers 70+ languages, including many lower-resourced ones, and is designed for downstream adaptation and long-context research. It is a base model that has not been instruction-tuned, and it is licensed under CC-BY-NC.
AIOpen model licenses are tightening at the Chinese frontier, with Zhipu's GLM-5.3 switching from MIT to a custom license requiring a security review for inference and fine-tuning providers with over $10 billion in annual revenue. Motif-3 ships under an MIT license with strong scores for its size, while Tencent's Hy4-preview is a competent model that currently overthinks. Western makers Google and Meta have moved to Apache 2.0.
AIUnsloth's Qwen3.8-27B GGUF has become the most-liked GGUF model of all time on Hugging Face, reaching 10 million downloads and 3.7K likes in 24 days. The post links the GGUF model page and a Qwen3.8 guide on the Unsloth docs site.
AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.
AITencent Hunyuan says its Hy4 preview has been upgraded to reduce long thinking and over-verification on complex tasks, which users had flagged. The company reports the same task quality with fewer turns and lower input and output tokens, confirmed by benchmark and human evaluation. The upgrade is live for all users, and Tencent says it will keep iterating based on feedback.
AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.
AIOpenBMB released MiniCPM5-2B-DSpark, a 323,776,001-parameter DSpark draft checkpoint with five layers that proposes seven draft tokens per forward pass for the MiniCPM5-2B target model. The model, trained on 7,054,154,509 tokens with an average acceptance length of 5.5174 at T=0 and 4.0514 at T=1.0, is served through SGLang with DSPARK speculative decoding. It is released in BF16 under the Apache-2.0 License.
AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.
Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.
AIv0 has made GPT-6 Astra available on its platform, according to the announcement. Users can try it at v0.app.