Skip to contentSkip to stories

Updated

All AI news

Sep 10

Sep 10Thu
  1. Microsoft Foundry BlogAI score41

    Azure AI Speech LLM 2607 adds multilingual accuracy gains and phrase list customization

    AIMicrosoft released Azure AI Speech LLM 2607, which improves multilingual recognition, mixed-language audio handling, and domain-specific entity accuracy, and runs up to 3x faster than the previous 2605 release. A new dedicated phrase list parameter lets developers supply domain vocabulary, supporting 2,000+ entities, without embedding it in a prompt. The model is available through the Fast API and Real-Time API, is testable in the Foundry Playground, and is deployed automatically with no customer action required.

  2. Google LabsAI score38

    Google Labs' Dreambeans personalized daily story app now available to all U.S. accounts

    AIGoogle Labs has made Dreambeans, its experimental app that creates personalized daily story collections, available to all U.S. accounts aged 18 and over on Android and iOS. Each daily collection combines personalized topics with information distilled from connected Google apps, including Calendar, Gmail, Photos, Search, YouTube, and Gemini. Users can dive deeper into stories, bookmark them, share them, and give feedback to improve future collections.

  3. Cognition Blog (Devin, Windsurf)AI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

  4. Cognition Blog (Devin, Windsurf)AI score22

    Cognition Welcomes Dioxus Team to Advance Open-Source Cross-Platform App Framework

    AICognition has welcomed Jonathan Kelley and the Dioxus team, whose framework Cognition used extensively to build and improve Devin's performance. Cognition plans to continue supporting Dioxus, Blitz, Taffy, and Subsecond while increasing investment in Dioxus-Native and Blitz. The Dioxus team will also work on Devin's virtual machine, computer use skills, and testing capabilities.

  5. Amazon ScienceAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

  6. Replit BlogAI score36

    Replit and Databricks Integration Becomes Generally Available with Lakebase Support

    AIReplit's integration with Databricks is now generally available, adding native Databricks Lakebase support that lets Replit Agent automatically provision a Lakebase database when an app is ready to deploy. Apps built with Replit can read live Databricks warehouse data while storing new app data in Lakebase, inheriting existing Unity Catalog security and governance controls. The update also adds automated preview deploys that keep test data isolated from live business data.

  7. Mistral AIAI score36

    Cloudera and Mistral Partner to Deliver Sovereign AI on Enterprise Data

    AIMistral AI and Cloudera announced a partnership that integrates Mistral's models with Cloudera's hybrid data platform for enterprise AI. Customers can run inference across private and public cloud, on-prem, and fully air-gapped environments, and train custom models on proprietary data while retaining ownership. Cloudera cited 30 exabytes of customer-managed data on its platform.

  8. DeepSeek API NewsAI score72

    DeepSeek releases V4.1-Flash with native multimodal support and API updates

    AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.

    Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.

Sep 9

Sep 9Wed
  1. BAAI · new models on Hugging FaceAI score24

    BAAI open-sources EPT, UniPath, and MiSI AIDD molecular and crystal modeling resources

    AIBAAI released open-source resources for three AIDD projects on Hugging Face: EPT, an equivariant pretrained transformer for unified 3D molecular representation learning, and UniPath, a learnable-time flow matching method for crystal structure and energy prediction. The repository mirrors their GitHub source code and READMEs, with setup, preprocessing, training, and evaluation documentation. The MiSI benchmark is released separately on Hugging Face.

  2. DeepSeek · new models on Hugging FaceAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

  3. Cursor ChangelogAI score73

    Cursor launches Projects for long-running, multi-agent coding work

    AICursor is launching Projects, a beta feature for larger work such as a feature, migration, or full app, rolling out to all users starting today. A coordinator agent plans the work, delegates it to implementing agents that can run in parallel, and runs on a cloud computer so it continues when the laptop is closed. Each Project keeps shared context files synced across cloud and local machines, and subscriptions let the coordinator act on Slack channels, schedules, or PRs without a prompt.

    Why it matters: The source details how a coordinator agent plans, delegates, and syncs shared context across cloud and local machines, useful for judging how long-running agent work might fit a team's workflow.

  4. Together AI BlogAI score42

    Together AI Launches Preemptible GPU Compute at 50% of On-Demand Price

    AITogether AI has launched a public preview of preemptible compute for Together GPU Clusters on Kubernetes in all regions, billed sub-hourly at a flat 50% of the on-demand rate. Preemptible nodes can be reclaimed when capacity is needed elsewhere, with a five-minute drain window for checkpointing before removal. The rate stays fixed rather than tracking a spot market, and the cluster automatically refills its preemptible target as capacity becomes available.

  5. Fireworks AI BlogAI score58

    Fireworks AI outlines a staged path from closed APIs to owned specialized models

    AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.

  6. Fireworks AI BlogAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.

  7. Microsoft Foundry BlogAI score62

    Microsoft Foundry's July and August 2026 updates bring Hosted Agents and Toolboxes to GA

    AIMicrosoft Foundry's July and August 2026 updates make Hosted Agents, Voice Live integration, and Toolboxes generally available. The post adds Claude tools on Azure, Model Router region and model pool changes, Foundry Local preview features, and updated Python, JavaScript, Java, and .NET SDK versions with migration notes.

    Why it matters: The roundup links each GA and preview change to code examples, migration notes, and runtime requirements, which helps developers judge what to upgrade and test first.

  8. Cognition Blog (Devin, Windsurf)AI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    AICognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    Why it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

  9. Google DeepMind · YouTubeAI score38

    How AI is transforming weather prediction, featuring WeatherNext 3

    AIGoogle DeepMind's Peter Battaglia discusses how machine learning is changing global weather forecasting, including early warnings for storms such as Hurricane Melissa. The episode covers traditional physics-based models versus AI models and probabilistic forecasting, and highlights WeatherNext 3 as Google DeepMind's most advanced global weather AI model yet.

  10. Mistral AIAI score54

    Mistral details how AI agents migrated 40,000 lines of Fortran to C++

    AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.

  11. Ai2 (Allen Institute for AI)AI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Google Developers BlogAI score72

    Google releases ADK for Kotlin 1.0 for building production AI agents

    AIGoogle announced general availability of ADK for Kotlin 1.0, a Kotlin Multiplatform framework for building AI agents on servers and Android. Version 1.0 reaches feature parity with ADK 1.0 Core and adds Android extensions for on-device models, cloud Gemini via Firebase AI Logic, and persistent sessions and memory with Room and AppSearch. The post includes a server-side incident triage example using KSP-generated tools and skills, plus an Android financial assistant example with human confirmation for transfers.

    Why it matters: The post names the new Android and server-side capabilities and the code setup, helping Kotlin developers judge whether ADK fits their agent projects.

  2. Factory NewsAI score34

    Factory Now on Claude Marketplace for Enterprise Autonomous Software Development

    AIFactory is now available on the Claude Marketplace, letting enterprise customers apply their committed Anthropic spend toward its autonomous software development platform. The platform automates the software development lifecycle, covering planning, implementation, testing, and security within one system, with enterprise deployment options that keep execution close to customers' code and infrastructure.

  3. Google Developers BlogAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

  4. Cohere · new models on Hugging FaceAI score38

    Cohere releases Tiny Aya Base 32K, a 3.35B multilingual model with 32K context

    AICohere Labs has released Tiny Aya Base 32K, an open-weights pretrained model with 3.35 billion parameters and a 32K context window. The model covers 70+ languages, including many lower-resourced ones, and is designed for downstream adaptation and long-context research. It is a base model that has not been instruction-tuned, and it is licensed under CC-BY-NC.

  5. Google DeepMind · YouTubeAI score60

    Google DeepMind launches AlphaGenome Atlas for mapping genetic variant effects

    AIGoogle DeepMind introduced AlphaGenome Atlas, an AI-powered database charting the molecular impact of every possible genetic variant. Scientists are already using it to investigate unsolved rare diseases and map rare mutations linked to complex traits.

    Why it matters: The source names a concrete use case, finding disease-causing DNA variants, which shows how the database could support rare disease research.

  6. Google DeepMindAI score74

    Google DeepMind launches AlphaGenome Atlas to predict 9 billion DNA variant effects

    AIGoogle DeepMind has introduced AlphaGenome Atlas, a platform with predicted molecular effects for 9 billion single-nucleotide variants in the human genome. It is free for academic research through a web portal, and the AlphaGenome Variant Impact score condenses predictions from AlphaGenome and AlphaMissense into one number for ranking variants. The source says collaborators used it to identify variants in unsolved rare disease cases and to find rare non-coding variants linked to traits.

    Why it matters: The source details how precomputed variant predictions, a single impact score, and linked feature attributions make genome-wide mutation effects searchable for researchers without coding skills.

  7. Google DeepMind · The KeywordAI score72

    Google DeepMind launches AlphaGenome Atlas, a database of DNA variant effect predictions

    AIGoogle DeepMind has released AlphaGenome Atlas, a web portal that predicts the regulatory effects of all 9 billion possible single-letter genetic changes in the human genome. The Atlas provides an AlphaGenome Variant Impact (AVI) score that combines coding and non-coding predictions to help researchers prioritize variants. The source says the portal requires no coding skills and is available to researchers and biologists worldwide.

    Why it matters: The source details how the Atlas's AVI score is used in real rare disease and UK Biobank analyses, showing a practical route for prioritizing non-coding variants.

  8. Google DeepMind · YouTubeAI score78

    DeepMind releases AlphaGenome Atlas, a predictive map of every possible DNA letter change

    AIGoogle DeepMind has used AlphaGenome to predict the molecular impact of every possible single-letter change in the human genome, around nine billion variants. The resulting AlphaGenome Atlas is a 1PB dataset that assigns each variant an AlphaGenome Variant Impact (AVI) score, covering both coding and non-coding variations, and is available to researchers worldwide. The video notes that AlphaGenome has not been validated or approved for any clinical use.

    Why it matters: The release supplies a precomputed impact score for every possible single-letter genome change, which lets researchers look up variants without running the model themselves.

  9. Mistral AIAI score62

    Mistral raises €3B Series D at over €21B valuation led by Samsung

    AIMistral announced a €3 billion Series D round at a post-money valuation of more than €21 billion, led by Samsung Electronics with co-leads Scaleup Europe Fund and PSG Equity. The company says the funding will expand frontier research, compute capacity, infrastructure, and international growth, and that it now operates in 20 countries with 125+ enterprise customers including Airbus, ASML, and HSBC.

    Why it matters: The round shows how a company frames sovereign, open-weight AI as a full stack spanning models, infrastructure, compute, and products, which is useful context for European enterprise AI strategy.

  10. NVIDIA · new models on Hugging FaceAI score46

    NVIDIA Releases NV-Reason-CT, a 3D Vision-Language Model for Chest and Abdominal CT

    AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.

  11. Suno BlogAI score46

    Suno Partners with Believe and TuneCore to Distribute Independent Artists' Music

    AISuno announced a global strategic partnership with Believe and its self-release platform TuneCore to give independent artists new ways to create, distribute and earn from music. Tracks made with Suno's new industry partner model will become eligible for distribution through Believe and TuneCore, and the companies plan new product experiences that let artists collaborate with fans and get paid. Suno said its audio watermarking, fingerprinting and download limits will also apply to music distributed through Believe and TuneCore.

Sep 7

Sep 7Mon
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.

Sep 6

Sep 6Sun
  1. Google DeepMind · The KeywordAI score24

    Google DeepMind backs 16 Asia-Pacific green AI projects in inaugural accelerator

    AIGoogle DeepMind selected 16 organizations for its inaugural Accelerator: AI for the Planet (APAC) cohort to scale AI-powered environmental solutions. Participants span biodiversity monitoring, sustainable agriculture, and climate and carbon projects across countries including New Zealand, Singapore, Indonesia, India, and Japan. Over three months, they will receive access to Google's AI stack, including frontier models, plus mentorship from company experts.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceAI score35

    UltraData-Code-L2-Classifier scores files for algorithmic code selection

    AIOpenBMB released UltraData-Code-L2-Classifier, a suite of language-specific file-level scorers for 11 programming languages in UltraData-Code-L1. The L2 corpus selected with these scorers contains approximately 400B tokens and retains about 12.23% of L1 files, and a 10B-token test on a 1B model raised EvalPlus pass@1 by 7.80 points over L1 training.

  3. OpenBMB (MiniCPM) · new models on Hugging FaceAI score32

    MiniCPM5-2B-DSpark draft model released for speculative decoding with MiniCPM5-2B

    AIOpenBMB released MiniCPM5-2B-DSpark, a 323,776,001-parameter DSpark draft checkpoint with five layers that proposes seven draft tokens per forward pass for the MiniCPM5-2B target model. The model, trained on 7,054,154,509 tokens with an average acceptance length of 5.5174 at T=0 and 4.0514 at T=1.0, is served through SGLang with DSPARK speculative decoding. It is released in BF16 under the Apache-2.0 License.

  4. OpenBMB (MiniCPM) · new models on Hugging FaceAI score62

    OpenBMB releases MiniCPM5-2B, a 2B open-source model with open training data

    AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.

    Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.

Sep 4

Sep 4Fri
  1. BAAI · new models on Hugging FaceAI score26

    ConsiSpace: BAAI and Peking University release geometry-consistent video spatial reasoning model

    AIBAAI and Peking University researchers released official weights for ConsiSpace, a geometry-consistent multimodal framework for spatial reasoning in long-form visual observations. The model is described in the paper "ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning" (arXiv:2607.17599).

  2. Tencent · new models on Hugging FaceAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.