Skip to contentSkip to stories

Updated

All AI news

Sep 4

Sep 4Fri
  1. Tencent · new models on Hugging FaceAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu
  1. Google Developers BlogAI score23

    Google's Gemini Enterprise DevEx sprint fixes governance setup friction for agents

    AIGoogle's Gemini Enterprise developer experience team tested agent governance workflows without internal shortcuts and fixed friction points across its agent governance products. Fixes included documentation stating that enabling the Identity-Aware Proxy API is a hard requirement, auto-allowing essential Google-managed platform APIs in the Agent Gateway, and adding Private Service Connect and Cloud DNS setup guidance for Semantic Governance. The team also published ready-made Logs Explorer queries for monitoring Agent Gateways and Content Security.

  2. xAI News (Grok)AI score44

    xAI's Haggle Bot finds over $100,000 in procurement savings across SaaS and supplies

    AIxAI built a Grok-powered procurement agent, Haggle Bot, that audits vendor spend, flags unused SaaS seats, and prepares renewal negotiations. The Bot has identified more than $100,000 in direct savings, including $14,220 from 43 unused seats of one SaaS product and $85,662 a year in unused SKUs from another. A person still approves any spending, contract terms, or messages sent to vendors.

  3. Midjourney UpdatesAI score52

    Midjourney's alpha adds v8.2 edit model with lightbox editor

    AIMidjourney's alpha site now runs the new v8.2 edit model, with an editor built into the lightbox. Users can edit images with plain-text instructions, attach up to 4 reference images, and view all session edits in one place. The update also adds an early Change Style feature, and the team says speed and error messaging have improved, while drag and drop and the prompt bar are still in progress.

  4. Engineering at MetaAI score34

    Meta's ZGateway Proxy Unifies ZippyDB Client Traffic to Cut Connection Sprawl

    AIMeta has introduced ZGateway, a stateless proxy tier that now carries about 40% of all ZippyDB traffic, projected to exceed 60%, and handles over 1 billion operations per second. The proxy collapses the many-to-many client-to-database connection mesh into two bounded hops, adding about 6% computational overhead in an average use case. It also enables admission control, load balancing, and cross-region resilience, which contain reconnection storms that previously caused host crashes.

  5. Google DeepMind · The KeywordAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

  6. Google DeepMind · YouTubeAI score72

    Google DeepMind's WeatherNext 3 offers hourly, 5km-resolution weather forecasts

    AIGoogle DeepMind introduced WeatherNext 3, a weather forecasting model that learns directly from satellite feeds and ground-level weather station data. It produces a fresh forecast every hour, compared with the six-hour refresh typical of traditional models, with native 5km resolution for temperature and humidity. It is available through Google Search, Gemini, Google Maps and more.

    Why it matters: The source shows a shift from six-hourly to hourly refresh and 5km local resolution, which matters for energy planning and local forecasting.

  7. BAAI · new models on Hugging FaceAI score25

    BAAI Releases Recon2Reason-Reasoning-4B, a Spatial Reasoning Vision-Language Model

    AIBAAI released Recon2Reason-Reasoning-4B, a 4,437,815,808-parameter vision-language model fine-tuned from Qwen3-VL-4B-Instruct for indoor spatial reasoning. The model handles metric distance, relative position, and object-relation questions from single or multiple images, and loads with the standard Qwen3VLForConditionalGeneration interface without trust_remote_code. The checkpoint is released under Apache-2.0 with BF16 Safetensors weights, and the retrieval-augmented scene-reconstruction extension ships separately.

  8. Prime Intellect BlogAI score59

    Prime Intellect rebuilds GLM-5.2 RL weight transfer on NIXL, cutting sync to 3.9 seconds

    AIPrime Intellect reports that rebuilding RL weight transfer for GLM-5.2 on NIXL and ModelExpress cut sync time from 86.1 seconds with NCCL to 3.9 seconds in its fastest setting. The method traces vLLM's loader to find each tensor's runtime layout, then reads only the needed source bytes over RDMA and replays the rest locally. Most remaining latency comes from vLLM's pause consensus, which the team reduced by syncing every wave instead of every 32.

Sep 2

Sep 2Wed
  1. ARC PrizeAI score77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    AIOpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  2. NVIDIA · new models on Hugging FaceAI score36

    NVIDIA Releases EgoHand-1.0 Model for Single-Image 3D Hand Pose Estimation

    AINVIDIA released EgoHand-1.0, a 883.5M-parameter DINOv3-based transformer that predicts SOMA hand pose, MHR shape coefficients, and camera translation from a single 256×256 hand crop. The model is evaluated on the HOT3D egocentric benchmark and is intended for research and demonstration rather than production use. Its outputs can supply hand trajectories for training robotic manipulation policies, and it runs on NVIDIA Ampere GPUs under Linux with PyTorch.

  3. NVIDIA · new models on Hugging FaceAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.

  4. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya En-Thinker, a 3.35B Multilingual Reasoning Model

    AICohere Labs released Tiny Aya En-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model with a 32K context length. It is trained on English reasoning traces for 44 languages plus English, with coverage extending to 20+ more languages through non-reasoning instruction data. The model is available under a CC-BY-NC license that also requires adherence to Cohere Labs' Acceptable Use Policy.

  5. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya L2-Thinker Multilingual Reasoning Model on Hugging Face

    AICohere Labs released Tiny Aya L2-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model that thinks in the same language as the user's prompt before answering. The model supports in-language reasoning for 44 languages plus English, with coverage extended to 20+ more languages through additional non-reasoning instruction data, and has a 32K context length. It is licensed under CC-BY-NC and is available on Hugging Face.

  6. Engineering at MetaAI score55

    Meta details an AI agent that learns from expert corrections without retraining

    AIMeta Engineering describes an AI agent for a compliance domain that stores expert knowledge in structured, auditable files and separates it from reasoning procedures called recipes. Expert feedback is diagnosed, compiled into verified text edits, tested against regression suites, and reviewed by humans, all without retraining the underlying model. Meta reports that domain experts rated outputs useful almost all the time and that assessment time fell from days to minutes.

Sep 1

Sep 1Tue
  1. Google Developers BlogAI score39

    Four engineering patterns behind top Google AI Agents Challenge submissions

    AIGoogle's AI Agents Challenge judges highlighted four engineering patterns in top-ranked submissions: bidirectional MCP, event-driven concurrency, same-bar fallback, and tiered routing. One team exposed its internal MCP tools as an external MCP server that other agents could call, with access control required once outside callers reach it. Another replaced a linear agent pipeline with an asyncio.Queue-based event bus so agents react to shared events in parallel rather than waiting in a call chain.

  2. Cursor ChangelogAI score62

    Cursor adds self-hosted machines that keep tool execution inside your network

    AICursor now supports self-hosted machines, so tool execution stays on your own infrastructure while the agent makes tool calls locally. Team pools are named worker queues that scale with requests and can hibernate idle machines, restoring them within a reconnect window. Cloud agents can also run on sandboxes such as AWS Lambda, Cloudflare, Modal, and Vercel, and self-hosted workers now support computer use on Linux and Mac.

    Why it matters: The update explains how self-hosted workers keep tool execution inside your network while pools scale and hibernate, which matters for teams with strict data controls.

  3. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  4. Anthropic · YouTubeAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

  5. Ai2 · new models on Hugging FaceAI score22

    Ai2 Releases Supplemental ACE2S-SHiELD+ Ablation Checkpoints on Hugging Face

    AIAi2 has published supplemental checkpoints for its ACE2S-SHiELD+ climate model on Hugging Face, covering four ablation configurations that test random CO2 data and energy conservation. Each configuration includes two random-seed models, and the repository recommends the main ACE2S-SHiELD+ checkpoint for most uses. The checkpoints are licensed under Apache 2.0 for research and educational use.

  6. Microsoft AI BlogAI score34

    Microsoft Publishes 2026 Responsible AI Transparency Report on Governance and Agentic AI Risks

    AIMicrosoft published its 2026 Responsible AI Transparency Report, its third annual edition, detailing updates to its governance and risk management. The company re-engineered its Responsible AI Standard to adapt to evolving technical risks and regulatory requirements, and is extending controls such as agent identities, tool permissions, and action monitoring to agentic AI systems.

  7. Google · new models on Hugging FaceAI score44

    Google Releases GNM v3.0, an Open 3D Parametric Model of the Human Head

    AIGoogle has released GNM v3.0, a parametric 3D statistical model of the human head, with weights published on Hugging Face and Kaggle under the Apache 2.0 license. The model gives controllable identity, expression, head pose, and internal anatomy including eyeballs, teeth, and tongue, and supports NumPy, JAX, PyTorch, and TensorFlow backends.

  8. Ai2 (Allen Institute for AI)AI score38

    Ai2 Panel Identifies Five Hard Challenges for AI-Assisted Science

    AIAt an August 27 Ai2 event on expanding its work with the Paul G. Allen Research Center at Providence Swedish Cancer Institute, panelists identified five persistent challenges for scientific AI. The main ones are keeping AI steerable as research evolves, deciding which tasks to delegate, and avoiding the amplification of weak study design or bad data.

  9. Ai2 (Allen Institute for AI)AI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

  10. OpenBMB (MiniCPM) · new models on Hugging FaceAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

  11. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

  12. Gemini API ChangelogAI score62

    Gemini API adds agentic video understanding for three Gemini models

    AIGoogle released agentic video understanding for Gemini 3.7 Flash, Gemini 3.6 Flash, and Gemini 3.5 Flash-Lite across the Interactions and GenerateContent APIs. The model dynamically navigates video timelines, requesting transcripts, frames, or audio tracks on demand. The source says this approach uses up to 88% fewer tokens for long-form content than static processing.

    Why it matters: The changelog names the affected models and API surfaces, and states a token-use figure that helps developers judge the cost of long video workloads.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

  2. Liquid AI NewsletterAI score46

    Liquid AI launches Pipette, an open-source benchmark for on-device foundation models

    AILiquid AI and Artificial Analysis released Pipette, an open-source benchmark platform for foundation models on edge devices, covering over 1,000 configurations across 30+ models. It measures five on-device metrics, including throughput, latency, context scaling, and memory use, on macOS, Windows, iOS, and Android. Liquid AI also said its updated LFM2.5 Q4_0 checkpoints, trained with Quantization-Aware Distillation, retain roughly 97% of BF16 baseline performance and suffer 73.4% less quality loss than standard post-training Q4_0 quantization.

  3. Microsoft ResearchAI score45

    GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Population-Scale Research

    AIMicrosoft Research released GigaPath-Flash and GigaTIME-Flash, efficient pathology foundation models built on a distilled ViT-S backbone and released under the Apache 2.0 license. GigaPath-Flash, with 22M-parameter tile and 21M-parameter slide encoders, reportedly scores within 3% of the original GigaPath on PANDA and EBRAINS benchmarks at roughly 50 times less compute. The models are research tools, not validated for clinical use.

  4. Amazon ScienceAI score45

    Amazon details using Verus to formally verify Rust code correctness

    AIAmazon Science explains Verus, an open-source automated program verifier for Rust that checks code against formal specifications for all possible inputs. Developers write specifications and proofs directly in Rust source using Rust-like syntax, and Verus returns feedback in under a second. Amazon says it has used Verus to prove the correctness of key primitives in the Nitro Isolation Engine and other infrastructure.

  5. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

  6. METR BlogAI score38

    METR Reports Two Security Incidents, Including Stolen API Key Used for Public Model Credits

    AIMETR disclosed two 2026 security incidents in which external attackers attempted unauthorized access, with no evidence of AI agents hacking third parties during its evaluations. In March, attackers stole an API key from a researcher's personal instance and consumed credits on public models that were worth about $600,000 but were granted to METR for free. METR says it found no evidence that sensitive information was accessed in either incident.