Skip to contentSkip to stories

Updated

#Model release

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 4

Sep 4Fri
  1. Mustafa SuleymanXAI score25

    Microsoft makes MAI image model 2.6 Flash available in Foundry

    AIMicrosoft's MAI-Image-2.6-Flash is now available in Microsoft Foundry and the MAI Playground. The post links to a Foundry model card for details, but gives no further specifications, pricing, or benchmarks.

  2. Mustafa SuleymanXAI score21

    MAI-Image-2.6-Flash generates images twice as fast as GPT-Image-2

    AIMicrosoft's MAI-Image-2.6-Flash generates images 2x faster than GPT-Image-2, which the post calls the best model in the world. It also uses 72% less GPU, enabling a lower price and what the post describes as the best price-performance score.

    Image from @mustafasuleyman's post
  3. Google AI StudioOfficialAI score44

    Google's Lyria 3.5 music model now available in AI Studio and Gemini

    AIGoogle has made Lyria 3.5, its best-sounding music generation model, available in AI Studio, through the Gemini API, and in the Gemini app. The model produces more expressive vocals and richer musical arrangements, enabling higher-fidelity tracks.

    Video from @GoogleAIStudio's post
  4. Daniel HanXAI score49

    Unsloth Desktop speeds up GLM-5.3-Flash GGUF local inference with MTP

    AIUnsloth Desktop now runs GLM-5.3-Flash GGUFs out of the box with faster inference, enabling MTP and faster long-context decoding. The quoted Unsloth post reports local GGUF inference 1.6–3.4× faster with optimized decoding and multi-token prediction, and 3-bit runs on 128GB setups.

  5. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  6. Tencent · new models on Hugging FaceOfficialAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu
  1. Mckay WrigleyXAI score26

    GPT-6 Astra release called a categorical step change in AI

    AIMckay Wrigley says the GPT-6 Astra release feels like the GPT-4 launch and marks a genuine step change where AI can do categorically new things. He suggests the model weights hold many capabilities that users will uncover through prompting.

  2. Midjourney UpdatesOfficialAI score52

    Midjourney's alpha adds v8.2 edit model with lightbox editor

    AIMidjourney's alpha site now runs the new v8.2 edit model, with an editor built into the lightbox. Users can edit images with plain-text instructions, attach up to 4 reference images, and view all session edits in one place. The update also adds an early Change Style feature, and the team says speed and error messaging have improved, while drag and drop and the prompt bar are still in progress.

  3. Mark ChenXAI score80

    Mark Chen announces GPT-6 Astra with computer use and agent oversight

    AIOpenAI researcher Mark Chen announced GPT-6 Astra, which he described as the company's most capable and aligned model yet. He said it can build and test software, work across apps on a computer, and help with open scientific problems. The post also highlights improved computer use compared with Operator and stronger monitoring that can stop potentially unauthorized agent actions.

    Why it matters: The post links a named model release to specific capabilities like computer use and aligned agent behavior, giving readers concrete claims to check against the model.

  4. Stephanie PalazzoloXAI score72

    OpenAI releases GPT-6 Astra; Greg Brockman suggests it could be AGI

    AIOpenAI has released GPT-6 Astra, according to a post by Stephanie Palazzolo. In a press briefing, president Greg Brockman suggested the model could be AGI. Executives also addressed a recent Information report about a technique that could make future models harder to monitor.

    Why it matters: The post links a model launch to executive comments on AGI and to a report about a technique that could make future models harder to monitor.

  5. Google DeepMind · The KeywordOfficialAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

  6. Google DeepMind · YouTubeOfficialAI score72

    Google DeepMind's WeatherNext 3 offers hourly, 5km-resolution weather forecasts

    AIGoogle DeepMind introduced WeatherNext 3, a weather forecasting model that learns directly from satellite feeds and ground-level weather station data. It produces a fresh forecast every hour, compared with the six-hour refresh typical of traditional models, with native 5km resolution for temperature and humidity. It is available through Google Search, Gemini, Google Maps and more.

    Why it matters: The source shows a shift from six-hourly to hourly refresh and 5km local resolution, which matters for energy planning and local forecasting.

  7. Microsoft AIOfficialAI score34

    Microsoft AI launches MAI-Transcribe-2 for fast, cheap transcription

    AIMicrosoft AI's MAI-Transcribe-2 is a transcription model that the company says offers the highest quality and lowest price at the fastest speed. Microsoft claims it runs 10x faster than GPT-Transcribe, and it is now available on Microsoft Foundry.

    Video from @MicrosoftAI's post
  8. BAAI · new models on Hugging FaceOfficialAI score25

    BAAI Releases Recon2Reason-Reasoning-4B, a Spatial Reasoning Vision-Language Model

    AIBAAI released Recon2Reason-Reasoning-4B, a 4,437,815,808-parameter vision-language model fine-tuned from Qwen3-VL-4B-Instruct for indoor spatial reasoning. The model handles metric distance, relative position, and object-relation questions from single or multiple images, and loads with the standard Qwen3VLForConditionalGeneration interface without trust_remote_code. The checkpoint is released under Apache-2.0 with BF16 Safetensors weights, and the retrieval-augmented scene-reconstruction extension ships separately.

  9. Gemini API ChangelogOfficialAI score50

    Google releases Lyria 3.5 music generation model with full-length song generation

    AIGoogle has made Lyria 3.5 generally available through the Gemini API, offering full-length song generation with improved musical coherence, natural vocals, and fine-grained control over duration and structure. The model model accepts text and image inputs and outputs high-fidelity 44.1 kHz stereo audio.

Sep 2

Sep 2Wed
  1. Noah ZwebenXAI score30

    Claude Tag paired with Fable 5.1 impresses in Slack demo

    AINoah Zweben of Anthropic shared a reaction to Claude Tag combined with Fable 5.1, calling the pairing impressive. The post background shows Claude Tag in Slack building a leadership deck from metrics data and flagging a vendor report that conflicts with those numbers, and Claude Tag is available in Slack on Team and Enterprise plans.

  2. NVIDIA · new models on Hugging FaceOfficialAI score36

    NVIDIA Releases EgoHand-1.0 Model for Single-Image 3D Hand Pose Estimation

    AINVIDIA released EgoHand-1.0, a 883.5M-parameter DINOv3-based transformer that predicts SOMA hand pose, MHR shape coefficients, and camera translation from a single 256×256 hand crop. The model is evaluated on the HOT3D egocentric benchmark and is intended for research and demonstration rather than production use. Its outputs can supply hand trajectories for training robotic manipulation policies, and it runs on NVIDIA Ampere GPUs under Linux with PyTorch.

  3. NVIDIA · new models on Hugging FaceOfficialAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.

  4. Meituan LongCatOfficialAI score29

    LongCat-2.0 now free to use in Command Code

    AIMeituan's LongCat-2.0 is now available free in Command Code, the AI coding tool. The model has 1.6T parameters, 48B activated, and a 1M-token context window. It is available on all Command Code plans to all subscribers.

  5. Google AI StudioOfficialAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

  6. Understanding AI (Timothy B. Lee)BlogAI score62

    How Google's RT-2 set the template for today's robotics models

    AIGoogle's RT-2 model, announced in July 2023, trained a multimodal LLM to output robot actions directly, and the article argues this approach launched the current robotics boom. The author follows later work from Physical Intelligence, including action chunking with flow matching, reinforcement learning on real robots, and visual subgoal generation, and notes that the field is debating whether vision-language-action models will give way to world models.

  7. Sundar PichaiXAI score62

    Google introduces Gemini 3.8 Flash Cyber, a cybersecurity model for vulnerability work

    AIGoogle introduces Gemini 3.8 Flash Cyber, which it describes as its most capable cybersecurity model. The company reports 86.2% on CyberGym, 47.2% on CWE-Bench for patching, and a 70%+ success rate in discovering vulnerabilities across 20 programming languages on its internal benchmark. Google says the model offers frontier-level performance at Flash-level speed and pricing.

    Why it matters: The post gives benchmark numbers for vulnerability discovery and patching, letting readers compare the model against its predecessor and rival systems in the chart.

    Image from @sundarpichai's post
  8. Sundar PichaiXAI score62

    Google introduces Gemini 3.8 Flash, its third Flash release in six weeks

    AIwith gains over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning. Sundar Pichai says it outperforms most larger frontier models on DeepSWE v1.1 at a fraction of the cost. The comparison table lists input at $0.75 and output at $3.75 per 1M tokens, with introductory pricing through December 31, 2026.

    Why it matters: The benchmark table compares Gemini 3.8 Flash with Gemini 3.7 Flash and rival models on price, coding, agent, and reasoning tasks, which helps readers judge the tradeoffs.

    Image from @sundarpichai's post
  9. Varun MohanXAI score57

    Gemini 3.8 Flash released with gains in agentic coding and knowledge work

    AIGoogle's Gemini 3.8 Flash is out, and Varun Mohan says it substantially improves on 3.7 Flash for agentic coding and general knowledge work. It is now available to everyone on Antigravity. The attached benchmark table lists Gemini 3.8 Flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, with introductory pricing through December 31, 2026.

    Image from @_mohansolo's post
  10. Josh WoodwardXAI score21

    Gemini 3.8 Flash Touted for Quality at a Great Price

    AIJosh Woodward, who leads Google's Gemini app, says Gemini 3.8 Flash delivers strong quality at a great price. The post gives no benchmark scores, pricing figures, or availability details.

    Image from @joshwoodward's post
  11. Google AI DevelopersOfficialAI score32

    Gemini 3.8 Flash builds interactive 3D hardware teardown visualizers with Three.js

    AIGoogle AI Developers says Gemini 3.8 Flash, built for complex reasoning, generated an interactive 3D visualizer using Three.js in Google AI Studio. The visualizer produces physically proportioned teardowns of hardware devices, automatically splitting each device into layers that users can explode and inspect with a deconstruction slider.

    Video from @googleaidevs's post
  12. Josh WoodwardXAI score44

    Gemini 3.8 Flash rolls out to the Gemini app

    AIGoogle is rolling out Gemini 3.8 Flash to the Gemini app, according to Josh Woodward's post. The post provides no further details on capabilities, benchmarks, pricing, or availability timing.

    Image from @joshwoodward's post
  13. koray kavukcuogluXAI score62

    Gemini 3.8 Flash claims stronger engineering results at lower cost than larger models

    AIGoogle's Koray Kavukcuoglu says Gemini 3.8 Flash is a major step up from Gemini 3.7 Flash and outperforms most larger frontier models on complex engineering problems at a fraction of the cost. The attached DeepSWE V1.1 chart, sourced to Datacurve AI, plots average cost per task against score for Gemini 3.8 Flash and other models. A link to Google's blog post with more details is included.

    Why it matters: The chart compares Gemini 3.8 Flash's DeepSWE score and average cost per task against several frontier models, showing where the cost-performance tradeoff lands.

    Image from @koraykv's post
  14. koray kavukcuogluXAI score62

    Google launches Gemini 3.8 Flash Cyber and Gemini 3.8 Flash models

    AIGoogle launches Gemini 3.8 Flash Cyber and Gemini 3.8 Flash. The post describes Flash Cyber as its most capable cybersecurity model for finding and fixing vulnerabilities, placing it on the Pareto frontier for patching on CWE-Bench. Flash Cyber is available to trusted defenders through the new Fairwind Program.

    Why it matters: The chart compares Pass@1 against cost per rollout, showing where Gemini 3.8 Flash Cyber sits relative to frontier and budget models on CWE-Bench.

    Image from @koraykv's post
  15. Logan KilpatrickXAI score60

    Google releases Gemini 3.8 Flash at the same price and similar speed as 3.7

    AIGoogle's Logan Kilpatrick says Gemini 3.8 Flash costs the same as Gemini 3.7 and runs at roughly the same speed. The model is available through the Gemini API, AI Studio, Antigravity, the Gemini App, and other surfaces.

    Why it matters: The post frames the new Flash version by comparing price and speed with its predecessor, which helps readers gauge the upgrade's practical trade-offs.

  16. Logan KilpatrickXAI score62

    Google releases Gemini 3.8 Flash with gains in agentic and coding tasks

    AIGoogle announced Gemini 3.8 Flash, its third updated Flash model in six weeks, citing improvements in agentic and coding capabilities. The benchmark table lists input at $0.75 and output at $3.75 per 1M tokens, with introductory pricing of $1.50 and $7.50 expiring December 31, 2026. Terminal-bench 2.1 shows 89.4% for Gemini 3.8 Flash against 85.8% for Gemini 3.7 Flash.

    Why it matters: The benchmark table compares Gemini 3.8 Flash against Gemini 3.7 Flash and rival models, showing where the gains and remaining gaps fall across coding and agent tasks.

    Image from @OfficialLoganK's post
  17. Google AI StudioOfficialAI score62

    Google releases Gemini 3.8 Flash with improved coding, agent, and reasoning

    AIGoogle AI Studio announced Gemini 3.8 Flash, which it calls its most intelligent workhorse model. The company says it brings significant improvements over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, through the Gemini API and AI Studio.

    Why it matters: The post gives specific pricing and access paths for the new model, letting developers compare it with 3.7 Flash on cost and availability.

    Image from @GoogleAIStudio's post
  18. Cohere · new models on Hugging FaceOfficialAI score44

    Cohere Releases Tiny Aya En-Thinker, a 3.35B Multilingual Reasoning Model

    AICohere Labs released Tiny Aya En-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model with a 32K context length. It is trained on English reasoning traces for 44 languages plus English, with coverage extending to 20+ more languages through non-reasoning instruction data. The model is available under a CC-BY-NC license that also requires adherence to Cohere Labs' Acceptable Use Policy.

  19. Cohere · new models on Hugging FaceOfficialAI score44

    Cohere Releases Tiny Aya L2-Thinker Multilingual Reasoning Model on Hugging Face

    AICohere Labs released Tiny Aya L2-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model that thinks in the same language as the user's prompt before answering. The model supports in-language reasoning for 44 languages plus English, with coverage extended to 20+ more languages through additional non-reasoning instruction data, and has a 32K context length. It is licensed under CC-BY-NC and is available on Hugging Face.

  20. Unsloth AIOfficialAI score62

    Qwen3.8-Flash-Next runs 1.3 to 1.7 times faster locally with MTP

    AIUnsloth says MTP enables Qwen3.8-Flash-Next to run about 1.3 to 1.7 times faster at inference with no accuracy change. GGUF versions can reach 170 tokens/s on an RTX PRO 6000, and the source lists memory requirements from 76 GB at 1-bit to 355 GB at BF16.

    Why it matters: The source gives concrete MTP speedup ranges, hardware memory requirements, and GGUF quantization sizes, which help readers judge whether local deployment fits their setup.

    Image from @UnslothAI's post