Skip to contentSkip to stories

Updated

#Model release

Showing low-relevance items too. Hide low-relevance items

Sep 2

Sep 2Wed
  1. Meituan LongCatOfficialAI score29

    LongCat-2.0 now free to use in Command Code

    AIMeituan's LongCat-2.0 is now available free in Command Code, the AI coding tool. The model has 1.6T parameters, 48B activated, and a 1M-token context window. It is available on all Command Code plans to all subscribers.

  2. Google AI StudioOfficialAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

  3. Understanding AI (Timothy B. Lee)BlogAI score62

    How Google's RT-2 set the template for today's robotics models

    AIGoogle's RT-2 model, announced in July 2023, trained a multimodal LLM to output robot actions directly, and the article argues this approach launched the current robotics boom. The author follows later work from Physical Intelligence, including action chunking with flow matching, reinforcement learning on real robots, and visual subgoal generation, and notes that the field is debating whether vision-language-action models will give way to world models.

  4. Sundar PichaiXAI score62

    Google introduces Gemini 3.8 Flash Cyber, a cybersecurity model for vulnerability work

    AIGoogle introduces Gemini 3.8 Flash Cyber, which it describes as its most capable cybersecurity model. The company reports 86.2% on CyberGym, 47.2% on CWE-Bench for patching, and a 70%+ success rate in discovering vulnerabilities across 20 programming languages on its internal benchmark. Google says the model offers frontier-level performance at Flash-level speed and pricing.

    Image from @sundarpichai's post
  5. Sundar PichaiXAI score62

    Google introduces Gemini 3.8 Flash, its third Flash release in six weeks

    AIwith gains over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning. Sundar Pichai says it outperforms most larger frontier models on DeepSWE v1.1 at a fraction of the cost. The comparison table lists input at $0.75 and output at $3.75 per 1M tokens, with introductory pricing through December 31, 2026.

    Image from @sundarpichai's post
  6. Varun MohanXAI score57

    Gemini 3.8 Flash released with gains in agentic coding and knowledge work

    AIGoogle's Gemini 3.8 Flash is out, and Varun Mohan says it substantially improves on 3.7 Flash for agentic coding and general knowledge work. It is now available to everyone on Antigravity. The attached benchmark table lists Gemini 3.8 Flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, with introductory pricing through December 31, 2026.

    Image from @_mohansolo's post
  7. Josh WoodwardXAI score21

    Gemini 3.8 Flash Touted for Quality at a Great Price

    AIJosh Woodward, who leads Google's Gemini app, says Gemini 3.8 Flash delivers strong quality at a great price. The post gives no benchmark scores, pricing figures, or availability details.

    Image from @joshwoodward's post
  8. Google AI DevelopersOfficialAI score32

    Gemini 3.8 Flash builds interactive 3D hardware teardown visualizers with Three.js

    AIGoogle AI Developers says Gemini 3.8 Flash, built for complex reasoning, generated an interactive 3D visualizer using Three.js in Google AI Studio. The visualizer produces physically proportioned teardowns of hardware devices, automatically splitting each device into layers that users can explode and inspect with a deconstruction slider.

    Video from @googleaidevs's post
  9. Josh WoodwardXAI score44

    Gemini 3.8 Flash rolls out to the Gemini app

    AIGoogle is rolling out Gemini 3.8 Flash to the Gemini app, according to Josh Woodward's post. The post provides no further details on capabilities, benchmarks, pricing, or availability timing.

    Image from @joshwoodward's post
  10. koray kavukcuogluXAI score62

    Gemini 3.8 Flash claims stronger engineering results at lower cost than larger models

    AIGoogle's Koray Kavukcuoglu says Gemini 3.8 Flash is a major step up from Gemini 3.7 Flash and outperforms most larger frontier models on complex engineering problems at a fraction of the cost. The attached DeepSWE V1.1 chart, sourced to Datacurve AI, plots average cost per task against score for Gemini 3.8 Flash and other models. A link to Google's blog post with more details is included.

    Image from @koraykv's post
  11. koray kavukcuogluXAI score62

    Google launches Gemini 3.8 Flash Cyber and Gemini 3.8 Flash models

    AIGoogle launches Gemini 3.8 Flash Cyber and Gemini 3.8 Flash. The post describes Flash Cyber as its most capable cybersecurity model for finding and fixing vulnerabilities, placing it on the Pareto frontier for patching on CWE-Bench. Flash Cyber is available to trusted defenders through the new Fairwind Program.

    Image from @koraykv's post
  12. Logan KilpatrickXAI score62

    Google releases Gemini 3.8 Flash with gains in agentic and coding tasks

    AIGoogle announced Gemini 3.8 Flash, its third updated Flash model in six weeks, citing improvements in agentic and coding capabilities. The benchmark table lists input at $0.75 and output at $3.75 per 1M tokens, with introductory pricing of $1.50 and $7.50 expiring December 31, 2026. Terminal-bench 2.1 shows 89.4% for Gemini 3.8 Flash against 85.8% for Gemini 3.7 Flash.

    Image from @OfficialLoganK's post
  13. Google AI StudioOfficialAI score62

    Google releases Gemini 3.8 Flash with improved coding, agent, and reasoning

    AIGoogle AI Studio announced Gemini 3.8 Flash, which it calls its most intelligent workhorse model. The company says it brings significant improvements over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, through the Gemini API and AI Studio.

    Image from @GoogleAIStudio's post
  14. Cohere · new models on Hugging FaceOfficialAI score44

    Cohere Releases Tiny Aya En-Thinker, a 3.35B Multilingual Reasoning Model

    AICohere Labs released Tiny Aya En-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model with a 32K context length. It is trained on English reasoning traces for 44 languages plus English, with coverage extending to 20+ more languages through non-reasoning instruction data. The model is available under a CC-BY-NC license that also requires adherence to Cohere Labs' Acceptable Use Policy.

  15. Cohere · new models on Hugging FaceOfficialAI score44

    Cohere Releases Tiny Aya L2-Thinker Multilingual Reasoning Model on Hugging Face

    AICohere Labs released Tiny Aya L2-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model that thinks in the same language as the user's prompt before answering. The model supports in-language reasoning for 44 languages plus English, with coverage extended to 20+ more languages through additional non-reasoning instruction data, and has a 32K context length. It is licensed under CC-BY-NC and is available on Hugging Face.

  16. Unsloth AIOfficialAI score62

    Qwen3.8-Flash-Next runs 1.3 to 1.7 times faster locally with MTP

    AIUnsloth says MTP enables Qwen3.8-Flash-Next to run about 1.3 to 1.7 times faster at inference with no accuracy change. GGUF versions can reach 170 tokens/s on an RTX PRO 6000, and the source lists memory requirements from 76 GB at 1-bit to 355 GB at BF16.

    Image from @UnslothAI's post

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeOfficialAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  2. Mike KriegerXAI score36

    Anthropic's new model works to targets and admits when stuck

    AIThe model keeps working until it reaches a given target and says so when it is stuck, rather than reporting false success. It is priced at $10/$50, the same as Fable 5, and cache reads on the API are cut 75% to $0.25/MTok.

  3. Alex AlbertXAI score62

    Alex Albert says Claude Fable 5.1 works from vague, messy instructions

    AIAlex Albert describes Claude Fable 5.1 as a model that fills in gaps from vague, messy instructions the way he would. He calls it impressive in many ways and encourages people to try it. The quoted post from @claudeai announces Claude Fable 5.1 and Claude Mythos 5.1 as the world's most advanced models for coding and knowledge work.

  4. Anthropic · YouTubeOfficialAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

  5. Ai2 · new models on Hugging FaceOfficialAI score22

    Ai2 Releases Supplemental ACE2S-SHiELD+ Ablation Checkpoints on Hugging Face

    AIAi2 has published supplemental checkpoints for its ACE2S-SHiELD+ climate model on Hugging Face, covering four ablation configurations that test random CO2 data and energy conservation. Each configuration includes two random-seed models, and the repository recommends the main ACE2S-SHiELD+ checkpoint for most uses. The checkpoints are licensed under Apache 2.0 for research and educational use.

  6. Fei-Fei LiXAI score62

    World Labs unveils Atlas, a multimodal world model with camera control

    AIWorld Labs has introduced Atlas, a multimodal world model it describes as trained from scratch. The post says Atlas generates frames with pixel-perfect camera control, reconstructs large scenes from as few as one input image, and outputs 3D spaces from one or more images. The author cites use cases including VFX and robotics.

  7. World LabsOfficialAI score30

    World Labs unveils Atlas, a scalable world foundation model

    AIWorld Labs announces Atlas, a scalable foundation model that can perceive, generate, reason, and interact with virtual and physical worlds. The company says early access opens in the coming weeks, with sign-up details on its blog.

  8. World LabsOfficialAI score34

    World Labs' Atlas generates explicit 3D from one or many images

    AIWorld Labs says its Atlas model outputs explicit 3D reconstructions from one or many input images, outperforming top open-source reconstruction models. It adds that passing more images gives Atlas more context, so it relies less on imagination as it sees more views.

    Video from @theworldlabs's post
  9. World LabsOfficialAI score35

    World Labs pre-trains Atlas to turn multimodal inputs into 3D views

    AIWorld Labs says it pre-trained Atlas from scratch to accept multimodal inputs, including camera movement, and convert them into 3D-grounded views. The company says Atlas lets users direct views, reconstruct real spaces from their inputs, and build explorable worlds.

  10. World LabsOfficialAI score46

    World Labs unveils Atlas, a multimodal world model with 3D reconstruction

    AIWorld Labs has introduced Atlas, which it describes as the first multimodal world model that generates image and video frames with pixel-perfect camera control. The model can also reconstruct the generated frames in 3D, letting users move the camera and simulate space and time.

    Video from @theworldlabs's post
  11. Google · new models on Hugging FaceOfficialAI score44

    Google Releases GNM v3.0, an Open 3D Parametric Model of the Human Head

    AIGoogle has released GNM v3.0, a parametric 3D statistical model of the human head, with weights published on Hugging Face and Kaggle under the Apache 2.0 license. The model gives controllable identity, expression, head pose, and internal anatomy including eyeballs, teeth, and tongue, and supports NumPy, JAX, PyTorch, and TensorFlow backends.

  12. Tencent HyOfficialAI score58

    Tencent Hy4 preview reports 31.8% throughput gain from self-found bottlenecks

    AITencent Hunyuan says its Hy4 preview model found inference bottlenecks on its own and raised end-to-end throughput by 31.8% through operator fusion and communication optimizations. The post says the gain holds across context lengths and concurrency levels. The release is listed at 770B total parameters with 49B active and a 1M context window, with links to the Hy blog, Hugging Face, and GitHub.

    Image from @TencentHunyuan's post
  13. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

  14. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesOfficialAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

  2. Microsoft ResearchOfficialAI score45

    GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Population-Scale Research

    AIMicrosoft Research released GigaPath-Flash and GigaTIME-Flash, efficient pathology foundation models built on a distilled ViT-S backbone and released under the Apache 2.0 license. GigaPath-Flash, with 22M-parameter tile and 21M-parameter slide encoders, reportedly scores within 3% of the original GigaPath on PANDA and EBRAINS benchmarks at roughly 50 times less compute. The models are research tools, not validated for clinical use.

  3. DeepSeek · new models on Hugging FaceOfficialAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.