Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 3

Sep 3Thu
  1. BAAI · new models on Hugging FaceAI score25

    BAAI Releases Recon2Reason-Reasoning-4B, a Spatial Reasoning Vision-Language Model

    AIBAAI released Recon2Reason-Reasoning-4B, a 4,437,815,808-parameter vision-language model fine-tuned from Qwen3-VL-4B-Instruct for indoor spatial reasoning. The model handles metric distance, relative position, and object-relation questions from single or multiple images, and loads with the standard Qwen3VLForConditionalGeneration interface without trust_remote_code. The checkpoint is released under Apache-2.0 with BF16 Safetensors weights, and the retrieval-augmented scene-reconstruction extension ships separately.

Sep 2

Sep 2Wed
  1. NVIDIA · new models on Hugging FaceAI score36

    NVIDIA Releases EgoHand-1.0 Model for Single-Image 3D Hand Pose Estimation

    AINVIDIA released EgoHand-1.0, a 883.5M-parameter DINOv3-based transformer that predicts SOMA hand pose, MHR shape coefficients, and camera translation from a single 256×256 hand crop. The model is evaluated on the HOT3D egocentric benchmark and is intended for research and demonstration rather than production use. Its outputs can supply hand trajectories for training robotic manipulation policies, and it runs on NVIDIA Ampere GPUs under Linux with PyTorch.

  2. NVIDIA · new models on Hugging FaceAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.

  3. Google AI StudioAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

  4. Sundar PichaiAI score62

    Google introduces Gemini 3.8 Flash Cyber, a cybersecurity model for vulnerability work

    AIGoogle introduces Gemini 3.8 Flash Cyber, which it describes as its most capable cybersecurity model. The company reports 86.2% on CyberGym, 47.2% on CWE-Bench for patching, and a 70%+ success rate in discovering vulnerabilities across 20 programming languages on its internal benchmark. Google says the model offers frontier-level performance at Flash-level speed and pricing.

    Image from @sundarpichai's post
  5. Varun MohanAI score57

    Gemini 3.8 Flash released with gains in agentic coding and knowledge work

    AIGoogle's Gemini 3.8 Flash is out, and Varun Mohan says it substantially improves on 3.7 Flash for agentic coding and general knowledge work. It is now available to everyone on Antigravity. The attached benchmark table lists Gemini 3.8 Flash at $0.75 per 1M input tokens and $3.75 per 1M output tokens, with introductory pricing through December 31, 2026.

    Image from @_mohansolo's post
  6. koray kavukcuogluAI score62

    Gemini 3.8 Flash claims stronger engineering results at lower cost than larger models

    AIGoogle's Koray Kavukcuoglu says Gemini 3.8 Flash is a major step up from Gemini 3.7 Flash and outperforms most larger frontier models on complex engineering problems at a fraction of the cost. The attached DeepSWE V1.1 chart, sourced to Datacurve AI, plots average cost per task against score for Gemini 3.8 Flash and other models. A link to Google's blog post with more details is included.

    Image from @koraykv's post
  7. Logan KilpatrickAI score62

    Google releases Gemini 3.8 Flash with gains in agentic and coding tasks

    AIGoogle announced Gemini 3.8 Flash, its third updated Flash model in six weeks, citing improvements in agentic and coding capabilities. The benchmark table lists input at $0.75 and output at $3.75 per 1M tokens, with introductory pricing of $1.50 and $7.50 expiring December 31, 2026. Terminal-bench 2.1 shows 89.4% for Gemini 3.8 Flash against 85.8% for Gemini 3.7 Flash.

    Image from @OfficialLoganK's post
  8. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash with improved coding, agent, and reasoning

    AIGoogle AI Studio announced Gemini 3.8 Flash, which it calls its most intelligent workhorse model. The company says it brings significant improvements over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash, $0.75 per million input tokens and $3.75 per million output tokens, through the Gemini API and AI Studio.

    Image from @GoogleAIStudio's post
  9. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya En-Thinker, a 3.35B Multilingual Reasoning Model

    AICohere Labs released Tiny Aya En-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model with a 32K context length. It is trained on English reasoning traces for 44 languages plus English, with coverage extending to 20+ more languages through non-reasoning instruction data. The model is available under a CC-BY-NC license that also requires adherence to Cohere Labs' Acceptable Use Policy.

  10. Cohere · new models on Hugging FaceAI score44

    Cohere Releases Tiny Aya L2-Thinker Multilingual Reasoning Model on Hugging Face

    AICohere Labs released Tiny Aya L2-Thinker, an open-weights 3.35 billion parameter multilingual reasoning model that thinks in the same language as the user's prompt before answering. The model supports in-language reasoning for 44 languages plus English, with coverage extended to 20+ more languages through additional non-reasoning instruction data, and has a 32K context length. It is licensed under CC-BY-NC and is available on Hugging Face.

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  2. Anthropic · YouTubeAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

  3. Ai2 · new models on Hugging FaceAI score22

    Ai2 Releases Supplemental ACE2S-SHiELD+ Ablation Checkpoints on Hugging Face

    AIAi2 has published supplemental checkpoints for its ACE2S-SHiELD+ climate model on Hugging Face, covering four ablation configurations that test random CO2 data and energy conservation. Each configuration includes two random-seed models, and the repository recommends the main ACE2S-SHiELD+ checkpoint for most uses. The checkpoints are licensed under Apache 2.0 for research and educational use.

  4. Google · new models on Hugging FaceAI score44

    Google Releases GNM v3.0, an Open 3D Parametric Model of the Human Head

    AIGoogle has released GNM v3.0, a parametric 3D statistical model of the human head, with weights published on Hugging Face and Kaggle under the Apache 2.0 license. The model gives controllable identity, expression, head pose, and internal anatomy including eyeballs, teeth, and tongue, and supports NumPy, JAX, PyTorch, and TensorFlow backends.

  5. Tencent HyAI score58

    Tencent Hy4 preview reports 31.8% throughput gain from self-found bottlenecks

    AITencent Hunyuan says its Hy4 preview model found inference bottlenecks on its own and raised end-to-end throughput by 31.8% through operator fusion and communication optimizations. The post says the gain holds across context lengths and concurrency levels. The release is listed at 770B total parameters with 49B active and a 1M context window, with links to the Hy blog, Hugging Face, and GitHub.

    Image from @TencentHunyuan's post
  6. OpenBMB (MiniCPM) · new models on Hugging FaceAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

  7. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Aug 31

Aug 31Mon
  1. Claude Apps Release NotesAI score72

    Anthropic launches Claude Fable 5.1 and Claude Mythos 5.1 models

    AIAnthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, which it describes as the world's most advanced models for coding and knowledge work. The release notes link to a blog post with more details, but the notes themselves give no benchmarks or specifications.

    Why it matters: The source names two new model versions and points to a companion blog post, so readers can compare the release details there.

  2. Microsoft ResearchAI score45

    GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Population-Scale Research

    AIMicrosoft Research released GigaPath-Flash and GigaTIME-Flash, efficient pathology foundation models built on a distilled ViT-S backbone and released under the Apache 2.0 license. GigaPath-Flash, with 22M-parameter tile and 21M-parameter slide encoders, reportedly scores within 3% of the original GigaPath on PANDA and EBRAINS benchmarks at roughly 50 times less compute. The models are research tools, not validated for clinical use.

  3. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.