Skip to contentSkip to stories

Updated

#Model release

Sep 17

Sep 17Thu
  1. inclusionAI (Ant Ling) · new models on Hugging FaceAI score46

    Ming-Image-0.1-Design-Layer splits flattened design images into RGBA layers

    AIinclusionAI has released Ming-Image-0.1-Design-Layer on Hugging Face, a model that decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan. The model runs at 1024 resolution (512 for faster processing) with 12 sampling steps, a CFG scale of 2.0, and BF16 precision on one CUDA GPU with 80 GiB VRAM. It is released under the MIT License.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceAI score42

    inclusionAI releases Ming-Image-0.1-Design, a 6B text-to-image model for text-rich designs

    AIinclusionAI has released Ming-Image-0.1-Design, a 6B text-to-image model for UI, infographics, and posters that outputs RGBA images with transparent backgrounds. The model is available on Hugging Face and ModelScope under the MIT License. It runs at 2048 x 2048 with 12 sampling steps and a CFG scale of 1.0, validated on one CUDA GPU with 80 GiB VRAM.

  3. Z.aiAI score40

    GLM-5.3 helped build the inference stack serving GLM-5.3-Flash

    AIZ.ai reports that GLM-5.3 helped build and optimize the inference infrastructure for GLM-5.3-Flash. The system went from first successful run to production readiness in under two weeks, with end-to-end throughput tripling over the initial baseline. The team credited dense feedback from local correctness tests, execution traces, microbenchmarks, and end-to-end measurements for enabling targeted hypothesis testing.

Sep 16

Sep 16Wed
  1. inclusionAI (Ant Ling) · new models on Hugging FaceAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Google AI StudioAI score72

    Google releases Gemini 3.8 Live and 3.5 Transcribe for real-time voice apps

    AIGoogle AI Studio released Gemini 3.8 Live, a native speech-to-speech model with an Extended Thinking variant, and made it available through the Live API. Gemini 3.5 Transcribe, released last month, supports 85+ languages with a reported 4.0% streaming and 2.6% non-streaming Word Error Rate, and accepts a custom vocabulary of up to 1,000 terms. Live API audio pricing is listed at $0.005/min for input and $0.018/min for output.

    Why it matters: The post lists concrete Live API capabilities, per-minute audio pricing, and transcription accuracy figures, helping developers weigh voice agent options against their own cascaded pipelines.

  4. Lewis TunstallAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.

  5. Google AI StudioAI score46

    Google launches Gemini 3.8 Live and Extended Thinking dialogue models

    AIGoogle introduced two live dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, available through AI Studio and the Gemini API. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. The Extended Thinking variant targets high-complexity tasks with increased intelligence and multi-step reasoning.

  6. Sundar PichaiAI score42

    Google outlines AI for science, weather, languages, and economic research

    AIGoogle says it is focusing AI efforts on health, disaster and weather resilience, learning, and economic opportunity. Recent examples include AlphaGenome Atlas, which maps all 9B possible single-letter genetic changes across the human genome and is openly available to researchers, and WeatherNext 3, described as its most accurate and capable global weather AI model to date. The post also cites AI & Economy ATLAS, an open-access look at global AI usage, and says its translation services now cover nearly 300 languages spoken by 7B people.

  7. Google AIAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    AIGoogle is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    Why it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  8. Google AIAI score62

    Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking audio models

    AIGoogle AI announces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced Gemini Audio models. Gemini 3.8 Live is built for scale, speed, and cost efficiency, handling mid-sentence interruptions, transitions across 97 languages, and visual context through Search Live. Gemini 3.8 Live Extended Thinking reasons and speaks in parallel, narrating its progress on multi-step tasks such as event planning.

  9. Google DeepMindAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  10. Google AI StudioAI score72

    Google launches Gemini 3.8 Live and Extended Thinking voice models

    AIGoogle introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents that reason and speak simultaneously. The Extended Thinking version scores 82.6 on Artificial Analysis' Speech to Speech Quality Index and 97.7% on Big Bench Audio, while 3.8 Live targets scale and cost efficiency. Developers can access both through the Gemini API in Google AI Studio, and enterprise and consumer rollouts vary by product.

    Why it matters: The source names the two models, their access paths, and specific benchmark results, showing how the voice agent capabilities differ between the two tiers.

  11. Google DeepMindAI score33

    Gemini 3.8 Live Extended Thinking adds upgraded reasoning for real-time programming tutoring.

    AIGoogle DeepMind demonstrated 3.8 Live Extended Thinking acting as a programming tutor in Gemini Live. Both 3.8 Live models feature upgraded reasoning, near real-time visual understanding, automatic detection across 97 languages, and background tool calling that doesn't interrupt the chat. The Extended Thinking variant adds higher performance and precision for harder tasks and narrates its progress, and it is available in Gemini Live in the Gemini app or through the Gemini API via Google AI Studio.

  12. NVIDIA · new models on Hugging FaceAI score34

    NVIDIA Releases RT-DETR Hand Detection v1.0 for Real-Time RGB Hand Localization

    AINVIDIA's RT-DETR Hand Detection v1.0 detects and localizes left and right hands in RGB images, outputting 2D bounding boxes with per-hand confidence scores in a single pass. The model, built on RT-DETRv2-S with HGNetv2-S backbone and about 20M parameters, is intended as a region-of-interest stage for downstream 3D hand pose estimation and is exported to ONNX. The source describes it as for demonstration purposes rather than production use, runs on NVIDIA Lovelace GPUs under Linux, and is licensed under the NVIDIA Software and Model Evaluation License.

  13. Baseten BlogAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.

  14. Gemini API ChangelogAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    AIGoogle has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    Why it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 14

Sep 14Mon
  1. Intern Large ModelsAI score23

    Intern-S2-397B from Intern Large Models gets SGLang Day-0 support

    AI🚀 SGLang has Day-0 support for Intern-S2-397B from @intern_lm, a 397B multimodal foundation model built for scientific intelligence and long-horizon agents. New pre-training paradigm: learns directly from raw scientific literature pages, no parsing needed. Scientific reasoning: RL across 20+ domains, from biomolecule design to material generation. Long-horizon agents: black-box agentic RL in large-scale sandboxed environments. Run it now with SGLang!

  2. NVIDIA · new models on Hugging FaceAI score40

    NVIDIA releases FoundationStereo small stereo depth model on Hugging Face

    AINVIDIA Research released FoundationStereo-small, a zero-shot stereo depth model that takes an RGB stereo pair and outputs a disparity map, on Hugging Face. The model has about 6.3×10^7 parameters and ships as ONNX files at fixed 576x960 and 320x736 resolutions, with TensorRT and ONNX runtime support. It is licensed under the NVIDIA Open Model License and is ready for commercial use.

  3. NVIDIA · new models on Hugging FaceAI score36

    NVIDIA's FoundationPose estimates 6-DoF object pose without fine-tuning given a CAD model

    AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.

  4. Google · new models on Hugging FaceAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model

    AIGoogle DeepMind released EmbeddingGemma 2, an open model under Apache 2.0 that maps text, images, video, and audio into one shared 768-dimensional vector space. The model has 740M total parameters and supports 8,192-token context, with Matryoshka truncation to 128d, 256d, and 512d. The source reports 14% better code-task performance than EmbeddingGemma 1 and says it is designed for consumer hardware such as phones and laptops.

    Why it matters: The release combines text, image, video, and audio retrieval in one 768-dimensional space at 740M parameters, a useful reference for on-device multimodal search design.

  5. Intern Large ModelsAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

  6. Tencent · new models on Hugging FaceAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

Sep 13

Sep 13Sun
  1. Qwen · new models on Hugging FaceAI score67

    Qwen releases open-source Qwen-Image-2.1 for generation and editing

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B parameters in its visual generation component. The model can generate regular or transparent RGBA images, supports up to 10 reference images for editing, and is licensed under the Qwen Research License Agreement.

    Why it matters: The source specifies the 7B visual component, transparent RGBA output, and up to 10 reference images, which helps readers judge its fit for generation and editing workflows.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceAI score36

    SingProbe adds a streaming guardrail to Step-3.7-Flash without a separate safety model

    AIinclusionAI released Step-3.7-Flash-singprobe, an 8.13M-parameter probe that reuses Step-3.7-Flash hidden states to score query intent, response unsafety, and hallucination risk at every generated token. The probe adds less than 0.5% decode-time overhead and reports 0.9858 R-AUC and 0.9295 T-AUC on streaming safety benchmarks. It is supported through SGLang and vLLM integration branches and loads from Hugging Face by checkpoint ID.

  3. inclusionAI (Ant Ling) · new models on Hugging FaceAI score38

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.8-27B

    AIinclusionAI has released Qwen3.8-27B-singprobe, a 10.1M-parameter intrinsic streaming guardrail that reuses Qwen3.8-27B hidden states to score query intent, response unsafety, and hallucination risk at every token. The probe adds less than 0.5% decode-time overhead and reports a 0.03% benign-response false-positive rate averaged across five datasets. It is supported through SGLang and vLLM integration branches, with training code available at inclusionAI/SingProbe.