Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 16

Sep 16Wed
  1. OpenBMBOfficialAI score20

    OpenBMB praises Dubedo's VoxCPM2-based voice cloning and dubbing studio

    AIOpenBMB says Dubedo is the kind of product it hoped VoxCPM2 would enable, citing speaker-aware cloning, multilingual generation, and an editing studio. The post praises @dubedostudio's work, while the background post describes Dubedo as dubbing into 30 languages with per-speaker voice cloning and a beta open for trial.

  2. Kling AIOfficialAI score38

    Fountain 0's ODYSSEUS: The Fall, fully generated with Kling 3.0, is out

    AIFountain 0's new feature film ODYSSEUS: The Fall, directed by Ash Koosha, is now available in full with every shot generated by Kling 3.0. The team previously premiered Dreams of Violets at the 2026 Tribeca Festival as the first AI feature film accepted into a major film festival.

    Video from @Kling_ai's post
  3. Google Developers BlogOfficialAI score38

    Google and Speakeasy open-source OpenAPI SDK generator suite under AGPLv3 license

    AISpeakeasy is open-sourcing its full OpenAPI client suite under the AGPLv3 license, including generators for seven languages (Python, TypeScript, Go, Java, C#, PHP, Ruby), an agent-native CLI generator, and a documentation MCP server generator. Google said the move followed the May 2026 shutdown of the SDK generation provider it had been using, which it cited as evidence that closed-source generators pose platform risk. Google's new Google GenAI SDKs for the Interactions, Agents, and Webhooks APIs were built with this pipeline across six targets.

  4. BAAI · new models on Hugging FaceOfficialAI score34

    BAAI and Peking University release Brainμ-Spike spike camera image reconstruction model

    AIPeking University's Yu Zhaofei team and the Beijing Academy of Artificial Intelligence (BAAI) released Brainμ-Spike, a small convolutional network for spike camera image reconstruction that is paired with the Brainμ model. The package includes weights, inference scripts, and evaluation tools, but the base large model and LoRA weights are not yet released, so the full generation pipeline cannot run from this repository alone.

  5. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceOfficialAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Zed BlogOfficialAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  4. Lewis Tunstall @ COLM 🌉XAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.

  5. Sundar PichaiXAI score42

    Google outlines AI for science, weather, languages, and economic research

    AIGoogle says it is focusing AI efforts on health, disaster and weather resilience, learning, and economic opportunity. Recent examples include AlphaGenome Atlas, which maps all 9B possible single-letter genetic changes across the human genome and is openly available to researchers, and WeatherNext 3, described as its most accurate and capable global weather AI model to date. The post also cites AI & Economy ATLAS, an open-access look at global AI usage, and says its translation services now cover nearly 300 languages spoken by 7B people.

    Image from @sundarpichai's post
  6. NVIDIA · new models on Hugging FaceOfficialAI score34

    NVIDIA Releases RT-DETR Hand Detection v1.0 for Real-Time RGB Hand Localization

    AINVIDIA's RT-DETR Hand Detection v1.0 detects and localizes left and right hands in RGB images, outputting 2D bounding boxes with per-hand confidence scores in a single pass. The model, built on RT-DETRv2-S with HGNetv2-S backbone and about 20M parameters, is intended as a region-of-interest stage for downstream 3D hand pose estimation and is exported to ONNX. The source describes it as for demonstration purposes rather than production use, runs on NVIDIA Lovelace GPUs under Linux, and is licensed under the NVIDIA Software and Model Evaluation License.

  7. RadixArkOfficialAI score42

    Periodic Labs builds Neon on SGLang and Miles for 2.5x faster inference

    AIPeriodic Labs chose SGLang and Miles to build Neon, an open-source model it says surpasses GPT-6 Astra on its analysis benchmark after mid-training and RL on 1,300 H200s. RadixArk says Periodic extended both frameworks for scientific RL at trillion-parameter scale, delivering more efficient training, lower memory use, and 2.5x faster inference. The work has been contributed back to both projects.

  8. LlamaIndex 🦙OfficialAI score22

    LlamaIndex Moves Off Stainless for LlamaParse SDK Generation

    AILlamaIndex says Stainless helped it keep LlamaParse SDKs current and pushed it to make the API's names and schemas more consistent. With the Stainless team joining Anthropic, George He and Yong Park explain what worked, what they learned, and why changing SDK generators needs careful handling.

    Image from @llama_index's post
  9. Google · Innovation & AIOfficialAI score52

    Google says its language technology now covers over 300 languages with new speech, data, and on-device tools

    AIGoogle reports that its technologies and products now power everyday interactions in more than 300 languages used by over 7 billion people, about 86% of the global population. The post describes new speech models, including Gemini 3.5 Live Translate and Gemini 3.5 Transcribe, plus the TranslateGemma open translation models trained across 55 languages.

  10. Leandro von WerraXAI score38

    Von Werra urges frontier AI labs to share small models and alignment recipes

    AIHugging Face's Leandro von Werra argues that frontier AI labs should release small variants of their models, share core parts of their alignment recipe, and publish tech reports with more than evaluations. He says these steps would let the wider community test model behavior and verify safety claims, rather than leaving the safety agenda to a few labs. He also calls for independent verification of alarming internal findings, with sensitive details disclosed first to an independent team.

Sep 14

Sep 14Mon
  1. Intern Large ModelsOfficialAI score23

    Intern-S2-397B, a scientific multimodal model, gets SGLang Day-0 support

    AISGLang announces Day-0 support for Intern-S2-397B from Intern Large Models, a 397B multimodal foundation model built for scientific intelligence and long-horizon agents. The model is pre-trained directly on raw scientific literature pages without parsing and uses reinforcement learning across more than 20 scientific domains, from biomolecule design to material generation. It also applies black-box agentic reinforcement learning in large-scale sandboxed environments.

  2. NVIDIA · new models on Hugging FaceOfficialAI score40

    NVIDIA releases FoundationStereo small stereo depth model on Hugging Face

    AINVIDIA Research released FoundationStereo-small, a zero-shot stereo depth model that takes an RGB stereo pair and outputs a disparity map, on Hugging Face. The model has about 6.3×10^7 parameters and ships as ONNX files at fixed 576x960 and 320x736 resolutions, with TensorRT and ONNX runtime support. It is licensed under the NVIDIA Open Model License and is ready for commercial use.

  3. NVIDIA · new models on Hugging FaceOfficialAI score36

    NVIDIA's FoundationPose estimates 6-DoF object pose without fine-tuning given a CAD model

    AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.

  4. vLLM BlogOfficialAI score53

    Novita AI open-sources Chord, a W4A16 MoE kernel for Kimi K2.x on vLLM

    AINovita AI has open-sourced Chord, a W4A16 MoE CUDA operator with BF16 activations, INT4 weights and group-32 scales, built for Kimi K2.x serving shapes. Measured per layer against public Humming, it reports 1.11–1.20x on H200 EP8 prefill, 1.17–1.33x on H200 TP8 serving, and 1.81–2.15x on B300 EP8 decode against an untuned Humming default. Integration of the grouped operators with vLLM's Humming backend is still a work in progress.

  5. Google Developers BlogOfficialAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    AIPart 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    Why it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.

  6. vLLM BlogOfficialAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  7. MiniMax (official)OfficialAI score41

    MiniMax H3 video generation exceeds 2× real-time on 8× B200

    AIMiniMax H3 with SGLang-Diffusion and VDN-H3 generates 14.4 seconds of 768p video in 9.0 seconds end-to-end after warmup on 8× B200 GPUs. Eight-step denoising takes 6.9 seconds, exceeding 2× real-time, with no measured quality regression versus dense 50-step H3 across 103 test prompts.

  8. InferactOfficialAI score42

    Inferact and Google Cloud partner to make TPUs first-class in vLLM

    AIInferact and Google Cloud announce a partnership to make Google TPUs a first-class platform in the vLLM open-source project. The collaboration targets production serving features, optimized kernels, a native PyTorch path via TorchTPU, and day-0 support for frontier model releases. A community program will offer shared TPU capacity and review and design help from vLLM core maintainers, with all outputs released as open source.

    Image from @inferact's post
  9. RadixArkOfficialAI score43

    SGLang-Diffusion runs MiniMax H3 video generation faster than playback

    AIRadixArk's SGLang-Diffusion, paired with VDN-H3, generates 14.4 seconds of 768p video in 9.0 seconds on 8× B200 GPUs. The 8-step denoising alone takes 6.9 seconds, which is over 2× real time, and the team reports no measured quality regression against dense 50-step MiniMax H3 across 103 test prompts.

  10. LlamaIndex 🦙OfficialAI score29

    LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines

    AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.

    Image from @llama_index's post
  11. Intern Large ModelsOfficialAI score62

    Intern-S2-397B released in BF16 and FP8 under Apache 2.0

    AIShanghai AI Laboratory's Intern Large Models announced Intern-S2-397B, available in BF16 and FP8 under Apache 2.0. The post reports 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both, and says it was jointly trained across 20+ scientific domains with long-horizon agent RL.

    Why it matters: The post names the benchmark scores and training scope behind Intern-S2-397B, letting readers compare its scientific and agentic claims against the table.

  12. Intern Large ModelsOfficialAI score25

    Intern-S2-397B gets Day-0 support in vLLM

    AIIntern-S2-397B, a model built for long-horizon scientific research, now has Day-0 support in vLLM. The model brings multimodal, reasoning, coding, and scientific agent capabilities, and vLLM has published a run recipe for it.

  13. Intern Large ModelsOfficialAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

    Why it matters: The post pairs a new open multimodal model with benchmark tables against named Qwen, DeepSeek, Kimi, GLM, GPT, Gemini, and Claude models, letting readers compare scientific and agentic results directly.

    Image from @intern_lm's post
  14. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

  15. MiniMax (official)OfficialAI score36

    MiniMax H3 community projects speed up open-source video generation

    AIMiniMax highlighted open-source community progress on its H3 video generation model, which it built with native stereo audio and multimodal reference control. Recent highlights include FastH3's 4-step distillation running on DGX Spark and Apple Silicon, and NVIDIA's Sol-H3 generating 15 seconds of 768p video with audio in 6.6 seconds on 8×B300 in a warm-inference benchmark. Other releases include VDN's faster-inference attention work with code and weights, and 8-step Acc-LoRAs from Alibaba PAI, with LightX2V offering 4- and 8-step Turbo LoRAs.

    Image from @MiniMax_AI's post

Sep 13

Sep 13Sun
  1. Qwen · new models on Hugging FaceOfficialAI score67

    Qwen releases open-source Qwen-Image-2.1 for generation and editing

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B parameters in its visual generation component. The model can generate regular or transparent RGBA images, supports up to 10 reference images for editing, and is licensed under the Qwen Research License Agreement.

    Why it matters: The source specifies the 7B visual component, transparent RGBA output, and up to 10 reference images, which helps readers judge its fit for generation and editing workflows.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score36

    SingProbe adds a streaming guardrail to Step-3.7-Flash without a separate safety model

    AIinclusionAI released Step-3.7-Flash-singprobe, an 8.13M-parameter probe that reuses Step-3.7-Flash hidden states to score query intent, response unsafety, and hallucination risk at every generated token. The probe adds less than 0.5% decode-time overhead and reports 0.9858 R-AUC and 0.9295 T-AUC on streaming safety benchmarks. It is supported through SGLang and vLLM integration branches and loads from Hugging Face by checkpoint ID.

  3. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score38

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.8-27B

    AIinclusionAI has released Qwen3.8-27B-singprobe, a 10.1M-parameter intrinsic streaming guardrail that reuses Qwen3.8-27B hidden states to score query intent, response unsafety, and hallucination risk at every token. The probe adds less than 0.5% decode-time overhead and reports a 0.03% benign-response false-positive rate averaged across five datasets. It is supported through SGLang and vLLM integration branches, with training code available at inclusionAI/SingProbe.