Skip to contentSkip to stories

Updated

#Open source/Repo

Sep 14

Sep 14Mon
  1. NVIDIA · new models on Hugging FaceAI score40

    NVIDIA releases FoundationStereo small stereo depth model on Hugging Face

    AINVIDIA Research released FoundationStereo-small, a zero-shot stereo depth model that takes an RGB stereo pair and outputs a disparity map, on Hugging Face. The model has about 6.3×10^7 parameters and ships as ONNX files at fixed 576x960 and 320x736 resolutions, with TensorRT and ONNX runtime support. It is licensed under the NVIDIA Open Model License and is ready for commercial use.

  2. NVIDIA · new models on Hugging FaceAI score36

    NVIDIA's FoundationPose estimates 6-DoF object pose without fine-tuning given a CAD model

    AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.

  3. vLLM BlogAI score53

    Novita AI open-sources Chord, a W4A16 MoE kernel for Kimi K2.x on vLLM

    AINovita AI has open-sourced Chord, a W4A16 MoE CUDA operator with BF16 activations, INT4 weights and group-32 scales, built for Kimi K2.x serving shapes. Measured per layer against public Humming, it reports 1.11–1.20x on H200 EP8 prefill, 1.17–1.33x on H200 TP8 serving, and 1.81–2.15x on B300 EP8 decode against an untuned Humming default. Integration of the grouped operators with vLLM's Humming backend is still a work in progress.

  4. LlamaIndexAI score29

    LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines

    AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.

  5. Tencent · new models on Hugging FaceAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

  6. MiniMaxAI score36

    MiniMax H3 community projects speed up open-source video generation

    AIMiniMax highlighted open-source community progress on its H3 video generation model, which it built with native stereo audio and multimodal reference control. Recent highlights include FastH3's 4-step distillation running on DGX Spark and Apple Silicon, and NVIDIA's Sol-H3 generating 15 seconds of 768p video with audio in 6.6 seconds on 8×B300 in a warm-inference benchmark. Other releases include VDN's faster-inference attention work with code and weights, and 8-step Acc-LoRAs from Alibaba PAI, with LightX2V offering 4- and 8-step Turbo LoRAs.

Sep 13

Sep 13Sun
  1. Qwen · new models on Hugging FaceAI score67

    Qwen releases open-source Qwen-Image-2.1 for generation and editing

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B parameters in its visual generation component. The model can generate regular or transparent RGBA images, supports up to 10 reference images for editing, and is licensed under the Qwen Research License Agreement.

    Why it matters: The source specifies the 7B visual component, transparent RGBA output, and up to 10 reference images, which helps readers judge its fit for generation and editing workflows.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceAI score38

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.8-27B

    AIinclusionAI has released Qwen3.8-27B-singprobe, a 10.1M-parameter intrinsic streaming guardrail that reuses Qwen3.8-27B hidden states to score query intent, response unsafety, and hallucination risk at every token. The probe adds less than 0.5% decode-time overhead and reports a 0.03% benign-response false-positive rate averaged across five datasets. It is supported through SGLang and vLLM integration branches, with training code available at inclusionAI/SingProbe.

  3. inclusionAI (Ant Ling) · new models on Hugging FaceAI score40

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.5-397B-A17B

    AIinclusionAI has released Qwen3.5-397B-A17B-singprobe, an intrinsic streaming guardrail built on Qwen/Qwen3.5-397B-A17B that scores query intent, response unsafety, and hallucination risk at every generated token using the base model's hidden states. The probe has 8.13M parameters, taps layers 18, 38, and 58, and adds less than 0.5% decode-time overhead. Training code is available at inclusionAI/SingProbe, and the probe runs through SGLang or vLLM integration branches.

  4. inclusionAI (Ant Ling) · new models on Hugging FaceAI score42

    SingProbe: inclusionAI releases streaming safety probe for gpt-oss-120b

    AIinclusionAI released SingProbe, a 5.8M-parameter intrinsic guardrail built on openai/gpt-oss-120b that scores query intent, response unsafety, and hallucination risk at every token. It reuses the base model's hidden states, adding less than 0.5% decode-time overhead, and reports a 0.06% benign-response false-positive rate. The probe is available on Hugging Face and supported through SGLang and vLLM integrations.

Sep 11

Sep 11Fri
  1. Baseten BlogAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

Sep 10

Sep 10Thu
  1. Tencent HunyuanAI score60

    Tencent Hunyuan releases open-source AuK audio model for speech generation and editing

    AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.

Sep 9

Sep 9Wed
  1. BAAI · new models on Hugging FaceAI score24

    BAAI open-sources EPT, UniPath, and MiSI AIDD molecular and crystal modeling resources

    AIBAAI released open-source resources for three AIDD projects on Hugging Face: EPT, an equivariant pretrained transformer for unified 3D molecular representation learning, and UniPath, a learnable-time flow matching method for crystal structure and energy prediction. The repository mirrors their GitHub source code and READMEs, with setup, preprocessing, training, and evaluation documentation. The MiSI benchmark is released separately on Hugging Face.

  2. DeepSeek · new models on Hugging FaceAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

Sep 8

Sep 8Tue

Sep 7

Sep 7Mon
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.

Sep 6

Sep 6Sun
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score32

    MiniCPM5-2B-DSpark draft model released for speculative decoding with MiniCPM5-2B

    AIOpenBMB released MiniCPM5-2B-DSpark, a 323,776,001-parameter DSpark draft checkpoint with five layers that proposes seven draft tokens per forward pass for the MiniCPM5-2B target model. The model, trained on 7,054,154,509 tokens with an average acceptance length of 5.5174 at T=0 and 4.0514 at T=1.0, is served through SGLang with DSPARK speculative decoding. It is released in BF16 under the Apache-2.0 License.

Sep 4

Sep 4Fri
  1. Tencent · new models on Hugging FaceAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  2. Tencent · new models on Hugging FaceAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu