Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Jun 15

Jun 15Mon
  1. ByteDance · new models on Hugging FaceOfficialAI score24

    Sa2VA-LLaVA-1.5-7B: ByteDance's SAM2-Grounded Segmentation and Chat Model

    AIByteDance has released Sa2VA-LLaVA-1.5-7B on Hugging Face, a model built on LLaVA-1.5-7B with a SAM2 grounding encoder that performs dense image and video referring segmentation alongside open-ended chat. The checkpoint is self-contained and loads with trust_remote_code=True without extra packages, and it is positioned as a LISA-comparable baseline within the Sa2VA family. Reported results include 80.3 cIoU on RefCOCO val and 54.8 J&F on MeViS (val_u).

Jun 13

Jun 13Sat
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score88

    Moonshot AI releases open-weight Kimi K3 with 2.8T parameters and 1M context

    AIMoonshot AI released Kimi K3 on Hugging Face as an open-weight, native multimodal agentic model with 2.8T total parameters and 104B activated parameters. It supports a 1-million-token context window and text and image input, with weights released under the Kimi K3 License. The model card reports benchmark results for coding, agentic, and vision tasks against several closed models, and recommends vLLM, SGLang, or TokenSpeed for inference.

    Why it matters: The release pairs open weights with a 2.8T-parameter MoE architecture and benchmark tables against several named closed models, useful for comparing frontier capability claims.

Jun 12

Jun 12Fri
  1. PaddlePaddleOfficialAI score41

    PaddleOCR releases PP-OCRv6 with models from 1.5M to 34.5M parameters

    AIPaddlePaddle has released PP-OCRv6, a new OCR model series in Tiny, Small, and Medium sizes at 1.5M, 7.7M, and 34.5M parameters. The models reportedly improve detection accuracy by 4.9% and recognition accuracy by 5.1% over PP-OCRv5, with up to 5.2× faster CPU inference via OpenVINO. The unified model supports 50 languages and new scenarios including PCB, CAD drawings, digital tubes, and dot-matrix text, under Apache 2.0.

    Image from @PaddlePaddle's post

Jun 11

Jun 11Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Jun 10

Jun 10Wed
  1. Xiaomi MiMoOfficialAI score67

    Xiaomi releases open-source MiMo Code V0.1 terminal coding assistant

    AIXiaomi MiMo has released MiMo Code V0.1, an open-source AI coding assistant for the terminal under the MIT license. It ships with MiMo V2.5, a multimodal model offered free for a limited time with a million-token context window. The tool automatically loads existing Claude Code skills, MCP servers and commands, and reuses API configuration, and it supports providers including Anthropic, OpenAI, DeepSeek, Kimi and GLM.

    Why it matters: The post specifies MiMo Code's Claude Code compatibility and MIT license, which bear directly on whether existing coding-agent setups can migrate without rework.

    Image from @XiaomiMiMo's post
  2. ByteDance · new models on Hugging FaceOfficialAI score52

    ByteDance open-sources Bernini-Diffusers for semantic video generation and editing

    AIByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.

  3. ByteDance · new models on Hugging FaceOfficialAI score34

    EvoQuality: ByteDance's self-evolving VLM for image quality assessment without human labels

    AIEvoQuality is a ByteDance vision-language model for no-reference image quality assessment that generates pseudo-ranking labels through pairwise majority voting and refines them with GRPO, requiring no human-annotated quality scores. On the paper's setting, it raised weighted-average PLCC from 0.615 to 0.770 and SRCC from 0.570 to 0.726 over its Qwen2.5-VL-7B backbone. The model is recommended for research and pre-production assessment, not as the sole criterion for high-stakes decisions.

  4. Xiaomi MiMoOfficialAI score82

    MiMo Code open-sources a terminal coding agent for long-horizon tasks

    AIXiaomi's MiMo team released MiMo Code, an MIT-licensed terminal coding agent built on OpenCode for long-horizon programming tasks. The design centers on three areas: Max Mode parallel sampling that generates five candidates per turn, Goal-based completion verification, and a memory system that checkpoints session state and rebuilds context. The article reports offline benchmark results and a double-blind A/B test with 1,213 pairs in which MiMo Code's win rate exceeded 65% beyond 200 execution steps.

    Why it matters: The article explains how MiMo Code handles long-horizon coding through computation, checkpointed memory, and cross-session evolution, useful for judging design tradeoffs in coding agents.

Jun 9

Jun 9Tue
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score52

    Z.ai releases SCAIL-2, an open-source end-to-end character animation model

    AIZ.ai released SCAIL-2, an open-source model that animates a reference character from a driving video without skeleton maps or inpainting masks. It also supports character replacement, multi-character scenes, and animal-driving, with 512p and 704p resolutions and inputs whose height and width are both divisible by 32.

Jun 8

Jun 8Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score70

    Cognition Introduces FrontierCode, a Benchmark for Mergeable Code Quality

    AICognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge. On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%. The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.

    Why it matters: The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.

  2. Xiaomi MiMoOfficialAI score62

    Xiaomi MiMo open-sources a 1T model running over 1,000 tps on 8 GPUs

    AIXiaomi MiMo and the TileRT team say a 1T model exceeds 1,000 tps on a single standard 8-GPU node using general-purpose GPUs. The speedup comes from FP4 quantization and DFlash, a block-masked parallel speculative decoding method that accepts more tokens per verification, with TileRT tailoring its compiler and kernels to these techniques. Open weights for the FP4 + DFlash checkpoint are available on Hugging Face.

    Why it matters: The post shows how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach over 1,000 tps on a standard GPU node.

  3. Xiaomi MiMoOfficialAI score62

    Xiaomi MiMo-V2.5-Pro UltraSpeed claims 1,000+ tokens/s on a 1T model

    AIXiaomi MiMo and TileRT released MiMo-V2.5-Pro-UltraSpeed, which the post says reaches output speeds above 1,000 tokens/s on a 1 trillion parameter MoE model. The post says this runs on a single standard 8-GPGPU node rather than wafer-scale or pure on-chip SRAM hardware. UltraSpeed access is application-based from Jun 8 to Jun 23 (PDT), and the UltraSpeed API costs 3x the standard price.

    Why it matters: The post specifies the hardware setup behind the claimed speed, which matters for judging whether the approach can be replicated on standard GPU nodes.

    Image from @XiaomiMiMo's post
  4. ByteDance · new models on Hugging FaceOfficialAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

  5. Xiaomi MiMo · new models on Hugging FaceOfficialAI score41

    Xiaomi releases MiMo-V2.5-Pro-FP4-DFlash, an FP4 model with block-diffusion decoding

    AIXiaomi MiMo has released MiMo-V2.5-Pro-FP4-DFlash, the FP4 backbone behind MiMo-V2.5-Pro-UltraSpeed, with MXFP4 quantization applied only to the MoE experts and a BF16 DFlash drafter for block-diffusion speculative decoding. The backbone has 1.02T total and 42B active parameters, and the drafter proposes blocks of up to 8 tokens per forward pass. The release is supported in SGLang, with example launch commands provided.

  6. Xiaomi MiMoOfficialAI score65

    Xiaomi MiMo-V2.5-Pro-UltraSpeed reaches 1000+ tokens/s on a 1T model

    AIXiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, reporting decode speeds above 1000 tokens/s on a 1-trillion-parameter model using a single standard 8-GPU node. The API is priced at 3x MiMo-V2.5-Pro and is available by application only from June 9 to June 23, 2026. The speedup relies on FP4 quantization of MoE Experts, DFlash speculative decoding with an average coding acceptance length of 6.30, and TileRT compute kernels.

    Why it matters: The post traces how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach 1000+ tokens/s on a single 8-GPU node, which is useful for teams weighing inference throughput.

Jun 5

Jun 5Fri

Jun 4

Jun 4Thu
  1. Cohere · new models on Hugging FaceOfficialAI score60

    Cohere releases North Mini Code 1.0, a 30B-A3B open-weights coding model

    AICohere and Cohere Labs released North Mini Code 1.0, an open-weights 30B-A3B mixture-of-experts model for code generation and agentic terminal tasks, under Apache 2.0. The model has 256K context and 64K max output, and is trained for tool use. Its benchmark table lists Terminal-Bench v2 at 36.0, SWE-Bench Verified at 67.6, and LiveCodeBench v6 at 70.3, below Qwen3.6 on several tasks.

    Why it matters: The card lists benchmark results against Qwen3.6, Gemma4, and other models, showing where North Mini Code trails on some coding and agentic tasks.

  2. Georgi GerganovXAI score44

    llama.cpp adds multi-GPU and tensor parallel support with NVIDIA

    AIMaintainers and NVIDIA engineers improved multi-GPU performance in ggml, the low-level engine behind llama.cpp, yielding significant gains on RTX systems. The work also lays groundwork for hardware-agnostic tensor parallelism in ggml. Details are in a technical blog from NVIDIA RTX Spark.

    Image from @ggerganov's post

Jun 3

Jun 3Wed
  1. PaddlePaddleOfficialAI score31

    Baidu CoBuddy, a free code-focused model, now live on Novita

    AIBaidu CoBuddy is now available for free on Novita AI, a code-focused model aimed at developers and AI agents. It offers a 131K context window, up to 65K output tokens, and native tool calling. The model is served through Novita's serverless API for high-throughput, low-latency inference.

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceOfficialAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. MiniMax · new models on Hugging FaceOfficialAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

  3. ByteDance · new models on Hugging FaceOfficialAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

Jun 1

Jun 1Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score50

    Cognition launches Devin Desktop, the next generation of Windsurf

    AICognition has announced Devin Desktop, the next generation of Windsurf, which makes the Agent Command Center the default IDE surface for managing local and cloud agents, PRs, and context. Spaces let related agents share context, and Agent Client Protocol (ACP) support lets any ACP-compatible agent run alongside Devin. The IDE remains fully backwards-compatible with Windsurf, including editor extensions, keybindings, LSPs, and terminal workflows.

  2. PaddlePaddleOfficialAI score36

    PaddleOCR and ERNIE Image now available as official Dify plugins

    AIPaddleOCR and ERNIE Image are now available as official Dify plugins, bringing document parsing and image generation into Dify's agent workflows. PaddleOCR, powered by PP-OCRv5, PP-StructureV3, and PaddleOCR-VL, turns images, scanned PDFs, and multilingual documents into structured data for chunking, vectorization, and RAG, with private or on-prem deployment supported. ERNIE Image offers free generation, a Turbo mode with 8-step inference, and an OpenAI-style API.

    Image from @PaddlePaddle's post

May 31

May 31Sun
  1. MiniMax BlogOfficialAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 30

May 30Sat
  1. Xiaomi MiMoOfficialAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 28

May 28Thu
  1. PaddlePaddleOfficialAI score36

    PaddleOCR-VL 1.6 released with 96.33% SOTA on OmniDocBench

    AIPaddlePaddle has released PaddleOCR-VL 1.6, which sets a new state-of-the-art score of 96.33% on OmniDocBench for text, formula, and table recognition. It ranks first on OmniDocBench v1.5 and Real5-OmniDocBench, with gains in table, classic text, rare character, seal, spotting, and chart recognition. The version is fully compatible with the v1.5 architecture, requiring no migration.

    Image from @PaddlePaddle's post

May 26

May 26Tue

May 24

May 24Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceOfficialAI score45

    Fun-ASR-Nano-2512-hf: Alibaba's Speech Recognition Model Gets Transformers Version

    AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.

May 22

May 22Fri
  1. ReflectionOfficialAI score38

    Reflection signs MOU with US Department of Energy for national labs

    AIReflection, an open-source AI lab, has signed a memorandum of understanding with the US Department of Energy to explore strategic collaborations under the Genesis Mission. Through the partnership, Reflection will provide open-weight models to the DOE's 17 National Laboratories, customizable on lab-specific scientific data and deployed into researcher workflows. The Genesis Mission focuses on advanced nuclear, fusion and grid modernization, quantum ecosystem growth, and national security AI.

May 20

May 20Wed
  1. Stability AIOfficialAI score62

    Stability AI releases Stable Audio 3.0 model family with open-weight music models

    AIStability AI released Stable Audio 3.0, a family of four audio models trained on fully licensed data. Three of them, Small SFX, Small and Medium, have open weights on Hugging Face, while Large is available through the Stability AI API and enterprise self-hosting. Outputs can be distributed and commercialized under the Stability AI Community License, and organizations with more than $1M in annual revenue can use the Enterprise License.

    Why it matters: The source specifies which models are open-weight, their licensing terms, and clip-length limits, which matters for anyone deciding whether to build on them.

  2. PaddlePaddleOfficialAI score36

    PaddleOCR 3.5 adds Hugging Face Transformers as inference backend

    AIPaddleOCR 3.5 now supports Hugging Face Transformers as an inference backend, letting users run PP-OCRv5 and PaddleOCR-VL 1.5 models directly within the Transformers ecosystem. Users can select it with engine="transformers" while keeping the same PaddleOCR pipeline, which the post says eases integration for RAG and Document AI applications.

May 19

May 19Tue
  1. Awni HannunXAI score46

    PyTorch models can now run on Apple silicon via MLX delegate

    AIAwni Hannun highlights a workflow to write PyTorch models, export them, and run them with MLX on Apple silicon. The linked PyTorch post says ExecuTorch now includes an MLX delegate that runs PyTorch models on Apple silicon GPUs, supporting LLMs, speech-to-text, and MoE models.

    Image from @awnihannun's post
  2. koray kavukcuogluXAI score72

    Google rolls out Gemini 3.5 Flash globally across consumer, developer, and enterprise platforms

    AIGoogle is rolling out Gemini 3.5 Flash globally for consumers in the Gemini app and Search AI Mode. It is also available to developers through the Gemini API, Google Antigravity, and Google AI Studio, and to businesses on the Gemini Enterprise Agent Platform.

    Why it matters: The post shows where each Gemini 3.5 Flash access path goes, from consumer apps to developer and enterprise platforms, which helps readers pick the right entry point.

May 17

May 17Sun