Skip to contentSkip to stories

Updated

#Open-source ecosystem

Jul 1

Jul 1Wed
  1. Mistral AI · new models on Hugging FaceAI score54

    Mistral AI releases Leanstral 1.5, an open-source Lean 4 code agent model

    AIMistral AI released Leanstral 1.5 on Hugging Face as an open-source code agent model for Lean 4 proof assistant tasks. The model uses 119B total parameters with 6.5B activated per token, a 256k context length, and accepts text and image input. The source gives setup paths through Mistral Vibe and a local vLLM server, with recommended settings of temperature 1.0 and reasoning effort set to high for complex prompts. The model is licensed under Apache 2.0.

  2. Jim FanAI score51

    Jim Fan introduces ASPIRE, a self-evolving robot skills library for continual learning

    AIJim Fan announces ASPIRE, a system where coding agents use multimodal sensory traces from simulation and real robots to run evolutionary search over control programs and add the results to a growing skills library. The post claims up to a roughly 10x reduction in transfer learning tokens for sim-to-real and single-arm to bimanual transfer, and says the full stack will be open-sourced.

Jun 30

Jun 30Tue
  1. Jim FanAI score60

    ASPIRE lets robots build an evolving skills library that transfers across tasks

    AIJim Fan introduces ASPIRE, a system in which coding agents observe multimodal sensory traces and run evolutionary search over control programs to distill skills into a growing library. The post says ASPIRE shares know-how rather than pixels or weights across the sim-to-real gap, reducing transfer learning tokens by up to about 10x. The author also says the full stack will be open-sourced and provides a gallery of 150+ tasks and 90+ skills.

  2. Xiaomi MiMoAI score22

    Xiaomi MiMo praised as developers build on open-weights models

    AIXiaomi MiMo's account celebrated growing developer adoption of its open-weights models, crediting Cline for building on MiMo. Cline's linked post announced a $9.99/month subscription offering 2-5x discounted access to GLM-5.2 and other open-weight models including DeepSeek, Kimi, MiniMax, MiMo, and Qwen, with a $1.99 promo for sign-ups via npm i -g cline.

Jun 28

Jun 28Sun
  1. PaddlePaddleAI score46

    PaddlePaddle announces Unlimited-OCR now runs in vLLM

    AIUnlimited-OCR, Baidu's long-context OCR model, now runs in vLLM, with a recipe provided for developers to try it. The background post says it parses entire books in one pass using Reference Sliding Window Attention (R-SWA), which keeps the KV cache fixed during decoding, and claims 35% faster throughput than DeepSeek-OCR at 6K output tokens.

Jun 27

Jun 27Sat
  1. PaddlePaddleAI score36

    PaddleFormers 1.2 adds DeepSeek-V4 training with 128K+ context support

    AIPaddleFormers 1.2 is released with support for training DeepSeek-V4 and 128K+ long-context training. The update adds Context Parallel, Packing, Document Mask Attention, and the Muon optimizer, plus ultra-fused mHC, CSA, and HCA operators, DeepEP/HybridEP communication, and lossless FP8 training with AutoSubbatch memory balancing. The project is presented as fully open-source and is available on GitHub.

  2. Ahead of AI (Sebastian Raschka)AI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.

Jun 26

Jun 26Fri
  1. PaddlePaddleAI score32

    PP-OCRv6 Ep.4 benchmarks show 3.9x CPU speedup and 0.13s A100 OCR

    AIPaddlePaddle's PP-OCRv6 Tech Deep Dive Ep.4 benchmarks the OCR models across A100, V100, Intel Xeon CPU, and Apple M4 setups. PP-OCRv6_tiny processes an image in 0.13s on A100, while PP-OCRv6_tiny with OpenVINO runs 3.9x faster than PP-OCRv5_mobile on Intel CPU. The post recommends Medium for high-concurrency APIs, Small for CPU document systems, Tiny for mobile or embedded devices, and Medium or Small for multilingual business use.

Jun 25

Jun 25Thu

Jun 23

Jun 23Tue
  1. PaddlePaddleAI score38

    PP-OCRv6 lightweight OCR model challenges large VLMs with 34.5M params

    AIPaddlePaddle introduced PP-OCRv6, a lightweight OCR architecture built on the LCNetV4 backbone, in the first episode of its tech deep dive series. The post says PP-OCRv6_medium reaches 86.2% detection Hmean and 83.2% recognition accuracy, surpassing PP-OCRv5_server while running faster. Three model specs—Tiny, Small, and Medium—target edge CPU devices, balanced deployment, and industrial high-accuracy pipelines.

Jun 20

Jun 20Sat

Jun 19

Jun 19Fri
  1. Andrew NgAI score72

    Andrew Ng says Anthropic and U.S. export controls on Fable expose AI access risks

    AIAndrew Ng argues that Anthropic's restrictions on building competing LLMs and a U.S. Commerce Department license requirement for foreign nationals led Anthropic to disable Fable access worldwide. He says this shows governments and providers can quickly cut off access to frontier AI, which may push nations and businesses toward sovereignty efforts and open-source alternatives, though training frontier models remains difficult.

Jun 16

Jun 16Tue
  1. Arthur MenschAI score44

    Mistral says its upcoming models will all be open-weight

    AIMistral states that this model and upcoming ones will be open-weight. The company argues that open weights are critical for customer confidence and for research and developer communities. It contends that systems reachable only through someone else's interface cannot be owned, inspected, audited, or improved, especially if data recording can no longer be turned off.

Jun 15

Jun 15Mon
  1. Zed BlogAI score38

    Zed Guild Cohort 1 Ends with 148 Merged Pull Requests from 33 Contributors

    AIZed's 12-week Guild program, its first cohort run this spring, had 33 active contributors merge 148 pull requests into the open-source editor. The top contributor, feitreim, merged 23 PRs, including fixes for Vim mode screen flickering and terminal ANSI rendering, and won a trip to Rust Week in Utrecht. Zed plans to organize Cohort 2 work into tighter groups around specific parts of the codebase.

  2. ByteDance · new models on Hugging FaceAI score24

    Sa2VA-LLaVA-1.5-7B: ByteDance's SAM2-Grounded Segmentation and Chat Model

    AIByteDance has released Sa2VA-LLaVA-1.5-7B on Hugging Face, a model built on LLaVA-1.5-7B with a SAM2 grounding encoder that performs dense image and video referring segmentation alongside open-ended chat. The checkpoint is self-contained and loads with trust_remote_code=True without extra packages, and it is positioned as a LISA-comparable baseline within the Sa2VA family. Reported results include 80.3 cIoU on RefCOCO val and 54.8 J&F on MeViS (val_u).

Jun 12

Jun 12Fri
  1. PaddlePaddleAI score41

    PaddleOCR releases PP-OCRv6 with models from 1.5M to 34.5M parameters

    AIPaddlePaddle has released PP-OCRv6, a new OCR model series in Tiny, Small, and Medium sizes at 1.5M, 7.7M, and 34.5M parameters. The models reportedly improve detection accuracy by 4.9% and recognition accuracy by 5.1% over PP-OCRv5, with up to 5.2× faster CPU inference via OpenVINO. The unified model supports 50 languages and new scenarios including PCB, CAD drawings, digital tubes, and dot-matrix text, under Apache 2.0.

Jun 11

Jun 11Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Jun 10

Jun 10Wed
  1. Xiaomi MiMoAI score82

    MiMo Code open-sources a terminal coding agent for long-horizon tasks

    AIXiaomi's MiMo team released MiMo Code, an MIT-licensed terminal coding agent built on OpenCode for long-horizon programming tasks. The design centers on three areas: Max Mode parallel sampling that generates five candidates per turn, Goal-based completion verification, and a memory system that checkpoints session state and rebuilds context. The article reports offline benchmark results and a double-blind A/B test with 1,213 pairs in which MiMo Code's win rate exceeded 65% beyond 200 execution steps.

    Why it matters: The article explains how MiMo Code handles long-horizon coding through computation, checkpointed memory, and cross-session evolution, useful for judging design tradeoffs in coding agents.

Jun 8

Jun 8Mon
  1. Cognition Blog (Devin, Windsurf)AI score70

    Cognition Introduces FrontierCode, a Benchmark for Mergeable Code Quality

    AICognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge. On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%. The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.

    Why it matters: The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.

  2. ByteDance · new models on Hugging FaceAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

Jun 4

Jun 4Thu

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. ByteDance · new models on Hugging FaceAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

Jun 1

Jun 1Mon
  1. Cognition Blog (Devin, Windsurf)AI score50

    Cognition launches Devin Desktop, the next generation of Windsurf

    AICognition has announced Devin Desktop, the next generation of Windsurf, which makes the Agent Command Center the default IDE surface for managing local and cloud agents, PRs, and context. Spaces let related agents share context, and Agent Client Protocol (ACP) support lets any ACP-compatible agent run alongside Devin. The IDE remains fully backwards-compatible with Windsurf, including editor extensions, keybindings, LSPs, and terminal workflows.

May 31

May 31Sun
  1. MiniMax BlogAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 30

May 30Sat
  1. Xiaomi MiMoAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 28

May 28Thu
  1. PaddlePaddleAI score36

    PaddleOCR-VL 1.6 released with 96.33% SOTA on OmniDocBench

    AIPaddlePaddle has released PaddleOCR-VL 1.6, which sets a new state-of-the-art score of 96.33% on OmniDocBench for text, formula, and table recognition. It ranks first on OmniDocBench v1.5 and Real5-OmniDocBench, with gains in table, classic text, rare character, seal, spotting, and chart recognition. The version is fully compatible with the v1.5 architecture, requiring no migration.

May 24

May 24Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score45

    Fun-ASR-Nano-2512-hf: Alibaba's Speech Recognition Model Gets Transformers Version

    AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.

May 17

May 17Sun

May 15

May 15Fri
  1. Intern Large ModelsAI score55

    Intern-S2-Preview: 35B Open Scientific Multimodal Model Released

    AIShanghai AI Laboratory's Intern Large Models introduces Intern-S2-Preview, a 35B scientific multimodal foundation model, and says it matches the trillion-scale Intern-S1-Pro on core scientific tasks. The post says it is the first open-source model with material crystal structure generation and strong general capabilities, with shared-weight MTP plus KL loss improving acceptance rate and speed. It is already supported by vLLM and SGLang, with weights on Hugging Face and ModelScope.

Apr 30

Apr 30Thu
  1. ARC PrizeAI score44

    GPT-5.5 and Opus 4.7 Fail ARC-AGI-3 Tasks Through Flawed World Models

    AIOpenAI's GPT-5.5 scored 0.43% and Anthropic's Opus 4.7 scored 0.18% on ARC-AGI-3, a set of 135 novel environments, according to ARC Prize's replay analysis of 160 runs. The analysis found three recurring failure modes: models perceived local action effects but failed to build global rules, mapped unfamiliar games onto known ones, and sometimes beat a level without learning the underlying mechanic. ARC Prize is open-sourcing its analysis package.

Apr 29

Apr 29Wed

Apr 28

Apr 28Tue

Apr 27

Apr 27Mon
  1. Xiaomi MiMo · new models on Hugging FaceAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

Apr 24

Apr 24Fri

Apr 22

Apr 22Wed