Skip to contentSkip to stories

Updated

#Open-source ecosystem

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 10

Apr 10Fri

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

Apr 7

Apr 7Tue

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score26

    LaSER-Qwen3-8B: Alibaba NLP's 8B dense retriever with latent reasoning released on Hugging Face

    AIAlibaba NLP released LaSER-Qwen3-8B, an 8B-parameter dense retriever built on Qwen/Qwen3-8B that internalizes explicit reasoning into latent space through continuous latent thinking tokens. The model scores 29.3 nDCG@10 on the BRIGHT benchmark, ahead of the rewrite-then-retrieve pipeline's 28.1, and carries a 4096-dimension embedding with an 8192-token maximum sequence length. It is licensed under MIT and adds about 1.7× latency over standard single-pass dense retrievers.

Mar 26

Mar 26Thu
  1. Guillaume LampleAI score62

    Mistral releases Voxtral TTS, its first open-weight speech model

    AIMistral's Voxtral TTS is its first speech model, presented as an open-weight text-to-speech model that reportedly delivers SOTA performance at significantly lower cost with very low latency. It combines autoregressive generation of semantic speech tokens with flow-matching for acoustic tokens, and a technical report on its training methodology is being released.

Mar 24

Mar 24Tue
  1. ARC PrizeAI score70

    ARC Prize announces ARC-AGI-3, an interactive benchmark for frontier agents

    AIARC Prize has released ARC-AGI-3, a set of hundreds of interactive, turn-based environments with thousands of game-style levels, with no instructions or stated goals. Humans score 100% while frontier AI scores 0.51%. ARC Prize 2026 offers over $2 million in prizes for open-source solutions to ARC-AGI-2 and ARC-AGI-3.

    Why it matters: The benchmark's human versus frontier AI gap and its interactive design show how agent evaluation is shifting from instruction-following toward exploration and adaptation.

  2. Jim FanAI score62

    Jim Fan warns that compromised LiteLLM package shows risks for AI agents

    AIJim Fan reposted a report that LiteLLM PyPI release 1.82.8 was compromised and contained a litellm_init.pth file that sends credentials to a remote server and self-replicates. He argues agents make this worse, since files like skills, configs, or PDFs read into context could spread malicious instructions. He concludes that agentic frameworks need guardrails and audited tooling.

Mar 20

Mar 20Fri
  1. Aman SangerAI score55

    Cursor's Composer 2 is built on Kimi k2.5 base model with added training

    AIAman Sanger says Cursor's team evaluated many base models on perplexity-based evals and found Kimi k2.5 the strongest. Composer 2 was then built with continued pretraining and a 4x scale-up of high-compute RL, with Fireworks providing inference and RL samplers. The author admits Cursor should have named the Kimi base in its launch blog and says it will do so for the next model.

Mar 19

Mar 19Thu
  1. Tri DaoAI score52

    Tri Dao Says Nonlinear RNNs Differ From Attention and Linear SSMs

    AITri Dao says nonlinear RNNs seem to do something genuinely different from attention and linear RNNs or SSMs. He reports they already perform well with the right parametrization, and adding just one nonlinear RNN layer substantially improves a transformer-Mamba/DeltaNet hybrid. The post quotes the M²RNN paper, which introduces non-linear RNNs with matrix-valued states for language modeling, with links to the paper, code, and models.

Mar 17

Mar 17Tue
  1. MiniMax BlogAI score63

    MiniMax M2.7 takes part in its own model and harness evolution

    AIMiniMax says M2.7 is its first model to deeply participate in its own evolution, building agent harnesses and running reinforcement learning experiment workflows. The post reports 56.22% on SWE-Pro, 55.6% on VIBE-Pro, 57.0% on Terminal Bench 2, and a 30% improvement on an internal evaluation set after more than 100 autonomous optimization rounds. It also states that M2.7 handles 30%-50% of its research team's workflow, though human researchers still make critical decisions.

    Why it matters: The post ties M2.7's self-evolution claims to specific benchmark numbers and workflow details, helping readers judge how much of the iteration loop is autonomous.

  2. Apple · new models on Hugging FaceAI score46

    Apple releases SimpleSD-4B-instruct, a self-distilled Qwen code model

    AIApple has released SimpleSD-4B-instruct on Hugging Face, a research checkpoint fine-tuned from Qwen3-4B-Instruct-2507 on its own sampled outputs to improve code generation. On LiveCodeBench, the model scores 41.5% pass@1 on LCBv6, up from the base model's 34.0%, and 45.7% pass@1 on LCBv5, up from 34.3%. The model is released under the Apple Machine Learning Research Model License and is intended for reproducibility rather than as an optimized Qwen release.

Mar 12

Mar 12Thu
  1. Intern Large ModelsAI score47

    InternVL-U: Open-Source 4B Unified Model for Reasoning, Generation, and Editing

    AIInternVL-U is a lightweight 4B unified multimodal model that combines reasoning, generation, and editing in one framework, according to Intern Large Models. The post says it uses unified contextual modeling, modality-specific modular design, and decoupled visual representations to balance performance and efficiency. It reportedly outperforms unified baselines more than 3× its size on text rendering, scientific reasoning, and spatially grounded generation and editing, and is open-source on GitHub and Hugging Face.

Mar 4

Mar 4Wed

Mar 2

Mar 2Mon

Feb 20

Feb 20Fri
  1. Jim FanAI score75

    DreamDojo: Open-source world model trained on 44K hours of human video

    AIJim Fan announced DreamDojo, an open-source interactive world model that takes robot motor controls and generates future frames in pixels. It is pre-trained on 44K hours of human egocentric video using latent actions, then post-trained onto specific robot hardware, and a real-time version runs at 10 FPS for live teleoperation, policy evaluation, and model-based planning. The author reports a +17% real-world success gain on a fruit packing task, and weights, code, datasets, and the whitepaper are released.

Feb 19

Feb 19Thu

Feb 4

Feb 4Wed
  1. Guillaume LampleAI score62

    Mistral's Voxtral Realtime streams speech with sub-200ms latency and open weights

    AIVoxtral Realtime is a natively streaming speech model for voice agents and live applications, with latency configurable down to sub-200ms. At 480ms it stays within 1-2% WER of the offline model, and the weights are released under Apache 2.0. The attached FLEURS chart compares word error rates across latency settings for ten languages, including Chinese.

  2. Guillaume LampleAI score62

    Mistral releases Voxtral 2 transcription models with real-time option

    AIMistral announces Voxtral 2 with two transcription models: Voxtral Realtime, released under an Apache 2 license with latency configurable to sub-200 ms, and Voxtral Mini Transcribe 2, which adds speaker diarization, word-level timestamps, and context biasing. The models support 13 languages and are available through the Mistral API, which the post describes as one of the most cost-effective transcription APIs on the market. The attached chart shows word error rates on FLEURS across Italian, Spanish, English, German, Portuguese, French, Russian, Dutch, and Chinese at several latency settings.

Jan 27

Jan 27Tue
  1. Tim DettmersAI score72

    Tim Dettmers Details How SERA Built an Open Coding Agent on 32 GPUs

    AIAi2's Open Coding Agents family, with SERA as its first release, was built by Tim Dettmers and collaborators on 32 GPUs. The method generates synthetic bug trajectories with soft verification, comparing patches by line overlap instead of running tests. The post reports that a 32B model fine-tuned on about 7,000 trajectories for one private repository matched its GLM 4.5-Air teacher, and that the baseline costs $500 to run.

Jan 26

Jan 26Mon
  1. BAAIAI score40

    BAAI RoboBrain 2.5 targets robot spatial and temporal reasoning gaps

    AIBAAI released RoboBrain 2.5, an embodied AI model that turns 2D scene understanding into actionable 3D trajectories and provides dense temporal value estimates for real-time progress feedback on long-horizon tasks. The post says it achieves SOTA across multiple spatial and temporal reasoning benchmarks, though it names no specific scores. Project page, paper, GitHub code, and model weights are linked.

Jan 22

Jan 22Thu
  1. BAAIAI score38

    BAAI releases RoboCOIN, a large bimanual robot manipulation dataset

    AIBAAI's RoboCOIN is a bimanual robot dataset with more than 180,000 trajectories across 421 tasks, collected from 15 robot platforms in 16 real-world scenarios. Its three-tier annotations at trajectory, segment, and frame levels help robots learn both what to do and how to do it. The post says integrating these annotations raised success rates on complex tasks by up to 50% for models such as π₀.

Jan 14

Jan 14Wed
  1. Black Forest Labs · new models on Hugging FaceAI score54

    Black Forest Labs releases FLUX.2 [klein] 4B Base on Hugging Face

    AIBlack Forest Labs has published FLUX.2 [klein] 4B Base, a 4 billion parameter text-to-image model that also supports multi-reference editing. The model is undistilled, is released with open weights under Apache 2.0, and is described as fitting in about 13GB VRAM on cards such as the RTX 3090 or 4070, with reference code available in its GitHub repository and support in ComfyUI and Diffusers.

  2. Black Forest Labs · new models on Hugging FaceAI score46

    FLUX.2 [klein] 9B Base Released on Hugging Face as Undistilled Open-Weight Model

    AIBlack Forest Labs has released FLUX.2 [klein] 9B Base, a 9 billion parameter undistilled rectified flow transformer with open weights for text-to-image generation and multi-reference editing. The model is intended for fine-tuning, LoRA training, and research, and fits in about 29GB VRAM on NVIDIA RTX 4090-class GPUs. A reference implementation is available on GitHub, and the model works with ComfyUI and Diffusers.

Dec 22, 2025

Dec 22, 2025Mon
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score58

    Alibaba's FunAudioLLM releases Fun-Audio-Chat-8B for low-latency voice interaction

    AIFunAudioLLM has released Fun-Audio-Chat-8B, a roughly 8B-parameter large audio language model for natural, low-latency voice interaction, under Apache 2.0. It uses Dual-Resolution Speech Representations with a 5Hz frame rate, which the source says reduces GPU hours by nearly 50%, and it supports English and Chinese.

Dec 14, 2025

Dec 14, 2025Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score38

    Alibaba Releases Fun-ASR-MLT-Nano-2512, an 800M Multilingual Speech Recognition Model

    AIAlibaba's FunAudioLLM released Fun-ASR-MLT-Nano-2512, an 800M-parameter multilingual speech recognition checkpoint on Hugging Face that supports 31 languages, with emphasis on East and Southeast Asian languages. It is trained on hundreds of thousands of hours of speech and is available through the FunASR toolkit. The source's benchmark tables cover the Fun-ASR family rather than this checkpoint, so no checkpoint-specific accuracy figures are reported.

  2. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score36

    Fun-ASR-Nano-2512 Speech Recognition Model Released by Tongyi Lab on Hugging Face

    AITongyi Lab has released Fun-ASR-Nano-2512, an end-to-end speech recognition large model trained on tens of millions of hours of real speech, supporting low-latency real-time transcription across 31 languages. The model, which has 800M parameters, targets industry use such as education and finance and claims 93% accuracy in far-field, high-noise conditions. It is available on Hugging Face and works with the FunASR toolkit.

Dec 12, 2025

Dec 12, 2025Fri

Dec 10, 2025

Dec 10, 2025Wed
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score42

    Fun-CosyVoice3-0.5B-2512 Released as Open-Source Multilingual Text-to-Speech Model

    AIAlibaba's FunAudioLLM has released Fun-CosyVoice3-0.5B-2512, a 0.5B-parameter LLM-based text-to-speech model on Hugging Face, with an RL variant also published. The model supports zero-shot voice cloning across 9 languages and 18+ Chinese dialects and accents, with streaming output at latency as low as 150ms. On the source's test-en benchmark, it reports a 2.24% WER and 71.8% speaker similarity, and the RL version reports 1.68% WER.

  2. Andrej KarpathyAI score34

    Karpathy Uses GPT-5.1 Thinking to Grade December 2015 Hacker News Discussions in Hindsight

    AIAndrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

Dec 2, 2025

Dec 2, 2025Tue
  1. Apple · new models on Hugging FaceAI score36

    Apple releases CLaRa-7B-E2E, an end-to-end RAG model with 16x and 128x compression

    AIApple's CLaRa-7B-E2E is a fully end-to-end unified RAG model that jointly optimizes retrieval and generation, with 16x and 128x document compression. It is trained with end-to-end finetuning using differentiable top-k retrieval and a unified language-modeling objective. The model is available on Hugging Face with example end-to-end inference code.

Oct 30, 2025

Oct 30, 2025Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score72

    Moonshot AI releases Kimi Linear 48B-A3B hybrid attention models on Hugging Face

    AIMoonshot AI has released Kimi-Linear-Base and Kimi-Linear-Instruct, both 48B total and 3B activated parameters with a 1M context length, on Hugging Face. The models use Kimi Delta Attention in a 3:1 hybrid ratio with global MLA, cutting KV cache by up to 75% and boosting decoding throughput by up to 6x at 1M tokens. The KDA kernel is open-sourced in FLA, and the checkpoints were trained on 5.7T tokens.

    Why it matters: The model card gives concrete throughput and KV cache figures for a hybrid attention design, which helps readers weigh its long-context tradeoffs against full attention.

May 18, 2025

May 18, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score22

    Cognition Revives Devin Open Source Initiative With $500 Credits for Projects

    AICognition is bringing back its Devin Open Source Initiative, offering $500 in Devin ACU credits to open-source GitHub projects with over 100 forks. Projects below that threshold will still be considered. Eligible projects must have an OSI-approved license and be actively maintained, and maintainers can apply through a linked form.

Jan 13, 2025

Jan 13, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score38

    Crossmint Uses Devin to Scale Open-Source Development of GOAT SDK

    AICrossmint said Devin became its top contributor to the open-source GOAT SDK during an initial trial, merging 8 pull requests versus 4 for the next contributor. Examples included a DEXScreener plugin built from a documentation URL and a harder Sui blockchain integration that needed three rounds of feedback and about an hour of human involvement. The company said its results depended on proper training, clear task context, and planned validation, not on treating Devin as superhuman.

Dec 11, 2024

Dec 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Devin Open Source Initiative Gives Maintainers 500 Free ACUs for Repo Work

    AICognition is launching the Devin Open Source Initiative, giving selected open source maintainers 500 free ACUs on a Devin Teams plan as part of Devin's general availability launch. The post shows Devin contributing pull requests to projects including Anthropic's MCP Inspector, Dagger, and nanoGPT, with maintainers still reviewing the results. Devin's GitHub integration forwards PR comments and CI checks to help refine changes, though the company warns a human should still verify final quality.