Skip to content

#Open source/Repo

Jul 27

Jul 27Mon
  1. Liquid AI BlogAI score49

    Liquid AI Releases LFM2.5-Encoders for Fast Long-Context Encoding on CPU

    Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.

  2. KimiAI score65

    Kimi K3 becomes available on Nebius Token Factory via API

    Kimi K3 is now available on Nebius Token Factory, which is named a Day 0 launch partner, through an OpenAI-compatible API and console. The quoted post says Artificial Analysis scores the open-weight model at 57 on its Intelligence Index, two points behind GPT-5.6 Sol (max), and lists up to 1M tokens of context.

    AIWhy it matters: The source names the cloud access route and an Artificial Analysis score of 57, letting readers compare Kimi K3 against GPT-5.6 Sol.

  3. KimiAI score31

    We've open-sourced MoonEP, our high-performance communication library for distributed MoE workloads. Built to make expert-parallel communication more efficient at scale, MoonEP helps reduce communication overhead in large MoE training and inference systems. Explore on GitHub: http://github.com/MoonshotAI/MoonEP

    We've open-sourced MoonEP, our high-performance communication library for distributed MoE workloads. Built to make expert-parallel communication more efficient at scale, MoonEP helps reduce communication overhead in large MoE training and inference systems. Explore on GitHub: http://github.com/MoonshotAI/MoonEP

  4. KimiAI score38

    We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, with fast snapshot, resume, and fork support for large-scale parallel agent workflows. Explore on GitHub: http://github.com/kvcache-ai/AgentEnv

    We've open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, with fast snapshot, resume, and fork support for large-scale parallel agent workflows. Explore on GitHub: http://github.com/kvcache-ai/AgentEnv

  5. KimiAI score40

    We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. It delivers 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on GitHub: http://github.com/MoonshotAI/FlashKDA

    We've open-sourced FlashKDA, our high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. It delivers 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20, and works as a drop-in backend for flash-linear-attention. Explore on GitHub: http://github.com/MoonshotAI/FlashKDA

  6. KimiAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    Moonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    AIWhy it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.

Jul 26

Jul 26Sun

Jul 24

Jul 24Fri

Jul 23

Jul 23Thu
  1. Matei ZahariaAI score36

    AI-based research and engineering is one of the main topics we're researching in the STAR Lab at Berkeley, and something I'm seeing used in industry more and more too. This is nice work from my grad students packaging multiple AI-based "autoresearch" algorithms into one API in the GEPA package so you can mix-and-match them to get the best results. These can be used for anything from writing prompts to designing agents to optimizing code.

    AI-based research and engineering is one of the main topics we're researching in the STAR Lab at Berkeley, and something I'm seeing used in industry more and more too. This is nice work from my grad students packaging multiple AI-based "autoresearch" algorithms into one API in the GEPA package so you can mix-and-match them to get the best results. These can be used for anything from writing prompts to designing agents to optimizing code.

  2. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-Base, a 122B deep research agent model

    BAAI has released AREX-Base, a 122B-total, 10B-activated Mixture-of-Experts deep research agent built on Qwen3.5-122B-A10B with a 262,144-token context. The model uses an inner research loop and an outer self-improvement loop, and the source reports it scoring 82.5 on BrowseComp and 85.4 on GAIA, under Apache 2.0.

    AIWhy it matters: The release pairs a 122B-parameter deep research agent with benchmark tables against frontier and open models, letting readers compare its search-agent results directly.

  3. BAAI · new models on Hugging FaceAI score47

    BAAI releases AREX-Turbo, a compact 4B recursive self-improving deep research agent

    BAAI's AREX-Turbo is a dense 4B deep research agent built on Qwen3.5-4B with a 262,144-token context length. It scores 70.7 on BrowseComp, 81.6 on GAIA and 40.6 on HLE with tools, versus 82.5, 85.4 and 52.4 for the 122B AREX-Base. The model is released under Apache License 2.0 and targets lower-cost research-agent deployment.

  4. Andrew NgAI score65

    Andrew Ng announces OpenWorker, an open-source agent that delivers finished work

    Andrew Ng and Rohit Prasad announced OpenWorker, an open-source agent that produces deliverables such as documents, Slack messages, and calendar updates across files and everyday tools. It checks in before consequential actions, runs on Mac with Windows support coming soon, and works with user-supplied API keys for models including GPT 5.6 Sol, Claude Fable, Gemini 3.6, open-weight models, or local Ollama models. Source code is available on GitHub, and the tool requires the user's own API key.

Jul 21

Jul 21Tue

Jul 20

Jul 20Mon

Jul 17

Jul 17Fri

Jul 16

Jul 16Thu
  1. Mistral AI · new models on Hugging FaceAI score46

    Mistral releases Shieldstral-1.0-3B, a policy-adaptive multimodal safety classifier

    Mistral AI released Shieldstral-1.0-3B, a 3B-parameter multimodal safety classifier that judges content against natural-language policies and outputs a continuous safety score. It moderates text, image, and text-plus-image content in a single forward pass and can be retargeted to new policies at inference time without retraining. The Apache 2.0 open-weight model is built on Ministral-3-3B-Base-2512 and trained on sequences up to 32k tokens.

Jul 15

Jul 15Wed
  1. John SchulmanAI score75

    Thinking Machines releases open-weights multimodal model Inkling

    Thinking Machines introduced Inkling, a model that reasons across text, image, and audio, and is making its full weights available. It is available today for fine-tuning on Tinker and can be tried in the Inkling Playground. John Schulman says pretraining began last winter and a small team added coding, reasoning, and agentic training starting in mid-January.

    AIWhy it matters: The post links an open-weights release to a stated training timeline, showing how a small team moved from pretraining to coding, reasoning, and agentic training.

  2. Liquid AI NewsletterAI score38

    Liquid AI Releases Antidoom and IFStruct to Fix Reasoning Loops and Schema Errors

    Liquid AI released Antidoom, an open-source method that retrains a single overtrained token to eliminate "doom loops" in small reasoning models. On LFM2.5-2.6B and Qwen3.5-4B, loop rates fell from 10.2% to 1.4% and from 22.9% to 1%, respectively. The company also released IFStruct, an open-source benchmark measuring whether model outputs satisfy a schema, where LFM2.5-350M rose from 21.10% to 44.90% after training.

Jul 14

Jul 14Tue
  1. Stability AIAI score16

    What Stability AI audio researcher CJ Carr (@cortexelation) is covering in the workshop: 🎵 Run the model locally (even on a MacBook Pro) 🎵 Train your own LoRA 🎵 Use Stable Audio with Ableton 🎵 Build your own audio tools on top of the model … and more

    What Stability AI audio researcher CJ Carr (@cortexelation) is covering in the workshop: 🎵 Run the model locally (even on a MacBook Pro) 🎵 Train your own LoRA 🎵 Use Stable Audio with Ableton 🎵 Build your own audio tools on top of the model … and more

Jul 12

Jul 12Sun
  1. ByteDance · new models on Hugging FaceAI score41

    ByteDance releases UniVR-34B-Planning for visual-space reasoning and planning

    ByteDance's UniVR-34B-Planning, built on Emu3.5 at 34B parameters, learns visual reasoning, physical dynamics, and long-term planning from visual demonstrations using a next-token objective and two-stage training on the VR-X dataset with VR-GRPO reinforcement learning. On the VR-X benchmark it scores 58.2 overall, up 18.4 points from the Emu3.5 34B baseline of 39.8. The Planning checkpoint is available on Hugging Face under CC BY 4.0, alongside a General checkpoint.

Jul 9

Jul 9Thu

Jul 8

Jul 8Wed
  1. Aman SangerAI score40

    Very excited for this model. It’s an enormous improvement over composer 2.5 and trained entirely from scratch. It’s been a pleasure working with the SpaceXAI team on it. Even more excited by the slope of the effort and the models to come.

    Very excited for this model. It’s an enormous improvement over composer 2.5 and trained entirely from scratch. It’s been a pleasure working with the SpaceXAI team on it. Even more excited by the slope of the effort and the models to come.

  2. Georgi GerganovAI score46

    llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort! https://github.com/ggml-org/llama.cpp/pull/22105

    llama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort! https://github.com/ggml-org/llama.cpp/pull/22105

Jul 5

Jul 5Sun
  1. ARC PrizeAI score47

    ARC Prize Awards First ARC-AGI-3 Milestone Prize to Tufa Labs' Open-Source Agent

    Tufa Labs won the first $37.5K ARC-AGI-3 milestone prize with "The Duck," a small open-source LLM that plays the games by writing and running Python in a live REPL. Reki placed second with a vision-language agent using Gemma-4-31B, and md Boktiar Mahbub Murad placed third with the "forge" framework. The second and final milestone prize ends September 30.

Jul 3

Jul 3Fri

Jul 1

Jul 1Wed
  1. Stability AIAI score38

    Most AI audio models have never heard a maqam. Team Motif fine-tuned Stable Audio 3.0 on Arabic maqam, built an Ableton plugin for microtonal style transfer, and won our Stable Audio 3.0 Challenge at Music Hackspace running locally on device. Watch Jad Al Masri break it down 👇

    Most AI audio models have never heard a maqam. Team Motif fine-tuned Stable Audio 3.0 on Arabic maqam, built an Ableton plugin for microtonal style transfer, and won our Stable Audio 3.0 Challenge at Music Hackspace running locally on device. Watch Jad Al Masri break it down 👇

Jun 30

Jun 30Tue
  1. Soumith ChintalaAI score32

    Bridgewater, one of the worlds largest hedge funds, a Tinker customer talks through how they've carefully fine-tuned a model focused on what makes interesting financial news. Their fine-tuned model is more effective and cheaper than any frontier model.

    Bridgewater, one of the worlds largest hedge funds, a Tinker customer talks through how they've carefully fine-tuned a model focused on what makes interesting financial news. Their fine-tuned model is more effective and cheaper than any frontier model.

Jun 28

Jun 28Sun