Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 13

Sep 13Sun
  1. Sebastian RaschkaXAI score35

    Raschka's Reasoning from Scratch Round 3 Builds a Math Verifier

    AISebastian Raschka's third "Reasoning from Scratch" video covers building a math verifier for evaluating language models and for later reinforcement learning with verifiable rewards (RLVR) training. The walkthrough covers extracting final answers from boxed outputs, normalizing them, checking mathematical equivalence, and running evaluation on the MATH-500 dataset.

    Video from @rasbt's post
  2. Satya NadellaXAI score36

    Nadella outlines principles for superintelligence, open ecosystems, and enterprise control

    AISatya Nadella says any pursuit of superintelligence must help humanity and remain under human control, and that AI benefits should spread across countries, communities, and companies. He argues for a frontier ecosystem where closed and open-source models both thrive, and that organizations should keep control of their tacit knowledge and learning loops without depending on a single model provider. Microsoft plans to publish its first-party MAI models' "Code of Conduct" for public consultation tomorrow.

Sep 12

Sep 12Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score58

    Shanghai AI Lab releases Intern-S2-397B, a 397B multimodal scientific model

    AIShanghai AI Lab's InternLM team released Intern-S2-397B, a multimodal foundation model for scientific intelligence and long-horizon agents. The model uses visual pre-training on raw scientific literature pages, multi-task reinforcement learning across more than 20 scientific domains, and agentic reinforcement learning in sandboxed environments.

  2. Mike KnoopXAI score46

    Mike Knoop urges keeping AI research open amid slowdown proposals

    AIMike Knoop says he sees a path to an ARC-AGI-4 benchmark focused on open-ended invention, which he calls the gating capability between zero-sum automation and positive-sum innovation. He argues that coordinated slowdown efforts would likely apply to everyone, including open-source work, and cites chain of thought and the transformer as inventions that grew out of open science research. He concludes the research frontier must stay open to keep humanity on a positive-sum path.

Sep 11

Sep 11Fri
  1. InferactOfficialAI score34

    vLLM adds day-0 support for DeepSeek v4.1 Flash across six NVIDIA GPUs

    AIInferact says vLLM now supports DeepSeek v4.1 Flash on day zero across H100, H200, B200, B300, GB200, and GB300 GPUs. SemiAnalysis independently verified the NVIDIA support, while the post notes AMD vLLM still does not work with the model. Serving recipes are available at recipes.vllm.ai.

  2. Baseten BlogOfficialAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

  3. LM StudioOfficialAI score40

    LM Studio Bionic now live with DeepSeek-V4.1-Flash support

    AILM Studio has launched its Bionic version, now live for users. The announcement is tied to DeepSeek-V4.1-Flash, the smallest model in DeepSeek's new architecture family, which adds native visual understanding and targets faster inference and higher throughput.

  4. Interconnects (Nathan Lambert)BlogAI score38

    Open-Source AI & Open Models Reading List Is Updated for Research and Policy Writing

    AINathan Lambert has compiled a reading list of open-model writing covering why labs release open weights, the open-versus-closed debate, and US-China competition, last updated 15 September 2026. The list includes pieces on open-model economics, safety and marginal-risk research, and recent Chinese releases such as Kimi K3 and GLM-5.2. It also cites lawmaker inquiries into Western companies' use of Chinese models.

  5. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score72

    Shanghai AI Lab releases Atria Dawn Preview, an agentic model built on GLM-5.2

    AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.

    Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.

Sep 10

Sep 10Thu
  1. hardmaruXAI score52

    Sakana Fugu releases Fugu Max and Fugu Ultra v2 multi-agent orchestration models

    AISakana AI released Fugu Max and Fugu Ultra v2, multi-agent orchestration systems that route tasks across a pool of open-weights and specialized models. The source says Fugu Max delivers performance within striking distance of elite models at two to six times lower cost, while Fugu Ultra v2 outperforms Opus 5 and Fable 5 on Chartography and outperforms models costing three to five times more per token on DeepSWE.

    Image from @hardmaru's post
  2. Together AI BlogOfficialAI score52

    Together AI expands Fine-Tuning with live metrics, expert LoRA, and early stopping

    AITogether AI expanded its Fine-Tuning service with support for newer open-weight models, live metrics tracking, and finer training controls. Expert LoRA adapters can be applied to Mixture-of-Experts expert layers, and early stopping keeps the checkpoint with the best validation loss. Dataset previews, sample weights, pre-flight validation, and lower prices on selected models are also included.

  3. Ai2 · new models on Hugging FaceOfficialAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  4. Google LabsOfficialAI score36

    Google Dreambeans now free for US users, connects to Gemini

    AIGoogle Labs' Dreambeans is now available free to all US users aged 18 and older on iOS and Android, with no subscription required. Users can connect the Gemini app to Dreambeans, which will use their Gemini chat context to surface more personalized daily stories.

    Video from @GoogleLabs's post
  5. Cognition Blog (Devin, Windsurf)OfficialAI score22

    Cognition Welcomes Dioxus Team to Advance Open-Source Cross-Platform App Framework

    AICognition has welcomed Jonathan Kelley and the Dioxus team, whose framework Cognition used extensively to build and improve Devin's performance. Cognition plans to continue supporting Dioxus, Blitz, Taffy, and Subsecond while increasing investment in Dioxus-Native and Blitz. The Dioxus team will also work on Devin's virtual machine, computer use skills, and testing capabilities.

  6. Tencent HyOfficialAI score60

    Tencent Hunyuan releases open-source AuK audio model for speech generation and editing

    AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.

    Why it matters: The release combines speech generation and editing under one natural-language interface, and its 4-step AuK-Flash variant reports about 4.5 times faster inference under matched conditions.

    Video from @TencentHunyuan's post
  7. RadixArkOfficialAI score22

    RadixArk publishes Miles cookbook for DeepSeek V4.1 Flash

    AIRadixArk has published a cookbook on its Miles documentation site covering how to run DeepSeek V4.1 Flash. The post itself contains only a link to the cookbook page, so no further details about features, figures, or setup steps are available.

  8. RadixArkOfficialAI score60

    Miles adds day-0 RL support for DeepSeek-V4.1-Flash

    AIRadixArk says Miles brings day-0 RL support to DeepSeek-V4.1-Flash, with SGLang providing inference support. The post says quantization-aware training mirrors SGLang's FP4/FP8 rounding, and that colocated training and rollout fit full-parameter RL on 16 GPUs. In a DAPO run over steps 0–80, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78.

    Why it matters: The post pairs day-0 inference and RL support with specific training-consistency details, showing how the new architecture is handled in practice across GPUs.

  9. LMSYS OrgOfficialAI score62

    SGLang adds day-0 inference and RL support for DeepSeek V4.1 Flash

    AISGLang and Miles ship day-0 inference and RL support for DeepSeek V4.1 Flash, with weights now available. The model is natively multimodal with 552B backbone parameters, 16B active during decode and 8B during prefill, and supports up to 1M context. V4.1 adds shared compressed KV across layers, a two-stage sparse indexer, and a 196B Engram lookup memory.

    Why it matters: The post lists the architecture changes and parameter counts for V4.1 Flash, giving readers concrete specs to compare against earlier DeepSeek V4 releases.

Sep 9

Sep 9Wed
  1. BAAI · new models on Hugging FaceOfficialAI score24

    BAAI open-sources EPT, UniPath, and MiSI AIDD molecular and crystal modeling resources

    AIBAAI released open-source resources for three AIDD projects on Hugging Face: EPT, an equivariant pretrained transformer for unified 3D molecular representation learning, and UniPath, a learnable-time flow matching method for crystal structure and energy prediction. The repository mirrors their GitHub source code and READMEs, with setup, preprocessing, training, and evaluation documentation. The MiSI benchmark is released separately on Hugging Face.

  2. DeepSeek · new models on Hugging FaceOfficialAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

  3. Fireworks AI BlogOfficialAI score58

    Fireworks AI outlines a staged path from closed APIs to owned specialized models

    AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.

  4. TinkerOfficialAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.

  5. RadixArkOfficialAI score38

    RadixArk's Miles integrates SGLang for fast, aligned post-training rollouts

    AIRadixArk says its Miles framework natively supports SGLang for fast rollouts while keeping rollout and training aligned for reliable post-training at scale. The post thanks the community for contributions and feedback shaping Miles. A related post from @adarshxs describes Miles v0.1 running fully async agentic RL on a 744B MoE across 64 GB300 GPUs.

  6. Ai2 (Allen Institute for AI)OfficialAI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Ian Johnson 🔬🤖XAI score23

    Ian Johnson builds a font generator from letter-cluster embeddings

    AIIan Johnson (@enjalot) built a font generator after finding a cluster for each letter of the alphabet in his dataset, with Astra helping write the code. The tool is available as a Hugging Face Space and on GitHub, and the dataset includes SigLIP2 embeddings that allow concept search and clicking a result to jump to similar blocks.

    Video from @enjalot's post
  2. Ian Johnson 🔬🤖XAI score22

    Latent Craft lets users explore a million book images via UMAP in browser

    AIIan Johnson introduced Latent Craft, a new way to explore large datasets with UMAP, letting users fly through and collect images from them. The demo covers all 1 million images explorable in the browser, drawn from a dataset of 1,080,814 public domain images, mostly from 19th-century books, shared on the Hugging Face Hub.

    Video from @enjalot's post
  3. Google Developers BlogOfficialAI score72

    Google releases ADK for Kotlin 1.0 for building production AI agents

    AIGoogle announced general availability of ADK for Kotlin 1.0, a Kotlin Multiplatform framework for building AI agents on servers and Android. Version 1.0 reaches feature parity with ADK 1.0 Core and adds Android extensions for on-device models, cloud Gemini via Firebase AI Logic, and persistent sessions and memory with Room and AppSearch. The post includes a server-side incident triage example using KSP-generated tools and skills, plus an Android financial assistant example with human confirmation for transfers.

    Why it matters: The post names the new Android and server-side capabilities and the code setup, helping Kotlin developers judge whether ADK fits their agent projects.

  4. InferactOfficialAI score42

    Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

    AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

  5. Werner VogelsXAI score50

    Werner Vogels highlights Kiro Crew's memory system drawing on brain evolution

    AIWerner Vogels says that after spending time with Kiro Crew since its launch, its memory system stands out for deciding what to keep, compress, and let go. He notes that Amazon engineers, starting from engineering constraints, arrived at an approach resembling the brain's evolved architecture. Per the referenced post, Kiro Crew is a persistent workspace that retains project context across sessions and runs scheduled jobs.

  6. Cohere · new models on Hugging FaceOfficialAI score38

    Cohere releases Tiny Aya Base 32K, a 3.35B multilingual model with 32K context

    AICohere Labs has released Tiny Aya Base 32K, an open-weights pretrained model with 3.35 billion parameters and a 32K context window. The model covers 70+ languages, including many lower-resourced ones, and is designed for downstream adaptation and long-context research. It is a base model that has not been instruction-tuned, and it is licensed under CC-BY-NC.

  7. BAAIOfficialAI score43

    FlagEval-Robo tests 12 open-weight embodied AI models across simulation and real robots

    AIBAAI introduces FlagEval-Robo, an open dual-track evaluation suite linking simulation with real-world execution. The team post-trained and stress-tested 12 leading open-weight embodied AI models under strictly aligned conditions. The post raises whether high benchmark scores reflect physical reality, though it does not yet report specific results.

    Image from @BAAIBeijing's post
  8. Daniel HanXAI score28

    Qwen3.8-27B GGUF becomes the most-liked GGUF on Hugging Face

    AIUnsloth's Qwen3.8-27B GGUF is now the most-liked GGUF ever on Hugging Face, with the post predicting it will soon enter the top 30 most-liked models overall. Unsloth reports it reached 10M downloads and 3.7K likes in 24 days, and credits the community, Hugging Face, and the Qwen team.