Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 20

Sep 20Sun
  1. QwenOfficialAI score34

    Qwen-Image-2.1 now supported in ComfyUI for image generation

    AIQwen-Image-2.1 is now supported in ComfyUI, and Qwen invites users to try it and share their creations. ComfyUI describes it as an open-weights 7B checkpoint that handles both generation and editing, with native 2K image generation and instruction editing from up to 10 reference images in one pass.

  2. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post
  3. QwenOfficialAI score56

    Qwen-Image-2.1 releases open weights for image generation and editing

    AIAlibaba's Qwen team released Qwen-Image-2.1 as an open-weights image model for both generation and editing, with a lightweight 7B architecture. The model natively generates and edits RGBA layers, supports up to 10 reference images for editing, and is available on GitHub, ModelScope, and Hugging Face.

    Image from @Alibaba_Qwen's post
  4. ModelScopeOfficialAI score62

    Qwen-Image-2.1 unifies image generation and editing with native transparency

    AIAlibaba's ModelScope introduces Qwen-Image-2.1, a model that handles image generation and editing together, with native transparency and a compact 7B visual generation component. It adds KV cache reuse to speed up generation and editing while reducing memory use, especially with multiple reference images. The model can combine up to 10 reference images, make targeted local edits, and preserve portrait identity and product details.

    Why it matters: The post names concrete capabilities and a 7B size, letting readers compare it against the larger image models in the accompanying chart.

    Image from @ModelScope2022's post
  5. OpenBMBOfficialAI score35

    OpenBMB's Augury model improves on-device plant ID for farmers

    AIA developer's Augury plant identification model, built on an OpenBMB model, raised photo top-1 accuracy from 71.8% to 80.2% by merging duplicate species keys and adding PCA whitening. The next steps are reaching 90%+ accuracy and building a phone GUI so farmers can use it on-device.

  6. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases Qwen-Image-2.1 prompt rewriter for image editing on Hugging Face

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B visual generation parameters. The Hugging Face page for Qwen-Image-2.1-PE-I2I is a fine-tuned Qwen3.5-VL 9B prompt rewriter that turns vague editing instructions and input images into precise editing prompts, supporting up to 10 reference images.

    Why it matters: The model card documents usage with transformers and diffusers, letting readers see how the editing prompt rewriter connects to the generation pipeline.

  7. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases open-source Qwen-Image-2.1 with a prompt rewriting model

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a 7B-parameter visual generation component. The release also includes Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B model that rewrites brief image requests in any language into detailed English prompts with a recommended aspect ratio.

    Why it matters: The release pairs a 7B visual generation component with a separate prompt rewriting model, showing how a brief image request becomes a detailed English prompt before rendering.

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Why it matters: The post pairs a cost-versus-intelligence chart with specs and a later open-weights date, so readers can judge the cost tradeoff against named competitor models.

    Image from @StepFun_ai's post
  2. OpenBMBOfficialAI score34

    OpenBMB's 2B MiniCPM5 powers a local personal news desk

    AIOpenBMB's 2B-parameter MiniCPM5 model runs as a local news desk on an older i5-9400F PC with 16GB RAM and no cloud API. The developer built a system that collects official sources hourly and sends a 24-hour Telegram recap with a lead story and links.

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score22

    Splash engine released as open source on GitHub

    AIThe Splash engine, posted by LM Studio, is now available as open source on GitHub. The post provides only a link to the incoai/splash repository and includes no further technical details.

  2. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

    Why it matters: The post reports a concrete decode speed on Apple silicon and compares it against named local inference tools, which helps readers judge local deployment performance.

  3. TinkerOfficialAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  4. Google GemmaOfficialAI score22

    DiffusionGemma runs as a parallel decision model, faster than autoregressive generation

    AIGoogle Gemma's account says DiffusionGemma, running in a Jev-style decision setup, denoises an open canvas in one step rather than generating tokens sequentially, taking about 0.2 seconds on a DGX Spark. It says full bidirectional attention lets every option attend to the full context at once, and that the model inherits Gemma 4's spatial vision capabilities for visual and text decisions.

    Video from @googlegemma's post
  5. LMSYS OrgOfficialAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

    Image from @lmsysorg's post
  6. SemiAnalysisBlogAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

  7. The Register · AINewsAI score34

    KDE turns 30 as Akademy weighs an AI-native desktop proposal

    AIKDE's Akademy conference in Graz, Austria, opens on September 19, where contributors Eva Brucherseifer and Jan Muehlig will present a talk proposing an "AI-native" KDE desktop built on a personal, encrypted "Kadai" kernel. The proposal's middle section is expected to divide attendees, while the project marks its 30th anniversary, with KDE 1.0 released in July 1998.

  8. Liquid AI · new models on Hugging FaceOfficialAI score55

    Liquid AI releases LFM2.5-VL-3B-DSpark drafter for faster vision-language decoding

    AILiquid AI released LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The source reports decoding up to 2.66× faster on a single H100 with SGLang, up to 3.13× on Apple M5 Max with MLX-VLM, and up to 2.14× on Apple M3 Ultra with llama.cpp, with output unchanged under greedy decoding.

  9. InferactOfficialAI score46

    Kimi K3 serving in vLLM is now 2.2–2.8× faster

    AIInferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.

Sep 17

Sep 17Thu
  1. OpenBMBOfficialAI score36

    OpenBMB's MiniCPM5-2B runs offline on-device with 128K context

    AIOpenBMB's MiniCPM5-2B is a 2.5B-parameter model with native 128K context, offering hybrid Think and No-Think modes in one checkpoint. Users can download it from Hugging Face and run it fully offline on-device, as RunAnywhere demonstrated. In a demo, the model first called a puzzle impossible, then corrected itself and wrote a working verifier.

  2. vLLM BlogOfficialAI score38

    vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

    AIvLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs. In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas. The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

  3. Sherwin WuXAI score34

    OpenAI Launches 47 Community-Built Legal Plugins in ChatGPT

    AIOpenAI launched 47 community-built plugins for legal work in ChatGPT, created by legal experts from LegalQuants, Skills.law, and LECG rather than white-labeled by OpenAI. The plugins are already live in the ChatGPT plugin store, alongside 26 partner-built plugins from companies including Thomson Reuters, Harvey, Legora, and iManage.

  4. Mark ZuckerbergXAI score60

    Meta releases Muse for Mac, an assistant that works across local apps and files

    AIMeta has released Muse for Mac, which works across apps, files, calendar, notes, and messages on the user's computer. Users control what Muse can access, and the download is available at

    Why it matters: The post names the Mac platform, the apps and data Muse can reach, and the user-controlled access setting, which clarifies how far the assistant operates on a personal computer.

  5. Google · AI blogOfficialAI score38

    UN System Data Commons unifies global statistics into an AI-ready open platform

    AIThe United Nations system launched UN System Data Commons, an open-source platform built on Data Commons by Google that integrates siloed global statistics into one AI-ready knowledge graph. Users can query it in natural language, browse by location or theme, and use MCP-enabled AI agents to fetch verified figures and draft charts or reports. The UN plans to add more datasets, aiming to include 80% of UN system statistical datasets by 2027.

  6. ChatGPTOfficialAI score38

    ChatGPT now available as an add-in inside Microsoft Word

    AIOpenAI's ChatGPT can now be added to Microsoft Word, letting users turn rough notes into first drafts, rework tangled paragraphs, proofread, and get suggested edits without leaving their document. The post also says it can flag formatting issues.

    Video from @ChatGPT's post
  7. Google AI StudioOfficialAI score58

    Google AI Studio open-sources Speakeasy's OpenAPI SDK generator suite

    AIGoogle AI Studio announced that Speakeasy is open sourcing its OpenAPI client generation suite under AGPLv3, following a May 2026 vendor shutdown that disrupted Google's SDK pipeline. The suite covers SDK generation for 7 languages, an agent-native CLI generator, and a documentation MCP server generator. Google says its pipeline now serves six targets with roughly one engineer maintaining it.

  8. Hacker News · Launch HN, YC launches (10+ points)BlogAI score58

    Skillsync launches tool to move AI chat sessions across coding agents

    AISkillsync, a Y Combinator W26 company, launched a tool that converts AI chat sessions between coding agents, including messages, reasoning and tool calls. The conversion engine txcript is open source, and the local-first app runs on Mac with a CLI and MCP support. Only sessions shared into team workspaces leave the user's machine.

  9. Daniel HanXAI score44

    Unsloth Desktop adds multi-user accounts and faster GRPO training

    AIUnsloth Desktop now supports multi-user accounts, alongside a revamped Docker image and custom Jupyter Notebook with custom themes, titles, and expandable cells. The update adds RDNA1+2 support, ARM64 Windows CUDA support, faster GRPO, and FP8/INT8 image diffusion support for 2x faster inference.

  10. Unsloth AIOfficialAI score60

    Unsloth Docker image lets users train and run 500+ models locally

    AIUnsloth announced that its Docker image now lets users train and run more than 500 models locally with no setup required. The image works on NVIDIA and AMD hardware and supports a new GUI or notebook workflow. The post links to an installation guide and the GitHub repository.

    Why it matters: The post explains how to run and train 500+ models locally with a no-setup Docker image, supported on NVIDIA and AMD, with a new GUI or notebook workflow.

    Image from @UnslothAI's post
  11. SenseTimeOfficialAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  12. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post
  13. OpenBMBOfficialAI score29

    Kahya-TTS: Turkish speech model fine-tuned from VoxCPM2 on 100 hours

    AIDeveloper Alican Kiraz fine-tuned OpenBMB's open-source VoxCPM2 voice model on nearly 100 hours of natural Turkish speech, creating Kahya-TTS for Turkish text-to-speech. The project shows how open-source voice models can be adapted to new languages and specialized datasets. The model is available on Hugging Face.

    Image from @OpenBMB's post
  14. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score46

    Ming-Image-0.1-Design-Layer splits flattened design images into RGBA layers

    AIinclusionAI has released Ming-Image-0.1-Design-Layer on Hugging Face, a model that decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan. The model runs at 1024 resolution (512 for faster processing) with 12 sampling steps, a CFG scale of 2.0, and BF16 precision on one CUDA GPU with 80 GiB VRAM. It is released under the MIT License.

  15. Ai2 (Allen Institute for AI)OfficialAI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

  16. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score42

    inclusionAI releases Ming-Image-0.1-Design, a 6B text-to-image model for text-rich designs

    AIinclusionAI has released Ming-Image-0.1-Design, a 6B text-to-image model for UI, infographics, and posters that outputs RGBA images with transparent backgrounds. The model is available on Hugging Face and ModelScope under the MIT License. It runs at 2048 x 2048 with 12 sampling steps and a CFG scale of 1.0, validated on one CUDA GPU with 80 GiB VRAM.