Skip to contentSkip to stories

Updated

Open source

Showing low-relevance items too. Hide low-relevance items

Sep 18

Sep 18Fri
  1. SemiAnalysisBlogAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

  2. The Register · AINewsAI score34

    KDE turns 30 as Akademy weighs an AI-native desktop proposal

    AIKDE's Akademy conference in Graz, Austria, opens on September 19, where contributors Eva Brucherseifer and Jan Muehlig will present a talk proposing an "AI-native" KDE desktop built on a personal, encrypted "Kadai" kernel. The proposal's middle section is expected to divide attendees, while the project marks its 30th anniversary, with KDE 1.0 released in July 1998.

  3. Liquid AI · new models on Hugging FaceOfficialAI score55

    Liquid AI releases LFM2.5-VL-3B-DSpark drafter for faster vision-language decoding

    AILiquid AI released LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The source reports decoding up to 2.66× faster on a single H100 with SGLang, up to 3.13× on Apple M5 Max with MLX-VLM, and up to 2.14× on Apple M3 Ultra with llama.cpp, with output unchanged under greedy decoding.

  4. InferactOfficialAI score46

    Kimi K3 serving in vLLM is now 2.2–2.8× faster

    AIInferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.

Sep 17

Sep 17Thu
  1. OpenBMBOfficialAI score36

    OpenBMB's MiniCPM5-2B runs offline on-device with 128K context

    AIOpenBMB's MiniCPM5-2B is a 2.5B-parameter model with native 128K context, offering hybrid Think and No-Think modes in one checkpoint. Users can download it from Hugging Face and run it fully offline on-device, as RunAnywhere demonstrated. In a demo, the model first called a puzzle impossible, then corrected itself and wrote a working verifier.

  2. vLLM BlogOfficialAI score38

    vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

    AIvLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs. In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas. The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

  3. Sherwin WuXAI score34

    OpenAI Launches 47 Community-Built Legal Plugins in ChatGPT

    AIOpenAI launched 47 community-built plugins for legal work in ChatGPT, created by legal experts from LegalQuants, Skills.law, and LECG rather than white-labeled by OpenAI. The plugins are already live in the ChatGPT plugin store, alongside 26 partner-built plugins from companies including Thomson Reuters, Harvey, Legora, and iManage.

  4. Mark ZuckerbergXAI score60

    Meta releases Muse for Mac, an assistant that works across local apps and files

    AIMeta has released Muse for Mac, which works across apps, files, calendar, notes, and messages on the user's computer. Users control what Muse can access, and the download is available at

    Why it matters: The post names the Mac platform, the apps and data Muse can reach, and the user-controlled access setting, which clarifies how far the assistant operates on a personal computer.

  5. Google · AI blogOfficialAI score38

    UN System Data Commons unifies global statistics into an AI-ready open platform

    AIThe United Nations system launched UN System Data Commons, an open-source platform built on Data Commons by Google that integrates siloed global statistics into one AI-ready knowledge graph. Users can query it in natural language, browse by location or theme, and use MCP-enabled AI agents to fetch verified figures and draft charts or reports. The UN plans to add more datasets, aiming to include 80% of UN system statistical datasets by 2027.

  6. ChatGPTOfficialAI score38

    ChatGPT now available as an add-in inside Microsoft Word

    AIOpenAI's ChatGPT can now be added to Microsoft Word, letting users turn rough notes into first drafts, rework tangled paragraphs, proofread, and get suggested edits without leaving their document. The post also says it can flag formatting issues.

    Video from @ChatGPT's post
  7. Google AI StudioOfficialAI score58

    Google AI Studio open-sources Speakeasy's OpenAPI SDK generator suite

    AIGoogle AI Studio announced that Speakeasy is open sourcing its OpenAPI client generation suite under AGPLv3, following a May 2026 vendor shutdown that disrupted Google's SDK pipeline. The suite covers SDK generation for 7 languages, an agent-native CLI generator, and a documentation MCP server generator. Google says its pipeline now serves six targets with roughly one engineer maintaining it.

  8. Hacker News · Launch HN, YC launches (10+ points)BlogAI score58

    Skillsync launches tool to move AI chat sessions across coding agents

    AISkillsync, a Y Combinator W26 company, launched a tool that converts AI chat sessions between coding agents, including messages, reasoning and tool calls. The conversion engine txcript is open source, and the local-first app runs on Mac with a CLI and MCP support. Only sessions shared into team workspaces leave the user's machine.

  9. Daniel HanXAI score44

    Unsloth Desktop adds multi-user accounts and faster GRPO training

    AIUnsloth Desktop now supports multi-user accounts, alongside a revamped Docker image and custom Jupyter Notebook with custom themes, titles, and expandable cells. The update adds RDNA1+2 support, ARM64 Windows CUDA support, faster GRPO, and FP8/INT8 image diffusion support for 2x faster inference.

  10. Unsloth AIOfficialAI score60

    Unsloth Docker image lets users train and run 500+ models locally

    AIUnsloth announced that its Docker image now lets users train and run more than 500 models locally with no setup required. The image works on NVIDIA and AMD hardware and supports a new GUI or notebook workflow. The post links to an installation guide and the GitHub repository.

    Why it matters: The post explains how to run and train 500+ models locally with a no-setup Docker image, supported on NVIDIA and AMD, with a new GUI or notebook workflow.

    Image from @UnslothAI's post
  11. SenseTimeOfficialAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    AISenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

    Image from @SenseTime_AI's post
  12. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post
  13. OpenBMBOfficialAI score29

    Kahya-TTS: Turkish speech model fine-tuned from VoxCPM2 on 100 hours

    AIDeveloper Alican Kiraz fine-tuned OpenBMB's open-source VoxCPM2 voice model on nearly 100 hours of natural Turkish speech, creating Kahya-TTS for Turkish text-to-speech. The project shows how open-source voice models can be adapted to new languages and specialized datasets. The model is available on Hugging Face.

    Image from @OpenBMB's post
  14. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score46

    Ming-Image-0.1-Design-Layer splits flattened design images into RGBA layers

    AIinclusionAI has released Ming-Image-0.1-Design-Layer on Hugging Face, a model that decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan. The model runs at 1024 resolution (512 for faster processing) with 12 sampling steps, a CFG scale of 2.0, and BF16 precision on one CUDA GPU with 80 GiB VRAM. It is released under the MIT License.

  15. Ai2 (Allen Institute for AI)OfficialAI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

  16. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score42

    inclusionAI releases Ming-Image-0.1-Design, a 6B text-to-image model for text-rich designs

    AIinclusionAI has released Ming-Image-0.1-Design, a 6B text-to-image model for UI, infographics, and posters that outputs RGBA images with transparent backgrounds. The model is available on Hugging Face and ModelScope under the MIT License. It runs at 2048 x 2048 with 12 sampling steps and a CFG scale of 1.0, validated on one CUDA GPU with 80 GiB VRAM.

Sep 16

Sep 16Wed
  1. OpenBMBOfficialAI score20

    OpenBMB praises Dubedo's VoxCPM2-based voice cloning and dubbing studio

    AIOpenBMB says Dubedo is the kind of product it hoped VoxCPM2 would enable, citing speaker-aware cloning, multilingual generation, and an editing studio. The post praises @dubedostudio's work, while the background post describes Dubedo as dubbing into 30 languages with per-speaker voice cloning and a beta open for trial.

  2. Kling AIOfficialAI score38

    Fountain 0's ODYSSEUS: The Fall, fully generated with Kling 3.0, is out

    AIFountain 0's new feature film ODYSSEUS: The Fall, directed by Ash Koosha, is now available in full with every shot generated by Kling 3.0. The team previously premiered Dreams of Violets at the 2026 Tribeca Festival as the first AI feature film accepted into a major film festival.

    Video from @Kling_ai's post
  3. Google Developers BlogOfficialAI score38

    Google and Speakeasy open-source OpenAPI SDK generator suite under AGPLv3 license

    AISpeakeasy is open-sourcing its full OpenAPI client suite under the AGPLv3 license, including generators for seven languages (Python, TypeScript, Go, Java, C#, PHP, Ruby), an agent-native CLI generator, and a documentation MCP server generator. Google said the move followed the May 2026 shutdown of the SDK generation provider it had been using, which it cited as evidence that closed-source generators pose platform risk. Google's new Google GenAI SDKs for the Interactions, Agents, and Webhooks APIs were built with this pipeline across six targets.

  4. Bryan CatanzaroXAI score13

    Bryan Catanzaro to speak at GTC Berlin on open models

    AINVIDIA's Bryan Catanzaro, VP of Applied Deep Learning Research, will present at GTC Berlin on building open models developers can inspect, adapt, and deploy. The post is a conference invitation, with GTC Berlin set for October 20–22, 2026, and no new model or product announced.

  5. BAAI · new models on Hugging FaceOfficialAI score34

    BAAI and Peking University release Brainμ-Spike spike camera image reconstruction model

    AIPeking University's Yu Zhaofei team and the Beijing Academy of Artificial Intelligence (BAAI) released Brainμ-Spike, a small convolutional network for spike camera image reconstruction that is paired with the Brainμ model. The package includes weights, inference scripts, and evaluation tools, but the base large model and LoRA weights are not yet released, so the full generation pipeline cannot run from this repository alone.

  6. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score55

    inclusionAI releases Realtime-Venus full-duplex audio-visual models on Hugging Face

    AIinclusionAI has published Realtime-Venus on Hugging Face with two 9B checkpoints: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for audio-only conversation. Both are built on MiniCPM-o 4.5 with a Qwen3-8B backbone and support full-duplex dialogue, proactive responses, and training-free long-video memory. The asynchronous Realtime-Venus-Harness runtime is hosted in a separate GitHub repository.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceOfficialAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Zed BlogOfficialAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.