Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 24

Sep 24Thu
  1. OpenBMBOfficialAI score34

    FIT-GGUF enables size-targeted mixed-precision quantization of MiniCPM5-2B

    AIDeveloper @Scorp1o_117 used FIT-GGUF to build four MiniCPM5-2B GGUF variants, ranging from about 1.14 GiB to 1.46 GiB, tuned to target file sizes or fidelity tiers. Instead of fixed presets, FIT-GGUF allocates precision tensor by tensor, with Quality, Balanced, Compact, and Mini options, and its generated files matched predicted sizes. Builds are evaluated with KL Divergence and Same-top metrics and are available on Hugging Face.

    Image from @OpenBMB's post
  2. KrASIA · Big TechNewsAI score55

    Mind Lab launches Mint Recursive, a post-training platform for companies

    AIMind Lab unveiled Mint Recursive, a post-training and inference platform for industry use, alongside Macaron-V1.1, a model post-trained entirely on it. Macaron-V1.1 is a 752-billion-parameter model built from GLM-5.3 with four two-billion-parameter LoRA expert modules for chat, agents, coding, and generation. The platform is serverless and bills by token usage, and it collects feedback from models in use to support continued training.

Sep 23

Sep 23Wed
  1. vLLM BlogOfficialAI score54

    vLLM adds distortion-free Gumbel-max watermarking for text provenance

    AIvLLM now supports Gumbel-max watermarking, which embeds a keyed signal into generated text without changing the expected token distribution. Detection requires the secret key and tokenizer, and the signal accumulates over longer outputs. Benchmarks on Qwen3.5-27B with MTP-3 show throughput changes between -1.1% and +2.0% across batch sizes, with no consistent slowdown.

  2. SemiAnalysisBlogAI score85

    SemiAnalysis releases ClusterMAX 3.0, rating 77 GPU clouds through hands-on testing

    AISemiAnalysis releases ClusterMAX 3.0, a rating of managed GPU clusters from neoclouds that covers 77 providers, with 323 in its market view. The rating is based on audit, performance, and reliability tests, along with interviews with over 200 end users. CoreWeave and Nebius hold the Platinum tier, Google Cloud and Oracle hold Gold, and only 19 providers earned a Medallion rating.

    Why it matters: The report shows how GPU cloud providers are ranked through hands-on tests, with details on benchmarks, reliability checks and SLA terms that buyers can reuse.

  3. InferactOfficialAI score49

    Inferact's TPU megakernel runs Kimi K3 at 709 tokens/s

    AIInferact says its first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode with DSpark speculative decoding, versus 450 tokens/s for its GB200 baseline. The company claims it is the first TPU inference megakernel, running the whole model in a single Pallas kernel, and says it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8 without speculative decoding. Inferact says it is open-sourcing the kernel today.

    Video from @inferact's post
  4. LM StudioOfficialAI score28

    Bionic adds a built-in interactive canvas for shared diagrams

    AIBionic now includes a built-in interactive canvas where users can create Excalidraw diagrams that both they and Bionic can view and edit. The canvas supports collaboration on mockups, system designs, and process maps, and users can ask Bionic to implement what is drawn.

    Video from @lmstudio's post
  5. Thomas DohmkeXAI score22

    Open-source 3D-printed Marvin robot connects to ChatGPT and Entire

    AIDeveloper Stefano (@spedemo) built a 3D-printed, remote-controlled robot named Marvin that uses voice detection and speech, with all of it open source. The post says Marvin connects to ChatGPT and Entire, and it will be shown at the WeAreDevs booth 753 in San Jose.

    Video from @ashtom's post
  6. Comfy BlogOfficialAI score62

    Comfy Router launches one API for frontier image, video, 3D, and audio models

    AIComfy Router is now live on the Comfy Developer Platform, giving developers one API to call frontier image, video, 3D, and audio models. Day one models include Seedance 2.5, MiniMax H3, Nano Banana Pro, GPT Image 2, Kling, and Black Forest Labs, and the provider for each job is selectable. Requests fail rather than silently switching providers, and inputs and outputs are deleted after 24 hours.

    Why it matters: The post shows how one API key and a provider parameter let developers swap routes for media models without rewriting calls, with failed requests reporting the provider.

  7. Mike KnoopXAI score57

    Tufa Labs reaches 83.06% on ARC-AGI-2, 2% short of the grand prize

    AIMike Knoop says the top ARC Prize 2026 ARC-AGI-2 score of 83.06% by Tufa Labs is only 2% short of the 85% grand prize threshold. The challenge runs under strict Kaggle compute limits with no internet access, and the winning solution is set to be open sourced. The image shows the leaderboard with RabbitHole at 76.94%, nvbanana at 74.17%, Yi-Chia Chen at 55.14%, and Kha Vo at 37.50%.

  8. ModelScopeOfficialAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post

Sep 22

Sep 22Tue
  1. Daniel HanXAI score22

    Unsloth Desktop hotfix adds Qwen-Image-2.1 image editing and fixes

    AIUnsloth Desktop received a hotfix update adding image editing for Qwen-Image-2.1. The update also fixes diffusers update issues, GGUF loading failures for Qwen-Image, and black artifacts during diffusion on A100 and consumer GPUs. Users should receive a banner prompting them to update.

  2. Unsloth AIOfficialAI score26

    Qwen-Image-2.1 FP8 and GGUF quants now run in Unsloth Desktop

    AIUnsloth announced that Qwen-Image-2.1 FP8 and GGUF quantized versions should now run properly in Unsloth Desktop. The app supports both image generation and image editing with these quants. Further details are available on the Unsloth GitHub repository.

    Image from @UnslothAI's post
  3. Tri DaoXAI score44

    Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs

    AIMayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.

  4. Comfy BlogOfficialAI score42

    ComfyUI Speeds Up MiniMax H3 Video VAE Encoding and Decoding

    AIComfyUI's update makes the MiniMax H3 video VAE encode up to about 2.2x faster and decode 1.4-2.7x faster, cutting a 1344x768, 129-frame round trip on an RTX 5090 from 24.3 to 12.7 seconds. The gains come from a fused encoder kernel enabled by default, fp16 accumulation support in a custom convolution, and an int8 decoder, and the source says the changes are visually lossless to the eye. Users need ComfyUI v0.36.0 or above, and the int8 VAE file is a drop-in replacement for the standard one.

  5. StepFunOfficialAI score43

    StepFun open-sources onPanda for token-level LLM annotation and inspection

    AIStepFun has open-sourced onPanda, a tool used internally for LLM data annotation and model inspection, letting users correct tokens and let models continue. The company reports a 52% lower median annotation time versus manual post-editing, with SFT and preference data combined in one workflow. It also supports token probability and top-k inspection, token-by-token decoding control, and browser-based testing across SVG generation, web development, and agent tasks.

  6. LlamaIndex 🦙OfficialAI score22

    LiteParse v2.14.6 parses text PDFs about 25% faster locally

    AILlamaIndex released LiteParse v2.14.6, an open-source PDF-to-Markdown parser that processes text-based PDFs about 25% faster. On realistic documents it handled pages at 2.8ms per page, 1.5 times faster than the next-fastest local parser. It runs locally in Python, Node.js, Rust, or directly in the browser.

    Image from @llama_index's post
  7. OpenBMBOfficialAI score59

    VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

    AIOpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2. The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook. VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

    Video from @OpenBMB's post

Sep 21

Sep 21Mon
  1. Kilo (acq. by Anaconda)OfficialAI score36

    Kilo says a newer Claude model breached OpenAI in three hours

    AIKilo's post says Hacktron spent hours failing to exploit a known flaw in an old image library, then a working exploit of OpenAI came within three hours after Claude Opus 5 shipped. The post argues that teams cannot afford model lock-in as frontier models change daily.

  2. vLLM BlogOfficialAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    AIvllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    Why it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  3. LlamaIndex 🦙OfficialAI score22

    LlamaIndex adds field-level confidence scores to Extract

    AILlamaIndex has added confidence scores to its Extract product, giving accuracy estimates field by field for extracted data. Developers can use these scores to decide which results their apps accept automatically and which need human review. The feature is available on the Cost Effective, Agentic, and Agentic Plus plans.

    Video from @llama_index's post
  4. LMSYS OrgOfficialAI score65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    AILMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    Why it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

    Image from @lmsysorg's post
  5. OpenBMBOfficialAI score23

    Developer builds local MiniCPM News Desk for traceable AI news briefings

    AIDeveloper Mark Fenner built MiniCPM News Desk, a local-first news briefing system powered by MiniCPM5-2B. The model selects key passages from official AI and technology sources, and a rule-based editorial layer preserves dates and context before assembling a daily recap. Invalid or incomplete outputs are rejected, and the full pipeline runs locally without a hosted-model fallback.

    Image from @OpenBMB's post
  6. KrASIA · Big TechNewsAI score44

    Aridge plans UAE flying car sandbox and Abu Dhabi–Ras Al Khaimah demo route

    AIAridge has signed a cooperation agreement with Ras Al Khaimah to set up a regulatory sandbox for advanced air mobility in the UAE, overseen by the UAE General Civil Aviation Authority. The program will test its Land Aircraft Carrier in desert conditions and plans a demonstration route between Abu Dhabi and Ras Al Khaimah. Regulatory approval for commercial passenger eVTOL flights remains pending.

  7. TechNode · AINewsAI score36

    IQAX Pushes eBLs, AI, and Digital Twins to Connect Global Trade Data

    AIIQAX has surpassed one million electronic Bills of Lading (eBLs), built on the GSBN blockchain and supporting DCSA, BIMCO, and ISO standards. The company is combining AI, IoT, and digital twins to move supply chain management from visibility toward predicting risks, with its AI-powered IoT platform covering more than 200 regions and 11,000 city pairs and about 92,000 connected devices.

Sep 20

Sep 20Sun
  1. OpenBMBOfficialAI score22

    MiniCPM5-2B on a local Mac correctly reconciles naproxen medication history

    AIOpenMed reports that MiniCPM5-2B, run on a local Mac, correctly kept naproxen in medication history rather than the current-medication export after a newer note said it was stopped. Every graph connection in the run links back to its source, using fictional clinical notes.

  2. LMSYS OrgOfficialAI score32

    RLinf adds Cosmos3 support with SGLang, boosting evaluation throughput 3.33x

    AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.

    Image from @lmsysorg's post
  3. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post
  4. OpenBMBOfficialAI score35

    OpenBMB's Augury model improves on-device plant ID for farmers

    AIA developer's Augury plant identification model, built on an OpenBMB model, raised photo top-1 accuracy from 71.8% to 80.2% by merging duplicate species keys and adding PCA whitening. The next steps are reaching 90%+ accuracy and building a phone GUI so farmers can use it on-device.

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score22

    Splash engine released as open source on GitHub

    AIThe Splash engine, posted by LM Studio, is now available as open source on GitHub. The post provides only a link to the incoai/splash repository and includes no further technical details.

  2. LM StudioOfficialAI score20

    LM Studio adds Splash engine for running incoai models on Mac

    AILM Studio users can enable the Splash engine under Settings > Runtime > Experimental backends and download supported models by searching "incoai." The post recommends an M3 or newer Mac running macOS 26.4 with 36GB+ RAM.

    Image from @lmstudio's post
  3. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.