Skip to contentSkip to stories

Updated

#Product update

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 21

Sep 21Mon
  1. Engineering at MetaOfficialAI score39

    Meta Open-Sources Rebalancer, a Library for Solving Assignment Problems

    AIMeta has open-sourced Rebalancer, an assignment-problem solver it has used for over nine years to allocate resources across its infrastructure. The library separates problem specification, in-memory storage, solving, and debugging, and translates problems into expression graphs solved via local search or mixed integer programs using FICO Xpress, Gurobi, or the open-source HiGHS solver.

  2. Gemini NotebookOfficialAI score34

    Live Chat rolls out to all Ultra users on mobile

    AIGoogle's Gemini Notebook says Live Chat is now fully rolled out to all Ultra subscribers on mobile. The feature enables real-time voice conversations with notebooks in about 100 languages, powered by the latest audio models. Users can ask questions about their sources and receive step-by-step guidance largely hands-free.

    Video from @Gemini_Notebook's post
  3. OpenBMBOfficialAI score23

    Developer builds local MiniCPM News Desk for traceable AI news briefings

    AIDeveloper Mark Fenner built MiniCPM News Desk, a local-first news briefing system powered by MiniCPM5-2B. The model selects key passages from official AI and technology sources, and a rule-based editorial layer preserves dates and context before assembling a daily recap. Invalid or incomplete outputs are rejected, and the full pipeline runs locally without a hosted-model fallback.

    Image from @OpenBMB's post
  4. MiniMax Design (H3)OfficialAI score22

    Community speeds up MiniMax H3 video generation with sparse attention

    AIA community developer integrated the Jev method into MiniMax H3 to sparsify attention, deciding per layer which parts to keep. On an RTX 4070, video generation time dropped from 6 min 7 sec to 3 min 34 sec, a 41.7% reduction. The post notes that Jev selected sparsity rates of 1%, 3%, 5%, and 10% across 49 layers in 4-step generation.

  5. WorkBuddyOfficialAI score26

    WorkBuddy adds GLM-5.3-Flash to its model lineup

    AIWorkBuddy has added GLM-5.3-Flash to its model lineup and made it available now on the platform. The post invites users to try the model on their next task, but gives no further details on capabilities, speed, pricing, or benchmarks.

    Image from @WorkBuddy_AI's post

Sep 20

Sep 20Sun
  1. QwenOfficialAI score38

    Qwen-Image-2.1 launches live Spaces demo for generation and editing

    AIQwen-Image-2.1 is now available as a live Hugging Face Spaces demo, letting users try image generation and editing in the browser without setup. The demo runs on a single checkpoint that handles both tasks. Background notes describe a 7B-parameter model supporting up to 10 image references, with integration into diffusers and ComfyUI.

  2. LMSYS OrgOfficialAI score32

    RLinf adds Cosmos3 support with SGLang, boosting evaluation throughput 3.33x

    AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.

    Image from @lmsysorg's post
  3. QwenOfficialAI score34

    Qwen-Image-2.1 now supported in ComfyUI for image generation

    AIQwen-Image-2.1 is now supported in ComfyUI, and Qwen invites users to try it and share their creations. ComfyUI describes it as an open-weights 7B checkpoint that handles both generation and editing, with native 2K image generation and instruction editing from up to 10 reference images in one pass.

  4. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post
  5. QwenOfficialAI score34

    Qwen-Image-2.1 edits three marked regions in one prompt

    AIQwen-Image-2.1 supports several ways to specify local edits, including circles marking regions on an image. In the example, the model removes a metal watch, changes hair color to black, and replaces clothing in three circled areas in a single pass.

    Image from @Alibaba_Qwen's post

Sep 19

Sep 19Sat

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

  2. Google GemmaOfficialAI score22

    DiffusionGemma runs as a parallel decision model, faster than autoregressive generation

    AIGoogle Gemma's account says DiffusionGemma, running in a Jev-style decision setup, denoises an open canvas in one step rather than generating tokens sequentially, taking about 0.2 seconds on a DGX Spark. It says full bidirectional attention lets every option attend to the full context at once, and that the model inherits Gemma 4's spatial vision capabilities for visual and text decisions.

    Video from @googlegemma's post
  3. Greg BrockmanXAI score44

    ChatGPT now lets users connect multiple accounts to most plugins

    AIChatGPT now supports connecting multiple accounts to most plugins, so users can bring work and personal context into one conversation. Connections are made in the plugin directory at Developers need no changes, though adding a profile tool to their MCP server lets ChatGPT label each account.

  4. Thomas DohmkeXAI score24

    Claude Code adds AGENTS.md support in version 2.1.277

    AIClaude Code version 2.1.277 now reads AGENTS.md when a folder has no CLAUDE.md, according to Anthropic engineer Thariq Shihipar's post. The behavior can be toggled in /config. Thomas Dohmke's main post jokingly says AI is finally aligned, with no further technical detail.

  5. Google AIOfficialAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

  6. LMSYS OrgOfficialAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

    Image from @lmsysorg's post
  7. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

  8. KrASIA · Big TechNewsAI score47

    Huawei unveils Atlas 960E superpod linking 4,096 NPUs with near-packaged optics

    AIHuawei unveiled the Atlas 960E superpod, which links up to 4,096 NPUs using near-packaged optics (NPO) and claims eight exaflops at FP8 precision and up to one petabyte of high-bandwidth memory. Its Hi-ONE engine provides 7.2 terabits per second of transmission capacity, and Huawei says 5,500 engines replace 48,000 conventional 800G optical modules, cutting power consumption by more than 550 kilowatts. Huawei has proposed its NPO implementation agreement to the Optical Internetworking Forum, though it has not yet become a standard.

  9. InferactOfficialAI score46

    Kimi K3 serving in vLLM is now 2.2–2.8× faster

    AIInferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.