Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 21

Sep 21Mon
  1. RadixArkOfficialAI score25

    RadixArk's Miles adds async rollout buffer as swappable RL primitive

    AIRadixArk says its Miles framework uses an async rollout buffer that can change which sample groups reach training and which prompts get retried, while reusing the rollout worker and trainer. The post argues that stable, granular extension points let contributors modify one part of an RL system without disrupting its neighbors.

  2. Amazon ScienceOfficialAI score47

    Amazon Bio Discovery's three AI methods accelerate antibody drug design

    AIAmazon Bio Discovery developed three AI approaches for antibody drug design: MochiBind for sequence-based affinity ranking, CA-MAP for developability prediction with batch effect correction, and an agent-guided design system. The agent-guided system produced 46 lab-validated hits against a novel cancer target.

  3. Mike KnoopXAI score38

    Mike Knoop says LLM logprobs are vanishing, yet they enable useful new patterns

    AIMike Knoop notes that logprobs used to be widely exposed by LLM inference APIs and sees the market maturing so that parts of the LLM stack can be packaged in new, useful ways. He links this to Bryan Helmig's post on prompting with max_tokens: 1 plus logprobs for fast, parallel judgments, which Helmig says has a lot more depth than he expected.

  4. SemiAnalysisBlogAI score62

    How MoE inference splits into prefill, midfill, and decode regimes

    AIThe article explains how Mixture of Experts models change inference by making prefill, midfill, decode attention, and decode experts distinct workloads. It describes how KV cache state, expert routing, and parallelism choices shape compute, memory, and network demands across an inference cluster.

  5. Amazon ScienceOfficialAI score60

    Amazon Science reports AI models for designing and characterizing antibodies

    AIAmazon Science describes three papers on AI for antibody discovery: MochiBind ranks antibody binding strength from sequence alone, CA-MAP predicts developability properties using batch-aware context, and an agent-guided pipeline designed nanobody binders against a novel cancer target. In the pipeline, 116 candidates survived lab screening, and 46 were identified as strong binders, which are being used to train the next design cycle.

    Why it matters: The source reports the method, benchmark setup, and experimental validation in a single design workflow, showing how predictors, agents, and lab screening connect in antibody discovery.

  6. ChatGPTOfficialAI score40

    ChatGPT now lets US Plus and Pro users track credit scores via Experian

    AIChatGPT users on Plus and Pro plans in the U.S. can now securely connect their Experian credit report and VantageScore 3.0 credit score in Finances. Once connected, the feature provides personalized insights into what is affecting their score and how it relates to their financial goals, with alerts when things change. It is available on web and the latest iOS and Android apps.

    Image from @ChatGPT's post
  7. LlamaIndex 🦙OfficialAI score22

    LlamaIndex adds field-level confidence scores to Extract

    AILlamaIndex has added confidence scores to its Extract product, giving accuracy estimates field by field for extracted data. Developers can use these scores to decide which results their apps accept automatically and which need human review. The feature is available on the Cost Effective, Agentic, and Agentic Plus plans.

    Video from @llama_index's post
  8. LMSYS OrgOfficialAI score30

    LMSYS Publishes Blog Post on NVFP4 KV Cache Quantization

    AILMSYS Org shared a blog post about NVFP4 KV cache, a topic linked from its 2026-09-16 article. The post itself contains only a link, so no further technical details, figures, or results can be confirmed from this source.

  9. LMSYS OrgOfficialAI score65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    AILMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    Why it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

    Image from @lmsysorg's post
  10. Engineering at MetaOfficialAI score39

    Meta Open-Sources Rebalancer, a Library for Solving Assignment Problems

    AIMeta has open-sourced Rebalancer, an assignment-problem solver it has used for over nine years to allocate resources across its infrastructure. The library separates problem specification, in-memory storage, solving, and debugging, and translates problems into expression graphs solved via local search or mixed integer programs using FICO Xpress, Gurobi, or the open-source HiGHS solver.

  11. Microsoft ResearchOfficialAI score50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    AIMicrosoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

  12. OpenBMBOfficialAI score23

    Developer builds local MiniCPM News Desk for traceable AI news briefings

    AIDeveloper Mark Fenner built MiniCPM News Desk, a local-first news briefing system powered by MiniCPM5-2B. The model selects key passages from official AI and technology sources, and a rule-based editorial layer preserves dates and context before assembling a daily recap. Invalid or incomplete outputs are rejected, and the full pipeline runs locally without a hosted-model fallback.

    Image from @OpenBMB's post
  13. TechNode · AINewsAI score36

    IQAX Pushes eBLs, AI, and Digital Twins to Connect Global Trade Data

    AIIQAX has surpassed one million electronic Bills of Lading (eBLs), built on the GSBN blockchain and supporting DCSA, BIMCO, and ISO standards. The company is combining AI, IoT, and digital twins to move supply chain management from visibility toward predicting risks, with its AI-powered IoT platform covering more than 200 regions and 11,000 city pairs and about 92,000 connected devices.

  14. MiniMax Design (H3)OfficialAI score22

    Community speeds up MiniMax H3 video generation with sparse attention

    AIA community developer integrated the Jev method into MiniMax H3 to sparsify attention, deciding per layer which parts to keep. On an RTX 4070, video generation time dropped from 6 min 7 sec to 3 min 34 sec, a 41.7% reduction. The post notes that Jev selected sparsity rates of 1%, 3%, 5%, and 10% across 49 layers in 4-step generation.

  15. Matei ZahariaXAI score32

    Matei Zaharia praises GEPA working with Jev

    AIMatei Zaharia, a prominent AI researcher, said it is very cool that GEPA works on Jev. The post is a short endorsement, linking to background about a test in which GEPA optimized Jev's prompts for extracting suspected adverse drug effects from medical sentences.

  16. WorkBuddyOfficialAI score26

    WorkBuddy adds GLM-5.3-Flash to its model lineup

    AIWorkBuddy has added GLM-5.3-Flash to its model lineup and made it available now on the platform. The post invites users to try the model on their next task, but gives no further details on capabilities, speed, pricing, or benchmarks.

    Image from @WorkBuddy_AI's post

Sep 20

Sep 20Sun
  1. OpenBMBOfficialAI score22

    MiniCPM5-2B on a local Mac correctly reconciles naproxen medication history

    AIOpenMed reports that MiniCPM5-2B, run on a local Mac, correctly kept naproxen in medication history rather than the current-medication export after a newer note said it was stopped. Every graph connection in the run links back to its source, using fictional clinical notes.

  2. QwenOfficialAI score38

    Qwen-Image-2.1 launches live Spaces demo for generation and editing

    AIQwen-Image-2.1 is now available as a live Hugging Face Spaces demo, letting users try image generation and editing in the browser without setup. The demo runs on a single checkpoint that handles both tasks. Background notes describe a 7B-parameter model supporting up to 10 image references, with integration into diffusers and ComfyUI.

  3. vLLM BlogOfficialAI score44

    vLLM Reports PD Serving Results for Qwen3.8-2.4T on GB300 NVL72

    AIvLLM achieved 5000 total token throughput per GPU in high-throughput PD serving of Qwen3.8-2.4T on a GB300 NVL72 cluster under an 8K/1K workload. The low-latency scenario reached 180 generated tokens per user, with both results shown on the Pareto frontier. The post also provides srt-slurm recipes and explains the tuning process used to create them.

  4. LMSYS OrgOfficialAI score32

    RLinf adds Cosmos3 support with SGLang, boosting evaluation throughput 3.33x

    AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.

    Image from @lmsysorg's post
  5. Sebastian RaschkaXAI score38

    Raschka Says Jev's Classifier Generalizes Well, Credits Data

    AISebastian Raschka argues that Jev is more than just a classifier, since it generalizes well where earlier encoder-style classification models were usually special-purpose and limited. He suggests the main advantage lies in its data rather than the training algorithm, along with a well-designed API.

  6. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score29

    StepFun opens Step 5 Preview for trial on its platform

    AIStepFun invites developers to try the Step 5 Preview through its online platform. The post links to platform documentation for the model and a Discord community for discussion. No specific capabilities, benchmarks, or pricing are stated in the post.

  2. StepFunOfficialAI score38

    StepFun previews Step 5 for large-scale research and analytical deliverables

    AIStepFun has previewed Step 5, an agent built for professional knowledge work spanning large-scale research, structured analysis, and interactive reporting. In one agent action, it coordinated 950 web fetches and assembled 300,000 monthly records across 1,000 locations over 25 years. In another, it produced a 17-sheet analytical workbook with source reconciliation, formulas, and trend models.

    Image from @StepFun_ai's post
  3. OpenBMBOfficialAI score34

    OpenBMB's 2B MiniCPM5 powers a local personal news desk

    AIOpenBMB's 2B-parameter MiniCPM5 model runs as a local news desk on an older i5-9400F PC with 16GB RAM and no cloud API. The developer built a system that collects official sources hourly and sends a 24-hour Telegram recap with a lead story and links.

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score22

    Splash engine released as open source on GitHub

    AIThe Splash engine, posted by LM Studio, is now available as open source on GitHub. The post provides only a link to the incoai/splash repository and includes no further technical details.

  2. LM StudioOfficialAI score20

    LM Studio adds Splash engine for running incoai models on Mac

    AILM Studio users can enable the Splash engine under Settings > Runtime > Experimental backends and download supported models by searching "incoai." The post recommends an M3 or newer Mac running macOS 26.4 with 36GB+ RAM.

    Image from @lmstudio's post
  3. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

    Why it matters: The post reports a concrete decode speed on Apple silicon and compares it against named local inference tools, which helps readers judge local deployment performance.

  4. Mark ZuckerbergXAI score38

    Meta opens developer access to build Muse connectors for its agent

    AIMeta is opening access for developers to build connectors for Muse, its agent platform. Developers supply the API, while Muse provides the agent, browser, and user context, so people can reach a service simply by asking and their agent handles the rest. New connectors are live today at

    Video from @finkd's post
  5. Google GemmaOfficialAI score22

    DiffusionGemma runs as a parallel decision model, faster than autoregressive generation

    AIGoogle Gemma's account says DiffusionGemma, running in a Jev-style decision setup, denoises an open canvas in one step rather than generating tokens sequentially, taking about 0.2 seconds on a DGX Spark. It says full bidirectional attention lets every option attend to the full context at once, and that the model inherits Gemma 4's spatial vision capabilities for visual and text decisions.

    Video from @googlegemma's post