Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. GoodfireOfficialAI score21

    Goodfire's probes run during inference with no added latency

    AIGoodfire reports that running its probes during model inference maintains the same throughput with no added latency. The company attributes this to infrastructure engineering, including kernel-level optimizations and a custom inference server.

  2. Daniel HanXAI score38

    Unsloth adds OS-level sandboxing for Linux, Mac, and Windows

    AIUnsloth now supports OS-level sandboxing on Linux via bwrap, on Mac via seatbelt, and on Windows via Microsoft's MXC. Per-tool-call latency is under 100ms across all three, and its software-style sandboxing with regex AST checks adds about 3ms. The Windows integration was built in collaboration with Microsoft.

  3. Unsloth AIOfficialAI score44

    Unsloth adds Windows OS-level sandboxing via Microsoft's mxc

    AIUnsloth now supports OS-level sandboxing on Windows by integrating Microsoft's open-source mxc repository for sandboxed code execution. The integration adds under 100 ms of overhead, according to the post. A setup guide is available in Unsloth's documentation.

    Image from @UnslothAI's post
  4. Arena.aiOfficialAI score55

    Arena raises $200M Series B and launches Alignment Index for AI agents

    AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

    Video from @arena's post
  5. elvisXAI score32

    Drama 3 voice model offers fine-grained tone and emotion control

    AIFish Audio's Drama 3 voice model lets users direct tone, emotion, and pacing in plain language, and can shift emotion mid-sentence. The poster, who found it remarkably effective in testing, says the control over delivery is unlike anything previously seen. A preview is available through the API as drama-3-preview.

  6. Augment Code BlogOfficialAI score62

    Augment Code sells Cosmos, Auggie CLI, and Context Engine assets to Harness

    AIAugment Code is selling select assets, including Cosmos, Auggie CLI, and the Code Context Engine, to Harness, and the product team is moving to Harness. The company says Harness's integrated platform delivers these capabilities to customers more effectively than building them independently. Harness describes itself as building the Autonomous SDLC Platform for shipping AI-written code across enterprises.

    Why it matters: The announcement shows how a coding AI company is folding its products into a larger software delivery platform, a shift that shapes how enterprise teams will buy these tools.

  7. SantiagoXAI score22

    Agent platform maps vulnerabilities and attack paths to protect systems

    AIA security platform uses agents to map a system's potential vulnerabilities and identify routes an attacker could take to reach sensitive data. It then recommends changes to close those paths. The quoted post cites a 700-agent swarm that breached Hugging Face with over 17,000 actions, and presents this tool, Cogent Attack Path Analysis, as the defensive counterpart.

  8. Tessl BlogOfficialAI score44

    Tessl Code Review Uses Repo-Owned Lenses to Make AI Review Context-Driven

    AITessl's Code Review defines review standards as skills in the repository, called lenses, routed to files by a repo-owned profile file. Because these team-visible standards are portable, lessons from review can feed back into code generation and maintenance, not only the next review.

  9. QbitAINewsAI score47

    Vidu Q4 Preview Offers 4K Video Generation at About 0.09 Yuan per Second

    AIShengshu Technology has opened a preview of its Vidu Q4 video generation model, which supports native 4K output and up to 15 reference images and three reference audio clips. Testers generated a one-minute video for about 5.4 yuan, roughly 0.09 yuan per second at 720P, which the article says is a starting price that varies by resolution and mode. The Vidu Q4 preview is available through the Vidu platform, with the MaaS API priced at about 0.6 yuan per second for 720P image-to-video.

  10. Understanding AI (Timothy B. Lee)BlogAI score67

    TypeSafe AI's Jev returns probabilities over fixed answers instead of text

    AITypeSafe AI released Jev, a model that answers yes/no, multiple-choice, or rating questions by outputting the estimated probability of each option. The author notes this design lets the model be served faster and more cheaply than LLMs and fits ordinary if-statement logic, and says he used it to flag spam comments on his blog in place of Gemini 3 Flash.

  11. QbitAINewsAI score34

    Physical AI firm Zhengxing Innovation unveils retail 24/7 human-robot collaboration solution

    AIZhengxing Innovation launched a Physical AI solution at APRCE 2026 for retail human-robot collaboration, built on its "embodied brain" and comprising the H1 humanoid and C1 wheeled-arm robots plus the M1 management platform. The company says the solution needs no store renovation, reports 99% autonomous task completion, and plans commercial service in 2027 via direct purchase or RaaS subscription.

  12. Artificial Analysis ArticlesOfficialAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    AIArtificial Analysis is adding trusted-access models to its Cyber Index, starting with GPT-6 Sol (Daybreak Blue, max), which is available only through OpenAI's Daybreak program. The model hits no safety blocks across the Index and scores 32 points higher overall than the publicly available GPT-6 Sol (max), with its largest gains on CyberGym-E2E.

    Why it matters: The source shows how safety refusals shape cyber benchmark scores, with the trusted-access model's gains concentrated on CyberGym-E2E, useful for comparing guarded and unguarded models.

Oct 7

Oct 7Wed
  1. Mastra BlogOfficialAI score60

    Mastra Connect adds ready-made tools for services like Linear and Notion

    AIMastra Connect is a public beta that lets Mastra projects connect providers such as Linear, Notion, and Slack, giving agents and workflows ready-made tools. Connect launches with 23 providers, almost 900 tools, and 7 hosted MCP providers, and it is free to use on Mastra platform during beta. Developers can add connections via the CLI or dashboard, limit tools with glob filters, and call a provider's SDK directly with credential() when a tool is missing.

    Why it matters: The post shows how connected services become agent tools, and how credentials and access limits are managed, which is useful for building agent workflows.

Oct 6

Oct 6Tue
  1. OpenRouter BlogOfficialAI score62

    ElevenLabs text-to-speech and speech-to-text models now available on OpenRouter

    AIElevenLabs now offers nine Text to Speech models and two Speech to Text models through OpenRouter, callable with an OpenRouter API key and no separate ElevenLabs plan. All ElevenLabs models are 50% off OpenRouter's list price through October 19, 8am PT, and Eleven v4, v4 Turbo, and Scribe v2 are recommended as starting points for narration, voice agents, and transcription.

    Why it matters: The source gives a concrete three-step build path and model selection guidance, showing how speech models plug into an existing text API for voice agents and transcription.

  2. Mastra BlogOfficialAI score67

    Mastra launches Agent Controller GA, a runtime for long-running agent sessions

    AIMastra has released Agent Controller in general availability, a runtime that hosts long-running agent sessions around the agent loop. The team says it was first built for Mastra Code and expanded to support Mastra Factory, which runs many concurrent sessions, and that memory usage in long-running Mastra Code processes dropped from 2–20 GB to 300–750 MB after optimizing UI state snapshots.

    Why it matters: The post explains how the controller evolved from one developer's session to many concurrent sessions, with measured memory and storage changes useful to engineers building multi-user agent apps.

Oct 2

Oct 2Fri
  1. Hugging Face BlogOfficialAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  2. Ai2 (Allen Institute for AI)OfficialAI score67

    Ai2 open-sources AstaBrief 8B, a fast open-weights scientific report model

    AIAi2 released AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. In Asta's Generate a report feature, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The model is built on Qwen3-8B with supervised fine-tuning and DPO, and institutions can run its open weights on their own infrastructure.

    Why it matters: The post explains the data filtering and one-pass generation choices behind a fast open-weights report model, showing what worked and what did not.

  3. Prime Intellect BlogOfficialAI score67

    Prime Inference launches serverless and reserved serving for open frontier models

    AIPrime Inference is a serving platform for frontier open-source models, offering serverless endpoints and reserved capacity on Prime's GPU infrastructure across multiple datacenters. Its first public deployment, GLM-5.3, went live on OpenRouter on September 22, and the post reports a near-zero tool-call error rate and 100% uptime since launch. The post also describes GLM-5.3 serving on GB200 NVL72 with prefill/decode disaggregation and NVFP4 KV compression.

    Why it matters: The post separates scheduler, KV-cache, and tool-call fixes, showing concretely which bottlenecks shape production serving of open frontier models.

Oct 1

Oct 1Thu
  1. Ai2 (Allen Institute for AI)OfficialAI score62

    Ai2 releases Olmo-core 3, an open framework for training large MoE models

    AIAi2 released Olmo-core 3, an open training framework redesigned to scale mixture-of-experts models into the trillion-parameter range. In one benchmark, expert count rose from 8 to 128 with about 3.2B active parameters per token, total capacity grew from 4.6B to 47B, and throughput fell by less than 5%. The framework is fully open, so researchers can train their own MoEs and experiment with routing and parallelism.

    Why it matters: The release documents concrete MoE scaling results and reported failure modes, useful for teams weighing training-stack tradeoffs before adopting an open framework.

Sep 30

Sep 30Wed
  1. Comfy BlogOfficialAI score60

    Comfy API launches to deploy ComfyUI workflows as autoscaling endpoints

    AIComfy API is now available to all users on a paid Comfy plan, letting them package a ComfyUI workflow with its custom nodes, LoRAs, models, and Python dependencies and deploy it as an autoscaling API endpoint. Builds capture the ComfyUI version and dependencies, and each immutable release gets its own URL, so the tested environment is the deployed one. Usage is billed separately, with GPU time charged by the second and storage prorated hourly.

    Why it matters: The post explains how a ComfyUI workflow is packaged into immutable releases and deployed as an autoscaling endpoint, showing a path from local graph to production service.

Sep 29

Sep 29Tue
  1. OpenClaw🦞OfficialAI score70

    OpenClaw Enterprise launches as an open-source control plane for persistent agents

    AIThe OpenClaw Foundation announced OpenClaw Enterprise, an open-source enterprise control plane for persistent agents, in collaboration with Red Hat, NVIDIA, and OpenAI. The product is built to run on an organization's own infrastructure and will always be free for organizations to use.

    Why it matters: The announcement names its collaborators and deployment model, which helps organizations judge how the enterprise control plane would fit their own infrastructure.

    Image from @openclaw's post
  2. BAAI · new models on Hugging FaceOfficialAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

Sep 23

Sep 23Wed
  1. Comfy BlogOfficialAI score62

    Comfy Router launches one API for frontier image, video, 3D, and audio models

    AIComfy Router is now live on the Comfy Developer Platform, giving developers one API to call frontier image, video, 3D, and audio models. Day one models include Seedance 2.5, MiniMax H3, Nano Banana Pro, GPT Image 2, Kling, and Black Forest Labs, and the provider for each job is selectable. Requests fail rather than silently switching providers, and inputs and outputs are deleted after 24 hours.

    Why it matters: The post shows how one API key and a provider parameter let developers swap routes for media models without rewriting calls, with failed requests reporting the provider.

  2. Prime Intellect BlogOfficialAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    AIPrime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    Why it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 21

Sep 21Mon
  1. vLLM BlogOfficialAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    AIvllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    Why it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  2. LMSYS OrgOfficialAI score65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    AILMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    Why it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

    Image from @lmsysorg's post

Sep 15

Sep 15Tue
  1. Zed BlogOfficialAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

Sep 11

Sep 11Fri
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score72

    Shanghai AI Lab releases Atria Dawn Preview, an agentic model built on GLM-5.2

    AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.

    Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.