Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 1

Oct 1Thu
  1. NVIDIA BlogOfficialAI score62

    NVIDIA Blackwell GPUs power OpenAI's GPT-6 Astra Ultrafast mode in API

    AIGPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is now available in the OpenAI API and to eligible ChatGPT Work and Codex users. The source says Ultrafast offers up to 8x faster token generation than Astra Standard mode, which can shorten coding agents' response times between tool calls. OpenAI also says it uses its own models to keep optimizing inference software on NVIDIA GPUs after deployment.

    Why it matters: The source ties a specific speed claim to coding agents' edit-test-debug loops, showing where faster token generation changes developer workflows.

  2. Harrison ChaseXAI score33

    Harrison Chase Argues Every Agent Harness Needs a Durable Runtime

    AIHarrison Chase argues that every agent harness requires a durable runtime, citing pi-durable as an example alongside deepagents built on LangGraph. The post frames durable execution as a basic requirement for agent systems rather than an optional feature. Pi 1.0 shipped with Pi Durable, which the referenced @pidotdev post invites users to customize.

  3. Boris ChernyXAI score54

    Claude Code adds mods that customize behavior and UI via plugins

    AIClaude Code now supports mods that change its behavior, customize the UI, or add features, written in a few lines of TypeScript or generated by Claude. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app, and the author says mods can be shared so others can try them.

  4. Alex HeathXAI score46

    OpenAI's Dots lead ChatGPT's always-on personal agent plans

    AIAlex Heath's podcast with OpenAI's @embirico covers Dots, a new always-on personal agent in ChatGPT that asks permission before acting by default. The episode also discusses Space, a workspace where people and agents collaborate on documents, data, and projects, along with pricing and possible access for free users.

    Video from @alexeheath's post
  5. Latent.SpaceXAI score55

    Anthropic's Thariq explains what Claude Code mods can access and control

    AIAnthropic's Thariq describes Claude Code mods, which can read conversation scope such as turn count and token usage. Mods run in process, so they can spawn subagents, parse their results with structured output, and modify the UI, which hooks cannot do. Mods ship inside plugins and are installed with /plugin in the CLI or desktop app.

    Video from @latentspacepod's post
  6. TypeSafe AIOfficialAI score20

    DSPy office hours recording shares Jev use cases and explainers

    AIDSPy's official account highlighted a recorded office hours session showcasing Jev use cases and explainers. The quoted context says the session covered DSPy's Jev/System One implementation, the ReAnchor optimizer, and design patterns for selective compaction, tool approvals, and subagent delegation.

  7. Comfy BlogOfficialAI score47

    Hakoniwa uses Comfy Agent to make the animated short YUI

    AIArtist 852 Hakoniwa made YUI, described as the first animated short created with Comfy Agent, which the ComfyUI team says took three days of focused work by one person at about 200,000 yen in total cost, excluding labor. The source says the film was made mostly with Seedance 2.5 and the making-of video with MiniMax H3, with Comfy Agent used to regenerate shots and compare video models.

  8. TypeSafe AIOfficialAI score34

    Jev outperforms LLM judge for research agent risk monitoring at 250x lower cost

    AIIn a business risk monitoring test by @edwardirby, the Jev judge matched the report quality of an ordinary LLM judge while missing no investigations, versus the LLM missing 5 of 11. The LLM's threat scores also flip-flopped from 0.35 to 0.68 to 0.50 on the same threat, while Jev was 250x cheaper and 3-6x faster.

  9. Microsoft CopilotOfficialAI score34

    Microsoft Copilot adds GPT-6.1 Sol and Claude Sonnet 5.5 models

    AIMicrosoft Copilot begins rolling out OpenAI's GPT-6.1 Sol and Anthropic's Claude Sonnet 5.5 today, joining Claude Opus 5.5 and GPT-6 Sol added earlier this month. Users can pick the model suited to each task, with Work IQ grounding responses in their files, meetings, and chats within existing permissions. The rollout starts today in Copilot Cowork and Copilot Studio, with Word, Excel, PowerPoint, and Chat following in phases over the coming week.

  10. Grok BotOfficialAI score36

    Grok Bot can now suggest ways to help proactively

    AIGrok Bot can now offer suggestions for ways to help without the user needing to ask first. The post does not provide further details about how these proactive suggestions work.

    Video from @bot's post
  11. DatabricksOfficialAI score20

    Databricks Smart Routing assigns each coding task to a suitable model

    AIDatabricks' Smart Routing evaluates each coding task separately and selects the lowest-cost model capable of handling it, balancing quality, latency, and cost. In a demo, Omnigent splits an app build into planning, backend, and frontend work, routes each part to a different model, and runs some tasks in parallel.

    Video from @databricks's post
  12. Harrison ChaseXAI score22

    Harrison Chase outlines a four-step approach to model routing

    AIHarrison Chase says model routing is a provocative term that lacks a clear definition, but offers a practical approach. His four steps are to understand tasks, understand the models, build the router inside the harness, and track outcomes, aiming to lower costs without a performance hit.

    Image from @hwchase17's post
  13. Google WorkspaceOfficialAI score32

    Google Sheets canvas turns spreadsheets into interactive mini-apps via prompts

    AIGoogle Workspace says Sheets canvas can turn static spreadsheet data into interactive tools such as Kanban boards, dashboards, and visual workflows from a simple prompt. Derek Snyder, Director of Product Marketing for Google Workspace, demonstrates the feature in the latest AI Boost Bite video.

    Video from @GoogleWorkspace's post
  14. Prime IntellectOfficialAI score32

    Extropic uses Prime Intellect to post-train Qwen3.6-35B-A3B for thermodynamic ML

    AIExtropic post-trained Qwen3.6-35B-A3B with Prime Intellect for thermodynamic ML research, nearly tripling its held-out eval results in about 100 GRPO steps. The team built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This let Extropic avoid managing multi-node GPU infrastructure and focus on research.

    Image from @PrimeIntellect's post
  15. Josh WoodwardXAI score34

    Google launches Stitch CLI to generate design ideas from terminal

    AIGoogle has introduced the @google/stitch CLI, letting users generate screens and design systems without leaving the terminal. It connects to local coding agents and can send a local dev server snapshot to Stitch. The tool complements the existing Stitch MCP and SDK, and can also be driven through agents such as Antigravity.

  16. ComfyUIOfficialAI score20

    ComfyUI introduces Comfy Agent, an AI agent for creative workflows

    AIComfyUI announces Comfy Agent, which its post presents as the first agent built for creative work. The post directs readers to a blog for more details, but its own text provides no further specifics on capabilities, availability, or pricing.

  17. Jerry LiuXAI score42

    LlamaIndex launches Extract v2.5 document extraction agents with improved accuracy

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction across cost-effective, agentic, and agentic plus tiers. The company reports the agents outperform Opus 5.5 and GPT-6 Sol while costing 30% to 4x less, with accuracy gains on long lists (86.1% to 95.5%), multi-page records (85.5% to 96.5%), and scanned forms (90.9% to 95.7%) on its agentic tier. The release adds advanced citations with bounding boxes and structural reasoning, and the agents are available on LlamaParse.

    Video from @jerryjliu0's post
  18. Lewis Tunstall @ COLM 🌉XAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  19. merveXAI score46

    Hugging Face clarifies ml-intern options, one trained model for $6

    AIHugging Face says ml-intern is an open-source ML engineering and research harness usable free on local setups, and it is also hosted on Hugging Chat with no-code access. A second hosted option runs on Hugging Face infrastructure, where ml-intern selects the cheapest GPU for a task so models can be trained for a few dollars. MaziyarPanahi reportedly trained a model by prompting alone for $6.60 on an NVIDIA A100 in 16 minutes.

  20. Comfy BlogOfficialAI score44

    Comfy Agent Launches in ComfyUI Cloud, Desktop Version Coming Weeks Later

    AIComfy Agent, an AI agent that builds, runs, and iterates on workflows from plain-language requests, is now available in Comfy Cloud and will arrive in Comfy Desktop in a few weeks. It can work directly on the canvas alongside users, support up to 5 parallel chats, and use public or private skills. Comfy Agent is in beta and uses existing Comfy Credits.

  21. Philipp SchmidXAI score25

    Unlimited human usage paired with capped agent usage

    AIThe post notes a new pricing pattern where human use is unlimited while agent usage is limited. It remarks that this is the first time the author has seen such a split, without naming the product or provider.

    Image from @_philschmid's post
  22. OpenRouter · New modelsBlogAI score36

    Pareto 26.10 Preview: A Multimodal Model for Research, Coding and Agents

    AIPareto 26.10 Preview is a multimodal composite model built for research, coding, and agentic workflows. It is described as delivering frontier-level performance across a broad range of general-purpose tasks, though the source excerpt is a preview and provides no benchmark scores, parameter counts, pricing, or availability details.