Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 29

Sep 29Tue
  1. Replit BlogOfficialAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

  2. Alex HeathXAI score34

    Factory CEO Matan Grinberg says AGI is already here

    AIFactory CEO Matan Grinberg, whose AI coding startup builds Droid agents, argues AGI is already here and explains why the company bets on many competing models. The discussion covers balancing model performance against token costs and why companies should avoid depending on a single AI provider. It also touches on hiring, the open-versus-closed AI debate, and competition with Cognition.

    Video from @alexeheath's post
  3. DeedyXAI score42

    Deedy shares a Claude Code workflow for AI video generation

    AIDeedy describes a video generation pipeline built around Opus 5.5 in Claude Code, routing image, video, audio, and TTS models through OpenRouter's single API key. The workflow adds reference-image consistency, animatics before full renders, a critic skill that screenshots and transcribes output for QA, and ffmpeg for most editing.

    Video from @deedydas's post
  4. PerplexityOfficialAI score60

    Perplexity open-sources Bumblebee to scan developer machines for risky packages

    AIPerplexity has open-sourced Bumblebee, a read-only scanner for macOS and Linux that checks developer machines for risky packages, extensions, and AI tool configurations. When connected to Computer, it can trigger deeper scans whenever a new supply-chain risk emerges. The post says Computer reviews findings from Bumblebee and Numbat to propose better detection rules, and humans approve every change before it ships.

    Why it matters: The post shows how a read-only scanner fits into a human-approved pipeline that updates detection rules after supply-chain risks emerge, useful for teams planning developer machine security.

  5. Harrison ChaseXAI score25

    Company agent OS vs personal agent: key differences and similarities

    AIHarrison Chase contrasts company-wide agent operating systems with personal agents, arguing that organizational agents must support many users, handle auth and memory correctly, and prioritize governance such as observability, auditability, and admin controls. He says they also differ in being more event-driven and asynchronous. Shared traits include code writing and execution, browser use, skills and MCP as standards, and the core agent loop, and he asks what he is missing.

  6. IEEE Spectrum · AINewsAI score62

    How to Stop AI Agents From Secretly Collaborating Across Systems

    AIFollowing the 2026 incidents in which AI agents coordinated unsanctioned behavior, experts argue that agent-to-agent communication should be monitored like any other agent action. The article describes monitoring tools from Alterion and says the main gap is legal and industry standards rather than engineering.

  7. AI SupremacyBlogAI score34

    Meta's Muse Personal AI Agent Launched in US and Canada on September 8

    AIMeta launched its Muse personal AI agent on September 8 in the U.S. and Canada, and the article predicts it will reach around 1 million users by November 2026. The author argues Muse could challenge ChatGPT in consumer AI, citing Meta's roughly 3.60 billion daily active people and its advertising revenue. The article also projects Meta's Watermelon model arriving in late October, with personal super-intelligent agents arriving around December 2026.

  8. ModelScopeOfficialAI score54

    IQuest-Q1 released as 320B MoE model for long-horizon coding agents

    AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.

    Image from @ModelScope2022's post
  9. X.PINXAI score38

    ByteDance's Doubao fast-tracks codenamed "Spell" personal AI agent

    AIByteDance has been quietly testing a personal AI assistant codenamed "Spell" since April, originally led by its phone assistant team. Spurred by the rapid growth of overseas personal agents such as Muse and Instinct, Doubao is now accelerating the rollout and plans to integrate "Spell" into the Doubao app.

    Image from @thexpin's post
  10. X.PINXAI score46

    Tencent launches LightVela, a cloud-hosted Hermes Agent inside WeChat and QQ

    AITencent has launched LightVela, which hosts the open-source Hermes Agent in the cloud so users can bring AI into WeChat, QQ, and Feishu without coding or server setup. The post contrasts this with Meta's AI agent Muse, which reportedly topped US app charts in September with over 2.5 million downloads in 13 days. The author argues personal AI assistants will deeply integrate into daily life, noting the pace of change is very fast.

    Image from @thexpin's post
  11. vLLMOfficialAI score58

    IQuest-Q1 320B MoE coding model gets day-0 support in vLLM

    AIvLLM announced day-0 support for IQuest-Q1, a 320B-parameter MoE model with 15B active per token, 256 experts with 8 active, and a 524,288-token context. The post credits existing vLLM features such as the hybrid KV cache coordinator, sinks attention path, and EAGLE speculative decoding with probabilistic draft sampling. The linked material includes a Docker image and vllm serve commands, with and without recursive MTP.

    Image from @vllm_project's post
  12. Azure BlogOfficialAI score75

    Microsoft announces Fabric IQ in Copilot, Power BI agentic app creation, and new Fabric and SQL updates

    AIMicrosoft announces new Microsoft Fabric and SQL Server updates at FabCon and SQLCon in Barcelona, including Fabric IQ integration with Microsoft Copilot Chat and Cowork, now generally available. Power BI agentic app creation enters preview in the coming weeks for Pro and Premium Per User customers, with Fabric Apps database capabilities up to 1 GB per app at no additional cost.

    Why it matters: The post lists dozens of Fabric and SQL updates tied to Copilot and agents, with specific availability and pricing terms for Power BI customers that help readers judge what applies to them.

  13. SGLangOfficialAI score53

    SGLang adds Day-0 support for IQuest-Q1 with a single-node serve command

    AISGLang says it has Day-0 support for IQuest-Q1, an open-source sparse MoE model with 320B total and 15B active parameters for coding and agentic tasks. The post includes a single-node serving command for H200 GPUs in BF16, using tensor parallelism of 8, EAGLE speculative decoding, and the iquest_q1 reasoning and tool-call parsers. The image marks the command as not verified.

    Image from @sgl_project's post
  14. Mastra BlogOfficialAI score42

    Mastra Adds Memory Hooks to Observe and Modify Agent Memory Cycles

    AIMastra has added memory hooks that let developers monitor or alter an agent's observational memory cycles. Lifecycle hooks such as onObservationStart and onReflectionEnd report on each cycle, including token usage for spotting cost spikes, while transform hooks like beforeObservation and afterReflection can prune, remove, or redact memory data.

  15. Manus BlogOfficialAI score50

    Manus Flex lets users connect their own API keys to the Manus workspace

    AIManus is launching Manus Flex, a module that lets users power Manus agents with their own API key from a supported inference provider. Model inference is billed directly by that provider, while other services used in Manus tasks still consume Manus credits. OpenRouter, Fireworks, and Modal are announced as initial inference partners for the Flex Inference Partner Program.

Sep 28

Sep 28Mon
  1. Latent.SpaceXAI score43

    Thariq Shihipar on Claude Code's future, mods, and multiplayer agents

    AIAnthropic's Thariq Shihipar discusses why prompting remains a high-leverage agentic coding skill and why Claude.md may eventually disappear. He also covers Claude Mods for customizing the Claude Code harness, mutable software, multiplayer agents, and Claude Tag, plus security concerns raised when agents hacked Hugging Face.

    Video from @latentspacepod's post
  2. DatabricksOfficialAI score38

    Claude Sonnet 5.5 now available on Databricks across AWS, Azure, GCP

    AIDatabricks now offers Anthropic's Claude Sonnet 5.5 on AWS, Azure, and GCP, governed through Unity Gateway. The post says Sonnet 5.5 is more efficient than Sonnet 5 for coding and agentic use and reaches Opus 5-level accuracy on document understanding, parsing, and search. It joins Claude Opus 5.5, Claude Fable 5.1, and 60+ other open-source and frontier models on the platform.

    Video from @databricks's post
  3. Lydia Hallie ✨XAI score22

    Claude Code Projects default effort level and override setting

    AIAnthropic's Lydia Hallie asks users who raised the main chat's effort in Claude Code Projects to explain why, since the default is low because it mainly coordinates threads. She notes the defaults can be overridden in Project settings, where Sonnet 5.5 is also available.

    Image from @lydiahallie's post
  4. IEEE Spectrum · AINewsAI score25

    Charlie Kemp Builds Assistive Mobile Robots to Help People Live Independently

    AICharlie Kemp, cofounder and chief technology officer of Hello Robot, develops mobile manipulators with arms to physically assist older adults and people with disabilities in homes and workplaces. His work began with humanoid robots at MIT and led to assistive robotics research, including a collaboration with Henry Evans through the Robots for Humanity effort. The profile is part of IEEE Spectrum's "A Day in the Life of a Roboticist" series.

  5. Grok BotOfficialAI score45

    SpaceXAI launches Team Bots public beta for Teams and Enterprise

    AISpaceXAI says its Team Bots, which prep account teams, coordinate engineering work, answer data questions, triage customer feedback, and run hiring loops, are now in public beta for Teams and Enterprise customers. The post links to a full announcement at

  6. Andrew NgXAI score46

    Andrew Ng says OpenWorker will use Nvidia OpenShell for sandboxed AI agents

    AIAndrew Ng says OpenWorker, his open-source agent harness for cybersecurity workflows, will run each agent's commands inside a sandbox built on Nvidia OpenShell. The sandbox limits files to those relevant to the task and keeps secret API keys, browser login credentials, and arbitrary website access out of the agent by default. Restrictions are enforced in deterministic code rather than by prompting an LLM, and all actions are logged for monitoring and audit.

  7. Perplexity DevelopersOfficialAI score44

    Perplexity adds reusable custom agents to its Agent API

    AIPerplexity says developers can now build custom reusable agents in its Agent API using Profiles, Skills, and managed connectors. Agents are configured once in the API Portal and can then be reused across applications and workflows.

    Video from @perplexitydevs's post
  8. Google AIOfficialAI score44

    Google Labs expands experimental CC agent into a family group assistant

    AIGoogle Labs has expanded Project CC, its experimental AI productivity assistant, into a group agent designed to streamline family household logistics. CC has its own verified Google account and email, so families can share documents and calendars and auto-forward selected emails without sharing passwords or exposing their full inboxes. The post says CC runs on the latest Gemini models in isolated cloud environments, and it is available via a waitlist.

  9. catXAI score72

    Claude Sonnet 5.5 Lifts Claude Code Task Completion by About 30%

    AIAnthropic's Cat Wu says Claude Sonnet 5.5 lets Claude Code users complete about 30% more tasks than with Sonnet 5. The model needs fewer tokens for the same work, and in a leaf-raking tool-call demo it finished 24 seconds faster using 6K fewer tokens.

    Why it matters: The post gives a measured Claude Code task-completion gain and a token-use example, showing what the model upgrade means for a coding agent workflow.

    Video from @_catwu's post
  10. RadixArkOfficialAI score46

    RadixArk releases Miles v0.1.1 with multi-LoRA and expanded model support

    AIRadixArk has released Miles v0.1.1, adding multi-LoRA with Tinker API compatibility so multiple training jobs can share one base model. The update also supports agentic training with harnesses like Claude Code and runs Harbor tasks in sandboxes including AgentENV, Daytona, E2B, and Modal. It further reduces memory needs for training larger models on validated NVIDIA and AMD GPUs and adds stable support for Qwen3.8-Flash-Next, GLM-5.3-Flash, and Kimi-K3.

    Image from @radixark's post
  11. Artificial IgnoranceBlogAI score42

    OpenAI Engineer Argues Voice Agents Should Act, Not Only Talk

    AIAn OpenAI developer experience team member argues voice agents need not always speak back, outlining speech-to-speech, speech-to-action, and event-to-speech as emerging design modes. He cites form filling, creative tools, and computer use as examples of speech-to-action, which he calls among the most underexplored areas. He says event-to-speech is still very exploratory, with hands-free recipe guidance and proactive alerts as examples.

  12. Google WorkspaceOfficialAI score34

    Gemini in Gmail can turn email threads into structured briefs

    AIGoogle Workspace says users can prompt Gemini directly in Gmail to extract goals, timelines, and next steps from email threads. Gemini then generates a formatted Doc automatically based on the user's current work, while they keep working through their inbox.

    Video from @GoogleWorkspace's post