Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 5

Oct 5Mon
  1. indigoXAI score42

    Five-step Grok Bot method for hiring and managing AI agents

    AIBrian's Grok Bot method treats each bot like a new hire: define the role, test it on text first, run three trials, escalate based on evidence, and add a second agent only after a bottleneck appears. Each bot's role is defined by five fields: a real name with a short label, a one-line job tied to an outcome, what it owns, its inputs, and what it may do freely versus what it must ask before doing. The post frames an Agent Team as the final result of this process, starting with one coordinator and three specialists.

    Image from @indigox's post
  2. meng shaoXAI score72

    Uber Designs an MCP Gateway to Expose Thousands of Internal APIs to AI Agents

    AIUber uses a control plane and data plane gateway to automatically convert its internal APIs into MCP tools, with 800+ MCP servers and 5,000+ tools hosted. The design includes an AutoCrawler that generates tool descriptions with an LLM, a default-disabled discover-not-expose security model, and techniques such as Omni MCP, Response Projection, and Code Mode to limit context bloat.

    Why it matters: The article details how Uber converts thousands of internal APIs into MCP tools, including discovery, permission, and context-size tactics that transfer to other enterprise agent deployments.

    Image from @shao__meng's post
  3. EveryBlogAI score22

    When Trying to Make AI Better Makes It Worse

    AIThe article argues that improving an AI setup can sometimes mean giving the AI fewer rules to follow, based on the author's experience across a million words of failed drafts. The source text provided is mostly paywall and subscription material, so no further specific figures, products, or benchmarks can be verified.

Oct 4

Oct 4Sun
  1. meng shaoXAI score44

    Baschez argues shared AI factories will outperform personal AI agents

    AINathan Baschez argues that the end state of AI work is not individual employees running personal agents like Codex or Claude Code, but shared, specialized "AI factories." He contends factories beat personal agents because they are shared, task-specific, and scrutinized, which creates feedback loops for systematic improvement. In a 100-person consulting firm comparison, concentrating about 18.3 hours of AI tuning per task on 1–2 tasks gives 5 times deeper learning than spreading 3.7 hours across 5–10 tasks.

    Image from @shao__meng's post
  2. Together AI BlogOfficialAI score38

    Together Link Routes Coding Agents to Open Models, Cutting Spend Over 50%

    AITogether Link connects coding agents such as Claude Code, Codex, OpenCode, and Pi to open models on Together AI, which the company says cuts spend by over 50%. Setup takes one command, and its "Auto" mode routes each session's first task to a fast low-cost model or a frontier model, with a per-session tracker comparing costs against Opus 5.5.

  3. PromptArmor Threat IntelligenceOfficialAI score47

    Databricks Genie Code Malicious Skill Enables Phishing and Data Exfiltration

    AIPromptArmor reports that a malicious Skill can make Databricks Genie Code display a phishing modal and exfiltrate tenant data without human approval. The attack exploits Skills loaded from users' personal workspaces and a display interface that lacks egress controls, and Databricks, after disclosure on August 16, 2026, said users are responsible for ensuring uploaded Skills contain no malicious content.

  4. Epoch AIOfficialAI score62

    OpenAI researchers' coding-agent usage is doubling about monthly, Epoch AI reports

    AIOpenAI researchers' daily coding-agent usage, valued at API prices, rose from under $1 in January 2026 to $601 for the median researcher by mid-August. The 90th-percentile researcher reached over $7,000 per day, and both groups show doubling times of roughly one month. Epoch notes these are API-list values, not OpenAI's internal costs.

    Why it matters: The figures show internal coding-agent usage growing fast enough to matter for research cost, though they measure API-list value rather than OpenAI's actual spending.

  5. Boris PowerXAI score40

    GPT-6 Astra tops Design Arena's 3D Design leaderboard in its first month

    AIBoris Power says GPT-6 can work autonomously on 3D design for hours while its results keep improving, a gap other models failed to match because they could not recover from mistakes. Design Arena reports GPT-6 Astra took #1 on four leaderboards, including 3D Design at 1484 and Frontend at 1397, a month after release.

  6. Guillermo RauchXAI score38

    fx.sh gets much faster as harness overhead matters more

    AIfx.sh has become much faster, with the v0.0.13 release reporting launches up to 23× faster, shell calls up to 8.6× faster, and exits up to 44× faster. Guillermo Rauch says that as models like Astra ultrafast speed up, harness overhead matters more, and the next release will improve session storage and retrieval.

  7. Jerry LiuXAI score23

    Jerry Liu Says ChatGPT/Codex Offers Best Agent Interface for Deep Work

    AIJerry Liu says ChatGPT/Codex currently has the best agent interface for deep work, unifying coding and knowledge work in one place with forking support that Claude's app lacks. He still prefers Claude Code CLI as close to the best a CLI can be, and uses Opus 5.5 mainly through it for product demos, while noting a GUI is sometimes nicer.

  8. Teknium 🪽XAI score20

    Teknium posts two eyes emojis, teasing an unexplained Hermes-related announcement.

    AITeknium, a researcher associated with the Hermes model family, posted only two eye emojis with no explanation of the main post's content. The post appears to be a teaser linked to a quoted post from @alexhvnsen describing a Hermes "Alan's way" companion app with Telegram-based control, a macOS and VM hybrid setup, and a proactive lead bot.

  9. Aravind SrinivasXAI score40

    Perplexity Computer builds custom GeoGuessr-style image location app

    AIPerplexity's Computer can build custom vertical AI apps, such as one that guesses an image's location using 3D and satellite views. The example app was built with the Perplexity SDK for web search, local place lookups, and visual clue extraction, and it uses Cesium for the 3D globe and satellite imagery.

  10. The SequenceBlogAI score57

    The Sequence reviews weekly AI news on agents, funding, and model releases

    AIThe Sequence's Issue 944 recaps a week of AI news, including OpenAI's Dots persistent agents and GPT-6.1 Sol, Google's Gemini 4 Argon, and Meta's Meta Enterprise Platform. It also covers AMD's roughly $8.2 billion all-stock deal for World Labs and Instinct's $1 billion Series C at a $10 billion valuation. The editorial argues that delegation to AI agents is the common theme across these announcements.

  11. Harrison ChaseXAI score31

    LangChain cuts coding agent costs with tracking, caps, and routing

    AILangChain says its coding agent costs fell significantly for a second straight month after adopting three steps. The steps are cost visibility through LangSmith tracing, per-user cost caps via its LLM gateway, and harness optimization such as model routing in its open-source OpenSWE cloud agent harness.

    Image from @hwchase17's post
  12. Aravind SrinivasXAI score20

    Perplexity's Decisions API clears Pokémon FireRed's Elite Four in one run

    AIPerplexity's Decisions API powered the decision-making in a one-shot clear of Pokémon FireRed's Elite Four and Champion. The run recorded a 592 ms median API response time, a 987 ms p95, and 96.4% of responses under one second. Estimated inference cost was $0.028 across 137 live API calls.

    Video from @AravSrinivas's post
  13. Yuchen JinXAI score23

    Yuchen Jin says AI agents are replacing terminals as the coding interface

    AIYuchen Jin argues that terminals, built around files, commands, and processes, are giving way to AI agents where users state intent and the agent operates the machine. He says understanding Linux and systems fundamentals remains valuable as a moat. In a follow-up, he calls the terminal era over for coding agents, saying persistent context matters more than tabs, and names the Codex desktop app as the best agentic UI for now.

  14. EveryBlogAI score57

    Dan Shipper Reviews OpenAI DevDay 2026 Releases for ChatGPT as Work OS

    AIOpenAI wants ChatGPT to become an operating system for work, and Dan Shipper sorted its 22 DevDay 2026 releases by how much each advances that goal. The five most important include Dots, an always-on agent, and Space, native documents the agent can edit, which form the workspace itself. After a week of use, Shipper concluded the ambition is big but the execution is not there yet, and even power users have a lot to figure out.

Oct 3

Oct 3Sat
  1. Hugging Face BlogOfficialAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

  2. Amjad MasadXAI score38

    Replit CEO proposes general AI models train smaller domain-specific replacements

    AIReplit CEO Amjad Masad argues that general models could train smaller, domain-specific successors on the fly when they detect a limited use case. He compares this to a just-in-time compiler that emits optimized code during execution. He says such specialized models could be cheaper, less vulnerable to prompt injection, and less harmful than general agents.

  3. Aravind SrinivasXAI score34

    Perplexity Computer adds inline interactive visualizations on request

    AIPerplexity's Computer can now generate inline visualizations when users ask it to "Visualize" a topic, producing interactive widgets and animations within the thread. The feature is best used on Standard or High effort, and an example given is an inline 3D cutaway of a jet engine.

  4. Amjad MasadXAI score42

    Amjad Masad and Alex Atallah discuss AI independence and specialized agents

    AIAmjad Masad of Replit and Alex Atallah of OpenRouter discuss why AI independence and model diversification matter for enterprises. They argue that depending on a single lab risks lock-in and that specialized agents may outperform one general superagent. The post presents the conversation as a podcast episode, the first Atallah has done since Stripe acquired OpenRouter.

  5. Yuchen JinXAI score22

    Yuchen Jin says terminals are wrong for coding agents

    AIYuchen Jin argues that the terminal is the wrong interface for coding agents, since managing many tabs creates cognitive overhead while context should persist. He says he rarely needs an IDE like Cursor because he seldom navigates the whole codebase now, calling the agent rather than the file the new primitive. He names the Codex desktop app as the best agentic UI for now, while noting the space is still early.

  6. PixVerseOfficialAI score22

    PixVerse launches Short Drama plugin for AI agent video creation

    AIPixVerse has introduced Short Drama, a plugin that lets an AI agent turn a user's scene description into a finished video. The plugin takes text covering characters, setting, action, and mood, and produces video through the agent workflow, giving teams concrete material to review and develop.

    Video from @PixVerse's post
  7. SantiagoXAI score23

    Consultant reports engineering teams gain speed by validating agent output

    AIA consultant helping several companies adopt AI in engineering workflows says teams become much more productive and ship better software faster once they ramp up. The shift he recommends is from prioritizing human-maintainable code to building strong processes that validate what agents do, and he rejects the view that such software will later prove worthless.

Oct 2

Oct 2Fri
  1. Prime IntellectOfficialAI score43

    CMU's SMDD-Bench adds 502 drug design tasks for RL training

    AICMU researchers released SMDD-Bench, a benchmark of 502 small-molecule drug design tasks that use RDKit, ADMET-AI, and Boltz-2 as feedback loops. The authors argue that long-horizon planning, exploration, and learning from imperfect feedback remain open problems beyond math and coding, and the benchmark is available in Prime Intellect's Environments Hub for training with prime-rl.

  2. IThome · AINewsAI score36

    Analyst Dumps Airbnb, Buys Meta After Testing Meta's Muse AI Agent

    AIIndependent analyst Mostly Borrowed Ideas said he sold his Airbnb stake and added to Meta after testing Meta's Muse AI agent for about 10 days. He said Muse browsed Airbnb like a human, then found a farmhouse stay about 60% cheaper by booking directly with the host, suggesting AI agents could bypass booking platforms. He acknowledged Muse is slow, with a five-hotel price comparison taking 14 minutes.

  3. Replit ⠕OfficialAI score40

    Replit adds interactive charts, new models, and Jev integration

    AIReplit chat now generates interactive charts when users ask Replit Agent to visualize data. Users can also choose GPT-6.1 Sol from OpenAI or Claude Sonnet 5.5 from Anthropic when building with Agent, or stay in auto mode. Jev is available through Replit AI Integrations for classifying content, routing requests, and scoring leads without managing API keys.

    Video from @Replit's post
  4. Prime IntellectOfficialAI score38

    GLM-5.3 served on GB200 NVL72 at 100+ tokens/s per user

    AIPrime Intellect served GLM-5.3 on GB200 NVL72 while targeting 100+ end-to-end tokens per second per user for concurrent agent tasks. At that interactivity bar, a 1:4 prefill-to-decode ratio delivered the most throughput, supporting 66 sessions per prefill group at 101 tokens/s per user and 100 output tokens/s per GPU.

    Image from @PrimeIntellect's post
  5. Prime IntellectOfficialAI score23

    Prime Intellect optimizes long-context agent serving across three paths

    AIPrime Intellect says long-context agent serving depends on retaining history, scheduling new work, and moving cached state efficiently. It optimized three paths separately: prefill topology and scheduling, compressed KV with a fused attention kernel, and a transfer-friendly cache layout.