Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Oct 4

Oct 4Sun
  1. Guillermo RauchXAI score38

    fx.sh gets much faster as harness overhead matters more

    AIfx.sh has become much faster, with the v0.0.13 release reporting launches up to 23× faster, shell calls up to 8.6× faster, and exits up to 44× faster. Guillermo Rauch says that as models like Astra ultrafast speed up, harness overhead matters more, and the next release will improve session storage and retrieval.

  2. Jerry LiuXAI score23

    Jerry Liu Says ChatGPT/Codex Offers Best Agent Interface for Deep Work

    AIJerry Liu says ChatGPT/Codex currently has the best agent interface for deep work, unifying coding and knowledge work in one place with forking support that Claude's app lacks. He still prefers Claude Code CLI as close to the best a CLI can be, and uses Opus 5.5 mainly through it for product demos, while noting a GUI is sometimes nicer.

  3. Teknium 🪽XAI score20

    Teknium posts two eyes emojis, teasing an unexplained Hermes-related announcement.

    AITeknium, a researcher associated with the Hermes model family, posted only two eye emojis with no explanation of the main post's content. The post appears to be a teaser linked to a quoted post from @alexhvnsen describing a Hermes "Alan's way" companion app with Telegram-based control, a macOS and VM hybrid setup, and a proactive lead bot.

  4. Aravind SrinivasXAI score40

    Perplexity Computer builds custom GeoGuessr-style image location app

    AIPerplexity's Computer can build custom vertical AI apps, such as one that guesses an image's location using 3D and satellite views. The example app was built with the Perplexity SDK for web search, local place lookups, and visual clue extraction, and it uses Cesium for the 3D globe and satellite imagery.

  5. PixVerseOfficialAI score18

    PixVerse launches ad variants plugin for AI agents to generate campaign versions

    AIPixVerse has introduced ad variants, a plugin that lets AI agents turn one existing ad into multiple versions. Users can swap talent, outfit, product, or background while keeping framing, camera movement, timing, and lighting locked. The plugin is aimed at producing new variants for different markets, seasons, and audiences.

    Video from @PixVerse's post
  6. The SequenceBlogAI score57

    The Sequence reviews weekly AI news on agents, funding, and model releases

    AIThe Sequence's Issue 944 recaps a week of AI news, including OpenAI's Dots persistent agents and GPT-6.1 Sol, Google's Gemini 4 Argon, and Meta's Meta Enterprise Platform. It also covers AMD's roughly $8.2 billion all-stock deal for World Labs and Instinct's $1 billion Series C at a $10 billion valuation. The editorial argues that delegation to AI agents is the common theme across these announcements.

  7. Harrison ChaseXAI score31

    LangChain cuts coding agent costs with tracking, caps, and routing

    AILangChain says its coding agent costs fell significantly for a second straight month after adopting three steps. The steps are cost visibility through LangSmith tracing, per-user cost caps via its LLM gateway, and harness optimization such as model routing in its open-source OpenSWE cloud agent harness.

    Image from @hwchase17's post
  8. Aravind SrinivasXAI score20

    Perplexity's Decisions API clears Pokémon FireRed's Elite Four in one run

    AIPerplexity's Decisions API powered the decision-making in a one-shot clear of Pokémon FireRed's Elite Four and Champion. The run recorded a 592 ms median API response time, a 987 ms p95, and 96.4% of responses under one second. Estimated inference cost was $0.028 across 137 live API calls.

    Video from @AravSrinivas's post
  9. Yuchen JinXAI score23

    Yuchen Jin says AI agents are replacing terminals as the coding interface

    AIYuchen Jin argues that terminals, built around files, commands, and processes, are giving way to AI agents where users state intent and the agent operates the machine. He says understanding Linux and systems fundamentals remains valuable as a moat. In a follow-up, he calls the terminal era over for coding agents, saying persistent context matters more than tabs, and names the Codex desktop app as the best agentic UI for now.

  10. EveryBlogAI score57

    Dan Shipper Reviews OpenAI DevDay 2026 Releases for ChatGPT as Work OS

    AIOpenAI wants ChatGPT to become an operating system for work, and Dan Shipper sorted its 22 DevDay 2026 releases by how much each advances that goal. The five most important include Dots, an always-on agent, and Space, native documents the agent can edit, which form the workspace itself. After a week of use, Shipper concluded the ambition is big but the execution is not there yet, and even power users have a lot to figure out.

Oct 3

Oct 3Sat
  1. Claude Code · GitHub ReleasesOfficialAI score7

    Claude Code v2.1.289 fixes plugin, sandbox, and terminal rendering bugs

    AIClaude Code v2.1.289 fixes a series of bugs, including deny and ask rules being bypassed on nested parts of compound shell commands on managed machines. It also fixes terminal freezes on short code blocks with unclosed tags, Read deny rules not applying to files reached through symlinks in the IDE, and plugin panes that drew nothing for certain link formats. A change to claude auth status that may have increased sign-outs in VSCode was reverted.

  2. Hugging Face BlogOfficialAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

  3. Amjad MasadXAI score38

    Replit CEO proposes general AI models train smaller domain-specific replacements

    AIReplit CEO Amjad Masad argues that general models could train smaller, domain-specific successors on the fly when they detect a limited use case. He compares this to a just-in-time compiler that emits optimized code during execution. He says such specialized models could be cheaper, less vulnerable to prompt injection, and less harmful than general agents.

  4. Aravind SrinivasXAI score34

    Perplexity Computer adds inline interactive visualizations on request

    AIPerplexity's Computer can now generate inline visualizations when users ask it to "Visualize" a topic, producing interactive widgets and animations within the thread. The feature is best used on Standard or High effort, and an example given is an inline 3D cutaway of a jet engine.

  5. Amjad MasadXAI score42

    Amjad Masad and Alex Atallah discuss AI independence and specialized agents

    AIAmjad Masad of Replit and Alex Atallah of OpenRouter discuss why AI independence and model diversification matter for enterprises. They argue that depending on a single lab risks lock-in and that specialized agents may outperform one general superagent. The post presents the conversation as a podcast episode, the first Atallah has done since Stripe acquired OpenRouter.

  6. Yuchen JinXAI score22

    Yuchen Jin says terminals are wrong for coding agents

    AIYuchen Jin argues that the terminal is the wrong interface for coding agents, since managing many tabs creates cognitive overhead while context should persist. He says he rarely needs an IDE like Cursor because he seldom navigates the whole codebase now, calling the agent rather than the file the new primitive. He names the Codex desktop app as the best agentic UI for now, while noting the space is still early.

  7. PixVerseOfficialAI score22

    PixVerse launches Short Drama plugin for AI agent video creation

    AIPixVerse has introduced Short Drama, a plugin that lets an AI agent turn a user's scene description into a finished video. The plugin takes text covering characters, setting, action, and mood, and produces video through the agent workflow, giving teams concrete material to review and develop.

    Video from @PixVerse's post
  8. SantiagoXAI score23

    Consultant reports engineering teams gain speed by validating agent output

    AIA consultant helping several companies adopt AI in engineering workflows says teams become much more productive and ship better software faster once they ramp up. The shift he recommends is from prioritizing human-maintainable code to building strong processes that validate what agents do, and he rejects the view that such software will later prove worthless.

Oct 2

Oct 2Fri
  1. Prime IntellectOfficialAI score43

    CMU's SMDD-Bench adds 502 drug design tasks for RL training

    AICMU researchers released SMDD-Bench, a benchmark of 502 small-molecule drug design tasks that use RDKit, ADMET-AI, and Boltz-2 as feedback loops. The authors argue that long-horizon planning, exploration, and learning from imperfect feedback remain open problems beyond math and coding, and the benchmark is available in Prime Intellect's Environments Hub for training with prime-rl.

  2. IThome · AINewsAI score36

    Analyst Dumps Airbnb, Buys Meta After Testing Meta's Muse AI Agent

    AIIndependent analyst Mostly Borrowed Ideas said he sold his Airbnb stake and added to Meta after testing Meta's Muse AI agent for about 10 days. He said Muse browsed Airbnb like a human, then found a farmhouse stay about 60% cheaper by booking directly with the host, suggesting AI agents could bypass booking platforms. He acknowledged Muse is slow, with a five-hotel price comparison taking 14 minutes.

  3. Replit ⠕OfficialAI score40

    Replit adds interactive charts, new models, and Jev integration

    AIReplit chat now generates interactive charts when users ask Replit Agent to visualize data. Users can also choose GPT-6.1 Sol from OpenAI or Claude Sonnet 5.5 from Anthropic when building with Agent, or stay in auto mode. Jev is available through Replit AI Integrations for classifying content, routing requests, and scoring leads without managing API keys.

    Video from @Replit's post
  4. Prime IntellectOfficialAI score38

    GLM-5.3 served on GB200 NVL72 at 100+ tokens/s per user

    AIPrime Intellect served GLM-5.3 on GB200 NVL72 while targeting 100+ end-to-end tokens per second per user for concurrent agent tasks. At that interactivity bar, a 1:4 prefill-to-decode ratio delivered the most throughput, supporting 66 sessions per prefill group at 101 tokens/s per user and 100 output tokens/s per GPU.

    Image from @PrimeIntellect's post
  5. Prime IntellectOfficialAI score23

    Prime Intellect optimizes long-context agent serving across three paths

    AIPrime Intellect says long-context agent serving depends on retaining history, scheduling new work, and moving cached state efficiently. It optimized three paths separately: prefill topology and scheduling, compressed KV with a fused attention kernel, and a transfer-friendly cache layout.

  6. Guillermo RauchXAI score34

    Muse Ships Open-Source ESP32 Firmware and Linux SDK for Gadgets

    AIMuse has released Muse Gadgets, an open-source ESP32 firmware and Linux SDK for building hardware devices that work with Muse. Developers can obtain an API token from gadgets.muse.ai and use a coding agent with the GitHub repo to build peripherals. Guillermo Rauch praised the team's rapid shipping.

  7. Baseten BlogOfficialAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  8. Aravind SrinivasXAI score44

    Perplexity Computer builds a 3D map of NYC restaurants

    AIPerplexity's Computer built a 3D map of nearly 26,000 restaurants and cafes across New York City's five boroughs. Users can search by dish or neighborhood and step inside places such as Peter Luger and Grand Central Oyster Bar. The post frames such projects as ones an agent can run for hours to produce something substantial.