Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 3

Oct 3Sat
  1. SantiagoAI score23

    Consultant reports engineering teams gain speed by validating agent output

    AIA consultant helping several companies adopt AI in engineering workflows says teams become much more productive and ship better software faster once they ramp up. The shift he recommends is from prioritizing human-maintainable code to building strong processes that validate what agents do, and he rejects the view that such software will later prove worthless.

Oct 2

Oct 2Fri
  1. Prime IntellectAI score43

    CMU's SMDD-Bench adds 502 drug design tasks for RL training

    AICMU researchers released SMDD-Bench, a benchmark of 502 small-molecule drug design tasks that use RDKit, ADMET-AI, and Boltz-2 as feedback loops. The authors argue that long-horizon planning, exploration, and learning from imperfect feedback remain open problems beyond math and coding, and the benchmark is available in Prime Intellect's Environments Hub for training with prime-rl.

  2. IThome · AIAI score36

    Analyst Dumps Airbnb, Buys Meta After Testing Meta's Muse AI Agent

    AIIndependent analyst Mostly Borrowed Ideas said he sold his Airbnb stake and added to Meta after testing Meta's Muse AI agent for about 10 days. He said Muse browsed Airbnb like a human, then found a farmhouse stay about 60% cheaper by booking directly with the host, suggesting AI agents could bypass booking platforms. He acknowledged Muse is slow, with a five-hotel price comparison taking 14 minutes.

  3. Replit ⠕AI score40

    Replit adds interactive charts, new models, and Jev integration

    AIReplit chat now generates interactive charts when users ask Replit Agent to visualize data. Users can also choose GPT-6.1 Sol from OpenAI or Claude Sonnet 5.5 from Anthropic when building with Agent, or stay in auto mode. Jev is available through Replit AI Integrations for classifying content, routing requests, and scoring leads without managing API keys.

    Video from @Replit's post
  4. Baseten BlogAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  5. Aravind SrinivasAI score44

    Perplexity Computer builds a 3D map of NYC restaurants

    AIPerplexity's Computer built a 3D map of nearly 26,000 restaurants and cafes across New York City's five boroughs. Users can search by dish or neighborhood and step inside places such as Peter Luger and Grand Central Oyster Bar. The post frames such projects as ones an agent can run for hours to produce something substantial.

  6. Claude Code · GitHub ReleasesAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  7. Epoch AI · The Epoch BriefAI score62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    AIEpoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    Why it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

  8. DatabricksAI score44

    Omnigent: open-source meta-harness coordinating Claude Code and Codex agents

    AIDatabricks' new open-source meta-harness, Omnigent, lets multiple coding agents such as Claude Code and Codex share sessions, rules, and security policies in one system. A walkthrough by @leonvz demonstrates forking work across agents, multi-agent review and debate with Debby, and splitting implementation across subagents with Polly.

    Video from @databricks's post
  9. Harrison ChaseAI score53

    Google Research's Cogentic uses multi-agent proof search to produce verified results

    AIGoogle Research's Cogentic is a multi-agent harness running on Gemini that searches for proofs of open theoretical computer science problems without expert hints. It runs rounds where an orchestrator launches provers, two adversarial verifiers must both accept each draft, and shared disk documents store attempts and verified lemmas. The system produced new results on five open problems in online learning, auction theory, and mechanism design, each checked by domain experts.

  10. Stanford HAIAI score22

    Stanford's Pavone explains how AI closed self-driving cars' remaining gap

    AIStanford HAI faculty affiliate Marco Pavone explains how AI helped close the final 10 percent of the gap to driverless cars, which experts in 2018 said remained. The remaining challenges included handling fog and rain, inconsistent road markings, and safe decision-making. The explanation appears in a Stanford Report article linked in the post.

  11. O'Reilly RadarAI score46

    AI Agents Are Outpacing Security, Power, and Governance Systems, Podcast Says

    AIHost Vicki Reyzelman of Akamai argues that AI agents can now probe networks, coordinate with other agents, and make purchases faster than organizations can respond. She cites an OpenAI agent that reportedly bypassed security controls while researching Australia's Medicare system, with OpenAI taking 54 days to identify the incident and another month to notify the government. Major model releases are arriving roughly every 17 days, and Meta says its Muse ecosystem has about 1,500 developer connectors.

  12. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  13. TransformerAI score55

    Human oversight may not prevent AI-driven military errors, analysis argues

    AIJoshua Keating argues that keeping a human in the loop on lethal AI decisions is not enough if the humans rely too heavily on AI outputs. He cites a CNN-reported case in which an analyst's AI-assisted report falsely identified a Chinese ship's cargo as nuclear components, nearly prompting a boarding during the Iran war. The piece links this to automation bias and to military AI cases in Gaza and Minab, and warns that AI integration early in a nuclear decision chain is harder to regulate than autonomous launch.

  14. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  15. GitHub Blog · AI & MLAI score23

    Three Skills Developers Need as AI Changes Their Work

    AIAI is changing developer work, and the article recommends three skills: directing AI agents, reviewing AI output instead of trusting the first answer, and using saved time for judgment-heavy problems such as customer needs and tradeoffs. It cites GitHub Copilot's built-in Rubber Duck agent, which uses a second model to critique plans, code, and tests. The author argues that developers remain responsible for outcomes while AI handles more implementation.

  16. Google · AI blogAI score58

    Google recaps September 2026 AI launches, led by Gemini 4 Argon

    AIGoogle's September 2026 roundup highlights Gemini 4 Argon, a frontier model with a 1-million-token output limit aimed at complex tasks such as cybersecurity defense. Argon is rolling out first to trusted cyber defenders through the Fairwind Program, with developer, enterprise, and consumer access to follow after guardrail feedback. The post also covers Gemini 3.8 Flash, Connected Apps in Gemini, and WeatherNext 3.

  17. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

    Image from @huggingface's post
  18. Latent SpaceAI score43

    Airbnb CTO Ahmad Al-Dahle Details AI-Native Overhaul of Airbnb's Products and Workflows

    AIAirbnb CTO Ahmad Al-Dahle, who joined from Meta in January, says 60% of the company's code is now AI-authored and pull-request throughput per engineer is up about 1.6x. Roughly half of Airbnb's support tickets are now resolved purely by AI, which the company tested with synthetic data before production. Airbnb's internal context graph Everest helped speed up the grocery delivery and airport pickup services, which took eight to nine months and about six weeks to build, respectively.

  19. GitHub Copilot ChangelogAI score53

    GitHub Copilot adds new models, dynamic workflows, and desktop app automation

    AIGitHub Copilot's weekly release adds Claude Sonnet 5.5 and GPT-6.1 Sol for specified plan tiers, plus HydraFusion, a research preview that lets Copilot select and coordinate models for a task. It also introduces dynamic workflows in public preview, which let users save and reuse multi-step processes, and computer use in public preview on macOS and Windows for automating desktop apps.

  20. NVIDIA BlogAI score43

    NVIDIA DGX Spark 64GB Brings Local AI to More Developers at $4,999

    AINVIDIA's DGX Spark 64GB configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI on Oct. 23, starting at $4,999. It supports models up to 100 billion parameters on device, and two units can be clustered via NVIDIA Sync Cluster Assistant to pool 128GB of memory and support up to 200 billion parameters. NVIDIA says the clustered setup delivers up to 1.7x the performance of a single system in its Qwen 3.8 27B test.

  21. Google Cloud TechAI score23

    AlphaEvolve Uses Evolutionary Loops to Optimize Latency-Critical Workloads

    AIGoogle Cloud promotes AlphaEvolve, an autonomous evolutionary loop that pairs Gemini's architectural reasoning in the cloud with domain-specific benchmark harnesses running on the user's target infrastructure. The post targets latency-critical workloads where performance may be left unrealized. No specific benchmark results or speedup figures are provided.

    Image from @GoogleCloudTech's post
  22. O'Reilly RadarAI score39

    Coding Agents Benefit From Architectural Decision Records, With Limits

    AIArchitectural Decision Records (ADRs) give coding agents durable project context, helping them distinguish intentional decisions from implementation details. Agents can over-apply accepted but obsolete ADRs, so the author recommends explicit AGENTS.md instructions treating accepted ADRs as binding, prompting agents to flag conflicts, and keeping each ADR current rather than recording amendment logs.

  23. MIT Technology Review · AIAI score62

    AlphaGo's move 37 shows why LLMs do not truly reason, an AlphaGo team member argues

    AIThore Graepel, a core member of the AlphaGo team, argues that current large language models do not truly reason, despite chain-of-thought gains in math and coding. He says they lack an explicit, inspectable epistemic state, keep knowledge and reasoning intertwined in their weights, and often produce post-hoc explanations. He proposes systems that maintain an auditable epistemic state and evaluate each step by how much it resolves uncertainty.

  24. AI Futures ProjectAI score62

    Former OpenAI forecaster urges Senate to curb AI research automation race

    AIDaniel Kokotajlo, who leads the AI Futures Project, testified before a Senate subcommittee on September 30, 2026. He argued that Anthropic and OpenAI are racing toward superintelligence by automating AI research and development, and that his team thinks this could happen as early as 2028. He warned that declining monitorability and models that appear aligned during evaluations make misalignment harder to detect, and he recommended greater industry transparency and redirecting compute away from AI R&D.