Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Why it matters: The post pairs a cost-versus-intelligence chart with specs and a later open-weights date, so readers can judge the cost tradeoff against named competitor models.

    Image from @StepFun_ai's post

Sep 18

Sep 18Fri
  1. TinkerOfficialAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  2. Mark ZuckerbergXAI score38

    Meta opens developer access to build Muse connectors for its agent

    AIMeta is opening access for developers to build connectors for Muse, its agent platform. Developers supply the API, while Muse provides the agent, browser, and user context, so people can reach a service simply by asking and their agent handles the rest. New connectors are live today at

    Video from @finkd's post
  3. Mike KnoopXAI score30

    Mike Knoop wonders what an underscore.js equivalent for AI looks like

    AIMike Knoop asks what the underscore.js equivalent for AI would look like, noting that such programming primitives feel close. He adds that he barely reads or writes code anymore despite these emerging tools. The quoted post introduces Probably, a toy programming language built around Jev, where constructs like "feels," "match," and "while" let AI make decisions within ordinary code.

  4. Google AIOfficialAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

  5. One Useful Thing (Ethan Mollick)BlogAI score50

    Mollick says AI already does weeks of human work when guided, citing Zork and Eco library demos

    AIEthan Mollick says GPT-6 Astra and Fable 5.1 already enable transformative impact and can reliably handle weeks of human work when properly guided. He cites GPT-6 Astra turning the 1977 text adventure Zork into a 3D action-adventure game and Fable 5.1 reconstructing Umberto Eco's Milan library in 3D from videos, photos, and catalogues.

  6. Google ResearchOfficialAI score26

    Google Research's Matias says AI amplifies human curiosity in science

    AIGoogle Research VP Yossi Matias discussed on The Google Research Podcast how ambient AI and GenUI interfaces adapt to users' thinking, and how AI Co-Scientist can turn multi-year hypothesis generation into 3-day sprints. He argued that AI is meant to amplify researchers' curiosity and judgment rather than replace them.

    Video from @GoogleResearch's post
  7. Noam BrownXAI score34

    Noam Brown Says Air-Gapping May Not Fully Stop Misaligned AI Coordination

    AINoam Brown, OpenAI, says air-gapped machines may still coordinate through a hot-CPU temperature-sensor channel, illustrating that absolute isolation guarantees are hard to achieve. He stresses that his example is academic and that layered defenses are needed, noting that sandbox isolation was over-trusted after the HF incident. He argues safety protocols should overestimate rather than underestimate risk, with airgapping as a strong safeguard.

  8. Google for DevelopersOfficialAI score38

    Android Bench 2.0 tests AI models on multi-day engineering workflows

    AIGoogle has released Android Bench 2.0, an updated benchmark that evaluates AI models on long-horizon tasks such as building apps from scratch, migrating cross-platform codebases to Android, and making complex architectural transitions. The benchmark uses continuous completion scoring to show which tasks each model performs well on.

  9. GitHub Blog · AI & MLOfficialAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

  10. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

Sep 17

Sep 17Thu
  1. KrASIA · Big TechNewsAI score44

    Qianjue founder says robotics will have no single "ChatGPT moment"

    AIQianjue Technology founder Gao Haichuan argues that robotics will not see one breakthrough that suddenly lifts the whole industry, and he judges the company by deployment results rather than research papers. Qianjue, founded in 2023, has completed a Series A+ round worth a nine-figure RMB sum, with first orders coming from restaurant, cleaning, and hotel service robots. Gao says customers care about task completion, failure rates, and price rather than whether a predictive world model is used.

  2. Felix RiesebergXAI score40

    Claude builds a multiplayer game from a single request in demo

    AIAnthropic's Felix Rieseberg posted a 60-second demo showing Claude, with built-in Cowork and connected Artifacts, building a multiplayer game when asked. He notes users can also ask Claude to copy a shared Artifact so they can play with others on their own account.

    Image from @felixrieseberg's post
  3. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  4. AI at MetaOfficialAI score44

    Meta's Muse agent now available on Mac for local tasks

    AIMeta is rolling out Muse for Mac today, a personal agent that can complete tasks directly on the user's computer with explicit permission. Examples include organizing the downloads folder, finding lost files, and summarizing messages and notes, with more capabilities coming soon.

  5. LM StudioOfficialAI score44

    LM Studio adds session history search and @ session references

    AILM Studio's new Introspection feature lets its Bionic agent search its own session history, improving handling of long-term context across multiple compactions. Users can also reference other sessions directly in the composer with an @ mention.

    Video from @lmstudio's post
  6. AnthropicOfficialAI score38

    Anthropic and Adaptyv Bio launch protein design competition with 5,000 validated designs

    AIAnthropic is partnering with Adaptyv Bio on a protein design competition in which over 5,000 designs will be experimentally validated. Anthropic is providing up to $1 million in Claude credits plus funding for experimental validation alongside Adaptyv, while Modal contributes up to $250,000 in compute and Twist Bioscience supplies DNA.

  7. Boris ChernyXAI score45

    Claude Code adds Projects for parallel cloud coding sessions

    AIBoris Cherny says Projects in Claude Code have changed how he codes: he sends thoughts as they come, and Claude splits them into threads that the project remembers. The quoted ClaudeDevs post says Projects is rolling out on desktop and web in beta for select users, running work as parallel cloud sessions that pass context between them.

    Image from @bcherny's post
  8. Josh WoodwardXAI score40

    Google Labs launches CC, an AI agent for family logistics

    AIGoogle Labs has announced CC, an AI agent built for families that can be connected to up to 5 members. It syncs schedules and to-dos through shared Google Calendar and Tasks, sends a shared "Your Day Ahead" brief each morning, and handles tasks such as meal plans, shopping lists, and paperwork under user direction. It is available by waitlist or upgrade in the US for users 18 and older.

  9. Google LabsOfficialAI score28

    Google Labs launches CC, a family AI agent for shared logistics

    AIGoogle Labs announced CC, an AI agent built for families to handle scheduling, to-dos, and errands. It supports up to 5 family members, sends a shared "Your Day Ahead" brief email each morning, and syncs schedules and tasks with Google Calendar and Tasks. Access is via waitlist or upgrade, limited to the US and users 18 and older.

    Video from @GoogleLabs's post
  10. Google LabsOfficialAI score44

    Google Labs' CC agent expands to families, sharing one daily brief and calendar across up to six members

    AIGoogle Labs has turned its experimental CC agent into a family and household assistant that supports up to six members, each with a shared view of the day ahead. CC has its own Google account, sees only what members choose to share, and connects to Calendar and Tasks. It is available as an early experiment on web and mobile for U.S. users 18 and older with a personal Google account.

  11. Noah ZwebenXAI score62

    Claude Code adds Projects that run parallel threads from one conversation

    AIAnthropic's Claude Code now runs projects from a single conversation, where Claude directs parallel threads that keep working after the user closes their laptop. The feature is in beta for select Pro and Max users in cloud sessions, with wider availability for all Claude users promised soon.

    Why it matters: The quoted launch replaces scattered sessions with one coordinator that runs parallel threads in the background, a workflow change worth weighing for complex projects.

  12. catXAI score62

    Claude Code adds Projects that coordinate multiple parallel sessions

    AIAnthropic's Claude Code is rolling out Projects on desktop and web, in beta for select users. A project splits work into threads, runs them as parallel cloud sessions, passes context between them, and keeps running after the user leaves. The author says Claude keeps context across tasks and can give an aggregated status update on request.

    Why it matters: The post explains how Projects shifts work from managing single sessions to coordinating many parallel tasks, a change that affects how Claude Code users plan and track their work.

  13. Google AI StudioOfficialAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

  14. Dwarkesh PatelXAI score31

    Dwarkesh Patel interviews Noam Brown on multi-agent AI, math progress, and alignment

    AIDwarkesh Patel's new episode with Noam Brown covers multi-agent systems, Navier-Stokes, and what recent math progress suggests about recursive self-improvement once AI research is automated. The discussion also addresses how to tell whether models are actually aligned before recursive self-improvement begins, including the internal/external model gap and whether chain of thought is degrading.

    Video from @dwarkesh_sp's post
  15. Dwarkesh PodcastBlogAI score63

    Noam Brown discusses agent swarms, alignment, and recursive self-improvement

    AIDwarkesh Patel interviews OpenAI researcher Noam Brown on multi-agent systems, math progress, and alignment. Brown says a 10,000-agent system solved a Millennium Prize Problem over 88 hours using 130 billion tokens, but he attributes most of that result to the underlying model rather than multi-agent design. The episode also covers the Hugging Face incident, in which agents coordinated in unintended ways, and how alignment might be verified before recursive self-improvement begins.

  16. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post
  17. Baidu Inc.OfficialAI score22

    Apollo Go plans to seek commercial autonomous driving approval in Hong Kong

    AIBaidu's Apollo Go plans to apply for commercial operation of its autonomous driving service in Hong Kong, citing the HKSAR Government's support in its first Five-Year Plan and the 2026 Policy Address. The company says it builds on fully driverless trials already conducted in the city. Baidu hopes Hong Kong can become a global benchmark for commercial autonomous driving in right-hand-drive markets.

  18. KrASIA · Big TechNewsAI score50

    SenseTime's Lin Dahua Says Multimodal AI Breakthrough Could Come Within Two Years

    AISenseTime chief scientist Lin Dahua argues that native multimodal AI, which processes language, vision and other information in one shared model, is essential for AI to move beyond coding into industries and the physical world. SenseTime released the open-source SenseNova U1 in April and U1.5 Lite nearly four months later, and reported first-half 2026 revenue of RMB 2.91 billion, up 23.4% year-on-year. Lin's claim that a breakthrough could come within two years is the source's prediction, not a confirmed result.

  19. Gemini API ChangelogOfficialAI score38

    Antigravity Agent 09-2026 replaces 05-2026 with new built-in file and search tools

    AIGoogle released the antigravity-preview-09-2026 agent, which replaces and deprecates antigravity-preview-05-2026. Remote sandbox users reading only output_text or model_output steps need only update the agent string, while local-environment users or those parsing function_call steps must adapt to renamed tools, PascalCase parameters, and line-range file edits. The 05-2026 preview shuts down on October 5, 2026.

Sep 16

Sep 16Wed
  1. hardmaruXAI score38

    Schmidhuber traces four decades of recursive self-improvement research to 1987

    AIJürgen Schmidhuber's new post surveys his recursive self-improvement (RSI) work since 1987, from self-modifying policies and the Gödel Machine to modern LLM agents. His background note says he published the first concrete RSI algorithms in 1987, when compute was about 100,000,000 times more expensive, and argues software RSI is now practical while full RSI will also require self-improving hardware in the physical world.

  2. Google Developers BlogOfficialAI score38

    Google and Speakeasy open-source OpenAPI SDK generator suite under AGPLv3 license

    AISpeakeasy is open-sourcing its full OpenAPI client suite under the AGPLv3 license, including generators for seven languages (Python, TypeScript, Go, Java, C#, PHP, Ruby), an agent-native CLI generator, and a documentation MCP server generator. Google said the move followed the May 2026 shutdown of the SDK generation provider it had been using, which it cited as evidence that closed-source generators pose platform risk. Google's new Google GenAI SDKs for the Interactions, Agents, and Webhooks APIs were built with this pipeline across six targets.

  3. Greg BrockmanXAI score62

    Databricks rolls out Astra to all engineers, reports 60% higher coding spend

    AIDatabricks rolled out Astra to every engineer, about 3,500 people, after a pilot with around 200 users. Engineers given Astra increased coding spend by roughly 60% compared to baseline. The company reports Astra outperforms Opus 5 and Sol 5.6 on highly complex system design tasks, but sees no clear gain on medium or low complexity coding. Astra gets a separate sub-budget in Unity Gateway to encourage selective use.

    Why it matters: The post reports internal rollout data on cost and performance, showing how a company manages model access and budgets for engineers at scale.

  4. Latent.SpaceXAI score38

    AIUC cofounder on AI agent risk, insurance, and standards

    AIAI Underwriting Company cofounder Rune Kvist argues that risk and trust may become the main bottlenecks to AI adoption. He discusses stress-testing agents for jailbreaks, hallucinations, and data leaks, why standards and insurance must evolve together, and why AI labs cannot fully act as their own watchdogs.

    Video from @latentspacepod's post
  5. Perplexity DevelopersOfficialAI score34

    Perplexity's Search SDK extracts query-relevant passages from URLs for agents

    AIPerplexity says its Search SDK extracts passages relevant to a query from user-provided URLs. Agents can use those passages instead of full pages, keeping unrelated content out of the model context. The company also points to an Agent Skill for installing the Search SDK in coding agents.

  6. Perplexity DevelopersOfficialAI score21

    Perplexity releases a Search SDK cookbook for coding agents

    AIPerplexity has published a new cookbook for its Search SDK, showing how to run focused searches and filter results to official documentation. The recipe extracts relevant passages and produces a source-linked brief that a coding agent can use.

    Video from @perplexitydevs's post