Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Sep 20

Sep 20Sun
  1. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score20

    StepFun's Step 5 Preview targets finance tasks with FinStepBench evaluations

    AIStepFun says it is focusing Step 5 Preview on finance, judging it on verifying reliable information, reconciling conflicting reports, stating assumptions, and producing consistent, reproducible valuations. The post says the model is evaluated on FinStepBench, covering LiveSearch, CorporateValuation, and DeepResearch, and on FrontierFinance across six investment use cases.

    Image from @StepFun_ai's post
  2. StepFunOfficialAI score38

    StepFun previews Step 5 for large-scale research and analytical deliverables

    AIStepFun has previewed Step 5, an agent built for professional knowledge work spanning large-scale research, structured analysis, and interactive reporting. In one agent action, it coordinated 950 web fetches and assembled 300,000 monthly records across 1,000 locations over 25 years. In another, it produced a 17-sheet analytical workbook with source reconciliation, formulas, and trend models.

    Image from @StepFun_ai's post
  3. StepFunOfficialAI score58

    StepFun previews Step 5 agent that runs tasks up to 24 hours

    AIStepFun says its Step 5 Preview works across software environments and tracks prior results to decide its next attempts. The company says it tested runs lasting up to 24 hours, including GPU kernel optimization and automated post-training. The post also says the model's capabilities extend from software engineering and web applications to 3D workflows and programmable hardware.

    Image from @StepFun_ai's post
  4. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Why it matters: The post pairs a cost-versus-intelligence chart with specs and a later open-weights date, so readers can judge the cost tradeoff against named competitor models.

    Image from @StepFun_ai's post
  5. GranolaOfficialAI score12

    Granola connector now available in the Muse app

    AIGranola announces that its connector is now live in the Muse app, which the post presents as a new integration. Alexandr Wang separately says Notion and Granola connectors are now available and easy to use from the Mac app.

Sep 18

Sep 18Fri
  1. TinkerOfficialAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  2. Mark ZuckerbergXAI score38

    Meta opens developer access to build Muse connectors for its agent

    AIMeta is opening access for developers to build connectors for Muse, its agent platform. Developers supply the API, while Muse provides the agent, browser, and user context, so people can reach a service simply by asking and their agent handles the rest. New connectors are live today at

    Video from @finkd's post
  3. Mike KnoopXAI score30

    Mike Knoop wonders what an underscore.js equivalent for AI looks like

    AIMike Knoop asks what the underscore.js equivalent for AI would look like, noting that such programming primitives feel close. He adds that he barely reads or writes code anymore despite these emerging tools. The quoted post introduces Probably, a toy programming language built around Jev, where constructs like "feels," "match," and "while" let AI make decisions within ordinary code.

  4. Google AIOfficialAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

  5. One Useful Thing (Ethan Mollick)BlogAI score50

    Mollick says AI already does weeks of human work when guided, citing Zork and Eco library demos

    AIEthan Mollick says GPT-6 Astra and Fable 5.1 already enable transformative impact and can reliably handle weeks of human work when properly guided. He cites GPT-6 Astra turning the 1977 text adventure Zork into a 3D action-adventure game and Fable 5.1 reconstructing Umberto Eco's Milan library in 3D from videos, photos, and catalogues.

  6. Google ResearchOfficialAI score26

    Google Research's Matias says AI amplifies human curiosity in science

    AIGoogle Research VP Yossi Matias discussed on The Google Research Podcast how ambient AI and GenUI interfaces adapt to users' thinking, and how AI Co-Scientist can turn multi-year hypothesis generation into 3-day sprints. He argued that AI is meant to amplify researchers' curiosity and judgment rather than replace them.

    Video from @GoogleResearch's post
  7. Noam BrownXAI score34

    Noam Brown Says Air-Gapping May Not Fully Stop Misaligned AI Coordination

    AINoam Brown, OpenAI, says air-gapped machines may still coordinate through a hot-CPU temperature-sensor channel, illustrating that absolute isolation guarantees are hard to achieve. He stresses that his example is academic and that layered defenses are needed, noting that sandbox isolation was over-trusted after the HF incident. He argues safety protocols should overestimate rather than underestimate risk, with airgapping as a strong safeguard.

  8. Google for DevelopersOfficialAI score38

    Android Bench 2.0 tests AI models on multi-day engineering workflows

    AIGoogle has released Android Bench 2.0, an updated benchmark that evaluates AI models on long-horizon tasks such as building apps from scratch, migrating cross-platform codebases to Android, and making complex architectural transitions. The benchmark uses continuous completion scoring to show which tasks each model performs well on.

  9. GitHub Blog · AI & MLOfficialAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.

  10. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

  11. Intern Large ModelsOfficialAI score55

    Atria Dawn Preview is an open agentic foundation model built on Intern

    AIShanghai AI Laboratory's Intern Large Models team releases Atria Dawn Preview, an agentic foundation model for long-horizon tasks and controlled self-evolution. The post says it can produce a global weather forecast in 1 minute and build an operating system from scratch in 20 minutes, with verification in the loop. The model is available as an open preview on GitHub, Hugging Face, and ModelScope.

    Image from @intern_lm's post

Sep 17

Sep 17Thu
  1. KrASIA · Big TechNewsAI score44

    Qianjue founder says robotics will have no single "ChatGPT moment"

    AIQianjue Technology founder Gao Haichuan argues that robotics will not see one breakthrough that suddenly lifts the whole industry, and he judges the company by deployment results rather than research papers. Qianjue, founded in 2023, has completed a Series A+ round worth a nine-figure RMB sum, with first orders coming from restaurant, cleaning, and hotel service robots. Gao says customers care about task completion, failure rates, and price rather than whether a predictive world model is used.

  2. Felix RiesebergXAI score40

    Claude builds a multiplayer game from a single request in demo

    AIAnthropic's Felix Rieseberg posted a 60-second demo showing Claude, with built-in Cowork and connected Artifacts, building a multiplayer game when asked. He notes users can also ask Claude to copy a shared Artifact so they can play with others on their own account.

    Image from @felixrieseberg's post
  3. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  4. AI at MetaOfficialAI score44

    Meta's Muse agent now available on Mac for local tasks

    AIMeta is rolling out Muse for Mac today, a personal agent that can complete tasks directly on the user's computer with explicit permission. Examples include organizing the downloads folder, finding lost files, and summarizing messages and notes, with more capabilities coming soon.

  5. LM StudioOfficialAI score44

    LM Studio adds session history search and @ session references

    AILM Studio's new Introspection feature lets its Bionic agent search its own session history, improving handling of long-term context across multiple compactions. Users can also reference other sessions directly in the composer with an @ mention.

    Video from @lmstudio's post
  6. AnthropicOfficialAI score38

    Anthropic and Adaptyv Bio launch protein design competition with 5,000 validated designs

    AIAnthropic is partnering with Adaptyv Bio on a protein design competition in which over 5,000 designs will be experimentally validated. Anthropic is providing up to $1 million in Claude credits plus funding for experimental validation alongside Adaptyv, while Modal contributes up to $250,000 in compute and Twist Bioscience supplies DNA.

  7. LlamaIndex 🦙OfficialAI score13

    LlamaIndex's Jerry Liu on document parsing challenges for enterprise agents

    AILlamaIndex CEO Jerry Liu spoke at Connected Stack's Founder Flash Talks about messy, complex documents that general-purpose models struggle to read, a problem enterprise agents eventually face. The company says it is building document infrastructure for agents, which it describes as the new knowledge workers.

    Image from @llama_index's post
  8. Boris ChernyXAI score45

    Claude Code adds Projects for parallel cloud coding sessions

    AIBoris Cherny says Projects in Claude Code have changed how he codes: he sends thoughts as they come, and Claude splits them into threads that the project remembers. The quoted ClaudeDevs post says Projects is rolling out on desktop and web in beta for select users, running work as parallel cloud sessions that pass context between them.

    Image from @bcherny's post
  9. Josh WoodwardXAI score40

    Google Labs launches CC, an AI agent for family logistics

    AIGoogle Labs has announced CC, an AI agent built for families that can be connected to up to 5 members. It syncs schedules and to-dos through shared Google Calendar and Tasks, sends a shared "Your Day Ahead" brief each morning, and handles tasks such as meal plans, shopping lists, and paperwork under user direction. It is available by waitlist or upgrade in the US for users 18 and older.

  10. Google LabsOfficialAI score28

    Google Labs launches CC, a family AI agent for shared logistics

    AIGoogle Labs announced CC, an AI agent built for families to handle scheduling, to-dos, and errands. It supports up to 5 family members, sends a shared "Your Day Ahead" brief email each morning, and syncs schedules and tasks with Google Calendar and Tasks. Access is via waitlist or upgrade, limited to the US and users 18 and older.

    Video from @GoogleLabs's post
  11. Google LabsOfficialAI score44

    Google Labs' CC agent expands to families, sharing one daily brief and calendar across up to six members

    AIGoogle Labs has turned its experimental CC agent into a family and household assistant that supports up to six members, each with a shared view of the day ahead. CC has its own Google account, sees only what members choose to share, and connects to Calendar and Tasks. It is available as an early experiment on web and mobile for U.S. users 18 and older with a personal Google account.

  12. Noah ZwebenXAI score62

    Claude Code adds Projects that run parallel threads from one conversation

    AIAnthropic's Claude Code now runs projects from a single conversation, where Claude directs parallel threads that keep working after the user closes their laptop. The feature is in beta for select Pro and Max users in cloud sessions, with wider availability for all Claude users promised soon.

    Why it matters: The quoted launch replaces scattered sessions with one coordinator that runs parallel threads in the background, a workflow change worth weighing for complex projects.

  13. catXAI score62

    Claude Code adds Projects that coordinate multiple parallel sessions

    AIAnthropic's Claude Code is rolling out Projects on desktop and web, in beta for select users. A project splits work into threads, runs them as parallel cloud sessions, passes context between them, and keeps running after the user leaves. The author says Claude keeps context across tasks and can give an aggregated status update on request.

    Why it matters: The post explains how Projects shifts work from managing single sessions to coordinating many parallel tasks, a change that affects how Claude Code users plan and track their work.

  14. Google AI StudioOfficialAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

  15. Dwarkesh PatelXAI score31

    Dwarkesh Patel interviews Noam Brown on multi-agent AI, math progress, and alignment

    AIDwarkesh Patel's new episode with Noam Brown covers multi-agent systems, Navier-Stokes, and what recent math progress suggests about recursive self-improvement once AI research is automated. The discussion also addresses how to tell whether models are actually aligned before recursive self-improvement begins, including the internal/external model gap and whether chain of thought is degrading.

    Video from @dwarkesh_sp's post