Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. Google AntigravityOfficialAI score36

    Antigravity builds and tests native Android apps from prompt to phone

    AIGoogle's Antigravity agent can take an Android app from prompt to a real device, using the Stitch MCP and Android CLI plugin. The agent pulls designs, builds native Jetpack Compose components, verifies them in the emulator, and runs the final build on a physical phone.

    Video from @antigravity's post
  2. GoogleOfficialAI score52

    Google Earth AI uses agents and satellite data to predict disease spread

    AIGoogle Earth AI combines environmental signals and other data sources with AlphaEarth Foundations, a Population Dynamics Foundation Model (PDFM), and a prototype Geospatial Reasoning agent. Researchers ask questions such as where a disease is likely to spread next, and the system automatically gathers relevant models and datasets to build a prediction model. By combining satellite views with population patterns, the tool aims to reveal hidden risk factors and identify issues earlier.

    Image from @Google's post
  3. TiboXAI score29

    OpenAI's Day 2 roundup adds auto-review, simplified API, and Decisions API

    AIOpenAI's Tibo announced that Approve for me (auto-review) is now included and does not consume usage, costing about 2-10% of a plan when used. The roundup also covers a simplified API for builders, meeting notes integration, and a Decisions API now live for builders, which the company will use in its own app.

  4. OpenAI DevelopersOfficialAI score13

    OpenAI's Decisions API powers routing, labeling, and screenshot-based actions

    AIDevelopers are using OpenAI's Decisions API to route requests to the right model, tool, or agent and to turn scaled inputs into labels, rankings, and scores. The post also lists uses including analyzing images and video frames, choosing buttons or form actions from screenshots, flagging risky tool calls, and categorizing large datasets.

    Video from @OpenAIDevs's post
  5. Gemini CLI · GitHub ReleasesOfficialAI score14

    Gemini CLI v0.63.0 released with retry indicator and auth loop fixes

    AIGemini CLI v0.63.0 adds a retry progress indicator during connection recovery and fixes an infinite authentication loop caused by file contention, headless keyring issues, and supervisor state drops. The release also bounds tool output size and cleans up temporary directories when background shell execution exits, alongside fixes for MCP enablement config handling and stdin restoration after capability detection.

  6. ChatGPTOfficialAI score44

    ChatGPT Meetings plugin takes notes and drafts follow-ups in beta

    AIOpenAI's ChatGPT Meetings plugin takes notes during meetings and saves a personalized summary and next steps in ChatGPT Space. Users can keep notes private or share them with their team, then ask ChatGPT to update a project plan or draft a follow-up. It is in beta for Pro and Business users in the ChatGPT desktop app on macOS, with Enterprise coming soon.

    Image from @ChatGPT's post
  7. elvisXAI score22

    Elvis Saravia urges learning to build good evals for domain edge

    AIElvis Saravia argues that building good evals on top of AI systems can put a practitioner at the frontier of their domain or task quickly. He advises readers to learn eval construction, calling it worth the investment. The post responds to Garry Tan's point that agents writing markdown skills on cron jobs can handle most knowledge work.

  8. AnthropicOfficialAI score49

    Anthropic expands Cyber Verification Program for verified security professionals

    AIAnthropic is expanding its Cyber Verification Program to give verified security professionals broader access to its most capable models. Through the program, they can use Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 with safeguards designed for defensive work. New tiers will also allow authorized offensive work such as penetration testing and red-teaming.

  9. Claude Code · GitHub ReleasesOfficialAI score40

    Claude Code v2.1.292 adds plugin marketplace flag and fixes security issues

    AIClaude Code v2.1.292 adds a --marketplace option to claude plugin install, which adds the marketplace if needed and then installs the plugin from it. The release also adds an effort parameter to the Agent tool and fixes several security issues, including permission prompts bypassed for network (UNC) file reads and a sandboxed read path that could return files outside approved access.

  10. Google ResearchOfficialAI score36

    Google Research demos Co-Director for coherent long-form AI video generation

    AIGoogle Research will demo Co-Director, a hierarchical multi-agent framework that optimizes video generation and consistency for long-form storytelling, at the #COLM2026 Google booth #107 today at 1:00 PM PT. The demo showcases interactive cinematic narratives, and the team's blog post details the approach.

    Image from @GoogleResearch's post
  11. elvisXAI score41

    Parsewave audit fixes 206 verifier bugs in AutomationBench

    AIParsewave audited all 600 public tasks in Zapier's AutomationBench and human review confirmed 206 real verifier bugs, all of which were fixed in AutomationBench Verified. Replaying 1,235 Kimi K3 runs on the old and fixed verifiers changed 27.9% of grades, with pass rate rising from 18.8% to 43.8% where verifiers were too strict and falling from 60.2% to 49.7% where they were too lenient.

  12. Matt ShumerXAI score34

    AgentID lets AI agents sign in to apps like Google login

    AIAgentMail has launched AgentID, a sign-in system that lets millions of AI agents with AgentMail accounts log into third-party apps within minutes. The post frames it as the agent-era equivalent of Sign in with Google, and offers three months of free AgentMail plans to developers who integrate it and share a screenshot.

  13. laurenXAI score42

    Developer automates releases and QA with Grok Bot agents in Slack

    AIA developer used Grok Bot to build two Slack team bots, sandcastle for release management and poteto for engineering, automating their release and QA process. Sandcastle DMs contributors PR links, kicks off builds, and runs a fuzz swarm of 10+ Grok 4.7 xhigh agents, while poteto triages and fixes issues via Cursor Projects.

    Image from @poteto's post
  14. Sierra BlogOfficialAI score62

    Sierra and Meta announce Personal Agent Protocol, an open standard for personal agents

    AISierra and Meta are developing Personal Agent Protocol, an open standard defining how personal agents interact with businesses, with industry partners including Genesys, Instinct, Rocket, Shopify, Stripe, and Walmart. The protocol uses OAuth sessions where consumers choose read-only or write access and companies choose whether agents reach them through websites, APIs via MCP and OpenAPI, or their own agents. The authors plan to publish the v0.1 specification later this month along with a reference implementation.

    Why it matters: The post specifies how personal agents would authenticate and reach businesses through websites, APIs, or company agents, which matters for anyone building agent integrations.

  15. LumaOfficialAI score13

    Luma argues no single model suits every creative stage

    AILuma Labs says there is no one model that excels at every task, since ideation and final-frame generation require different strengths. The post's point is that creative work should be matched to the best model for each stage rather than a single model for all.

    Video from @LumaLabsAI's post
  16. ClaudeOfficialAI score62

    Claude now works inside Google Docs, Sheets, and Slides

    AIand files from those apps can also be opened within Claude. In Google Workspace, Claude appears in a sidebar next to the open file, reads its contents, and edits it in place, with each edit available for user approval before it is applied.

    Why it matters: The source describes Claude working within Google Workspace files and approving edits, a direct change to how users in those apps can collaborate with the model.

    Video from @claudeai's post
  17. Allie K. MillerXAI score13

    Give your AI agent its own email to filter junk signups

    AIAllie K. Miller suggests giving an AI agent a separate email address, which Instinct does automatically, and using it for junk signups so the agent filters that mail away from your main inbox. She compares it to Google Voice for email and argues retail emails will get less attention unless they give people a reason to reach the human inbox.

  18. eric zakariassonXAI score32

    Grok turns one product photo into an 8-second vertical video ad

    AIA pipeline built with Grok takes a single product photo and produces an eight-second vertical ad, writing the brief, generating three scenes, selecting the best one, animating it, and adding a voiceover. Each run costs about $1.35, and the code is available in the xai-cookbook repository on GitHub.

    Video from @ericzakariasson's post
  19. eric zakariassonXAI score29

    Cursor adds five SpaceXAI TypeScript SDK demo apps to cookbook

    AICursor has added five apps to its cookbook, all built on the new SpaceXAI TypeScript SDK. The apps cover premise-to-short-film, picture-to-video ads, screenshot-to-React-component, X posts-to-sentiment dashboard, and link-to-podcast conversion. Demos and code are available in the post.

    Video from @ericzakariasson's post
  20. William ZhangXAI score15

    Zeroset raises $5.2M pre-seed to deploy enterprise world models

    AIZeroset announces a $5.2M pre-seed round, co-led by GradientVC and 2048VC with participation from Leblon Capital, to deploy world models in enterprises. The company says it will build models of each enterprise's decisions, dependencies, and unwritten processes from the traces people and agents create across systems, and that companies should own those models. Its first system, Nebula, captures state and traces as work happens.