Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. Google GeminiOfficialAI score12

    Gemini can draft and send PandaDoc contracts from a simple prompt

    AIGoogle Gemini now integrates with PandaDoc, letting users draft, customize, and send client-ready contracts. Users simply ask Gemini to create a document, such as a Services Agreement for Acme Corp with a $50,000 contract value, using PandaDoc.

    Image from @GeminiApp's post
  2. Google GeminiOfficialAI score10

    Gemini can create and update Monday.com tasks via natural language commands

    AIGoogle Gemini can manage projects on Monday.com when users give plain-language instructions, such as creating a task named "Finalize Q4 Marketing Plan" on a board and setting its status to "To Do." The post provides this single example and does not describe further capabilities or availability details.

    Image from @GeminiApp's post
  3. Azure BlogOfficialAI score40

    Azure resilience now requires continuous validation, not just architecture diagrams

    AIMicrosoft's Azure Blog argues that resilience drifts as workloads change, so architecture diagrams cannot prove a system is resilient. It says roughly 70 percent of cloud outages are related to change, and that teams need health modeling and resiliency goals measured against live signals. The article is the first in a series on validating resilience at scale.

  4. QwenOfficialAI score10

    Qwen's Mobile Creative Agent Performance Is Highlighted

    AIQwen shared performance results for its Mobile Creative Agent, though the post provides no specific figures or details. The brief announcement offers limited information beyond the performance label.

    Image from @Alibaba_Qwen's post
  5. QwenOfficialAI score17

    Qwen shares performance results for its mobile-use agent

    AIQwen posted performance results for its mobile-use agent, though the post itself provides no benchmark scores or details. The announcement is a brief teaser with no further specifics on the agent's capabilities.

    Image from @Alibaba_Qwen's post
  6. QwenOfficialAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

    Why it matters: The post names three mobile agents and their benchmark results, while also releasing the benchmark suite, so readers can check the claims against the reported figures.

    Image from @Alibaba_Qwen's post
  7. ModelScopeOfficialAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

    Why it matters: The post links benchmark results, parameter scale, and a multi-agent RL training run, giving readers concrete figures to compare against other open models.

    Image from @ModelScope2022's post
  8. Anthropic NewsroomOfficialAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    AIAnthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    Why it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

  9. Prime Intellect BlogOfficialAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    AIPrime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    Why it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 22

Sep 22Tue
  1. Redwood Research BlogBlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  2. PlatformerBlogAI score15

    Muse Is Having a Moment: Consumer AI Agents Explored

    AIPlatformer asks whether consumer AI agents are the future or a mirage, but the source text provided is only that single question. No further details about Muse's features, performance, or availability are available, so no additional claims can be made.

  3. Google Developers BlogOfficialAI score62

    Antigravity SDK adds local Gemma 4 26B agent support via LiteRT

    AIGoogle announced that the Antigravity SDK supports local agent workflows, with initial support for Gemma 4 26B A4B through Google AI Edge's LiteRT. The post includes Python setup steps and says a recommended machine has more than 24GB VRAM or unified memory. It also describes a hybrid pattern in which a cloud Gemini 3.8 Flash planner hands work to local Gemma 4 26B models, with 97.2% of tokens in one recorded run staying local.

    Why it matters: The source shows how to run an agent with a local Gemma 4 26B model using LiteRT, plus a hybrid cloud-planner pattern that keeps most tokens on-device.

  4. Tibor BlahoXAI score88

    OpenAI launches GPT-6 Sol and Luna while Anthropic releases Claude Opus 5.5

    AIOpenAI released GPT-6 Sol and Luna, with API prices cut in half, while Anthropic released Claude Opus 5.5 at roughly Fable 5.1 level for 40% less than Opus 5. GPT-6 Sol and Luna cost $2/$10 and $0.10/$0.50 per million tokens, versus GPT-5.6 promotional pricing, and Opus 5.5 costs $4/$20 per million tokens. Sonnet 5.5 and Haiku 5.5 are announced for the coming weeks.

    Why it matters: The post links OpenAI's GPT-6 Sol and Luna pricing with Anthropic's Claude Opus 5.5 launch, which helps readers compare the two vendors' current frontier offerings.

    Image from @btibor91's post
  5. eric zakariassonXAI score13

    Cursor's scrappy support system evolved into a scalable, self-improving operation

    AICursor's support team says a scrappy version built about 1.5 years ago to handle heavy volume taught them to crawl, walk, then run, and to identify flywheels for continuous improvement. The background post notes the company rebuilt customer support around Grok Bot to handle operations without adding headcount, with the bot responding to customers, resolving tickets, and managing the queue autonomously.

  6. Sierra BlogOfficialAI score34

    Sierra Lets Companies See, Edit, and Export Their AI Agents' Logic and Data

    AISierra says its platform makes enterprise AI agents visible and editable, with journeys, policies, and actions viewable in Agent Studio and testable through Simulations and Experiments before rollout. Customers can export agent logic in a portable structured format, access conversation logs and performance data through export APIs, and manage the agent's code in a Git repository. Sierra agents also connect to existing systems through MCP, REST, GraphQL, or custom integrations.

  7. ChatGPTOfficialAI score62

    OpenAI rolls out GPT-6 Sol and GPT-6 Luna in ChatGPT Work and Codex

    AIOpenAI announced GPT-6 Sol and GPT-6 Luna, rolling out today in ChatGPT Work and Codex. The rollout covers Plus, Pro, Business, Enterprise, and Edu users.

    Why it matters: The post names two new GPT-6 variants and their rollout to specific ChatGPT and Codex plan tiers, which shows how access is being staged.

    Video from @ChatGPT's post
  8. Alex AlbertXAI score60

    Claude Cowork and chat are merging into a single Claude

    AIAnthropic's Claude Cowork and chat are being merged into one Claude, which can take a question or report and keep working after the user closes their laptop. When something is unclear, Claude asks, and the user keeps the final say. The update rolls out to Pro and Max over the next few weeks, and the author notes it is coming to everyone soon.

    Why it matters: The quoted announcement merges Claude Cowork and chat into one interface, which changes how users hand off longer tasks and asks for their input.

  9. Hacker News · Launch HN, YC launches (10+ points)BlogAI score18

    Coverage Cat launches umbrella insurance quotes through personal AI agents

    AICoverage Cat, a licensed insurance brokerage from YC S22, lets users and their AI agents compare umbrella, home, auto, and renters coverage through a portal or Agent API. The company says it sells no leads and earns no commission on the policies its agents recommend, and it is live for shoppers in California, Florida, New York, Texas, and Washington.

  10. Cognition Blog (Devin, Windsurf)OfficialAI score26

    Cognition Expands to Latin America, Launching São Paulo Hub for Devin Software Engineering

    AICognition announced its expansion into Latin America at MASP in São Paulo, starting with a local team to help companies build more of their software in the region. Itaú reports more than 75% of its technology teams use Devin, with legacy .NET services migrated to Java 6x faster and about 70% of security vulnerabilities resolved automatically. Nubank says Devin cut a multi-million-line monolith migration from years to weeks, at over 20x lower cost.

  11. Mike KriegerXAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  12. catXAI score62

    Claude Opus 5.5 becomes the default model in Claude Code and Claude app

    AIClaude Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. Anthropic is defaulting to effort medium across products, which it says is comparable to Fable 5.1 on intelligence but faster. Rate limits will go 25% further on Opus 5.5 compared to Opus 5.

    Why it matters: The post names concrete default changes across Claude Code and the Claude app, plus a specific effort setting and rate-limit difference, useful for judging day-to-day cost and speed.

  13. StepFunOfficialAI score27

    StepFun's Step Code tops Terminal-Bench 2.1 and Multi-Frame with fewer tokens

    AIStepFun's Step Code passed 72 of 89 tasks (80.9%) on Terminal-Bench 2.1, tying for the highest pass rate among evaluated harnesses while using fewer tokens than the other tied leaders. On Multi-Frame, it passed 110 of 150 tasks (73.3%) and averaged 5.09M tokens per task, the highest pass rate and lowest token use among six harnesses evaluated.

    Image from @StepFun_ai's post
  14. StepFunOfficialAI score52

    StepFun releases Step Code v0.1.0 as an open-source coding CLI

    AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

    Image from @StepFun_ai's post
  15. WorkBuddyOfficialAI score18

    HKUST students build two AI workbenches with WorkBuddy, win Game Track

    AIHKUST's Anchor team used WorkBuddy to build two production-ready workbenches and won the Game Track championship. Kaiwu Producer creates a complete FPS game in 8 hours through full-pipeline 3D generation with an AI-driven narrative memory engine and zero human intervention. Anchor is a de-labeling narrative engine that automatically detects stereotypical dependencies.

    Video from @WorkBuddy_AI's post
  16. WorkBuddyOfficialAI score13

    Hong Kong Teens Build Interactive AI Cinematic Game With WorkBuddy and Miora

    AIThree 15-year-old Hong Kong students won the Animation Track at the Tencent Cloud Hackathon Global Finals with an interactive cinematic game. The game has 37 scenes, 10 choices, and 7 endings, built with Miora for style and asset consistency and WorkBuddy for visual narrative mapping.

    Video from @WorkBuddy_AI's post
  17. Sebastian RaschkaXAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  18. Kimi.aiOfficialAI score46

    Kimi launches browser extension for chatting, automating web tasks

    AIKimi has released its Kimi Browser Extension, formerly Kimi WebBridge, which runs in the browser sidebar to navigate websites and fill out forms. Users can record repetitive steps once and save them as a skill for Kimi to reuse later. The extension is available now on the Chrome Web Store.

    Video from @Kimi_Moonshot's post