Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Sep 22

Sep 22Tue
  1. Hacker News · Launch HN, YC launches (10+ points)BlogAI score18

    Coverage Cat launches umbrella insurance quotes through personal AI agents

    AICoverage Cat, a licensed insurance brokerage from YC S22, lets users and their AI agents compare umbrella, home, auto, and renters coverage through a portal or Agent API. The company says it sells no leads and earns no commission on the policies its agents recommend, and it is live for shoppers in California, Florida, New York, Texas, and Washington.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score26

    Cognition Expands to Latin America, Launching São Paulo Hub for Devin Software Engineering

    AICognition announced its expansion into Latin America at MASP in São Paulo, starting with a local team to help companies build more of their software in the region. Itaú reports more than 75% of its technology teams use Devin, with legacy .NET services migrated to Java 6x faster and about 70% of security vulnerabilities resolved automatically. Nubank says Devin cut a multi-million-line monolith migration from years to weeks, at over 20x lower cost.

  3. Mike KriegerXAI score67

    Anthropic launches Claude Opus 5.5, leading in coding and knowledge work

    AIAnthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 family. According to the quoted launch post, it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. The author says it leads in coding and knowledge work and praises its writing quality.

    Why it matters: The quoted launch post gives a concrete cost comparison, useful for weighing Opus 5.5 against earlier Opus and Fable 5.1 models for routine work.

  4. catXAI score62

    Claude Opus 5.5 becomes the default model in Claude Code and Claude app

    AIClaude Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. Anthropic is defaulting to effort medium across products, which it says is comparable to Fable 5.1 on intelligence but faster. Rate limits will go 25% further on Opus 5.5 compared to Opus 5.

    Why it matters: The post names concrete default changes across Claude Code and the Claude app, plus a specific effort setting and rate-limit difference, useful for judging day-to-day cost and speed.

  5. StepFunOfficialAI score27

    StepFun's Step Code tops Terminal-Bench 2.1 and Multi-Frame with fewer tokens

    AIStepFun's Step Code passed 72 of 89 tasks (80.9%) on Terminal-Bench 2.1, tying for the highest pass rate among evaluated harnesses while using fewer tokens than the other tied leaders. On Multi-Frame, it passed 110 of 150 tasks (73.3%) and averaged 5.09M tokens per task, the highest pass rate and lowest token use among six harnesses evaluated.

    Image from @StepFun_ai's post
  6. StepFunOfficialAI score52

    StepFun releases Step Code v0.1.0 as an open-source coding CLI

    AIStepFun has released Step Code v0.1.0, an open-source command-line tool under the MIT License that covers reading and editing code, running tests, and shipping from one CLI. The post reports 80.9% on Terminal-Bench 2.1 and 73.3% on Multi-Frame, a 150-task long-horizon benchmark from StepFun. It also includes one-command static site publishing with StepPage and links the GitHub repository.

    Image from @StepFun_ai's post
  7. WorkBuddyOfficialAI score18

    HKUST students build two AI workbenches with WorkBuddy, win Game Track

    AIHKUST's Anchor team used WorkBuddy to build two production-ready workbenches and won the Game Track championship. Kaiwu Producer creates a complete FPS game in 8 hours through full-pipeline 3D generation with an AI-driven narrative memory engine and zero human intervention. Anchor is a de-labeling narrative engine that automatically detects stereotypical dependencies.

    Video from @WorkBuddy_AI's post
  8. WorkBuddyOfficialAI score13

    Hong Kong Teens Build Interactive AI Cinematic Game With WorkBuddy and Miora

    AIThree 15-year-old Hong Kong students won the Animation Track at the Tencent Cloud Hackathon Global Finals with an interactive cinematic game. The game has 37 scenes, 10 choices, and 7 endings, built with Miora for style and asset consistency and WorkBuddy for visual narrative mapping.

    Video from @WorkBuddy_AI's post
  9. Sebastian RaschkaXAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  10. Kimi.aiOfficialAI score46

    Kimi launches browser extension for chatting, automating web tasks

    AIKimi has released its Kimi Browser Extension, formerly Kimi WebBridge, which runs in the browser sidebar to navigate websites and fill out forms. Users can record repetitive steps once and save them as a skill for Kimi to reuse later. The extension is available now on the Chrome Web Store.

    Video from @Kimi_Moonshot's post
  11. Lovable BlogOfficialAI score38

    Lovable joins Blueprint Alliance to advance an open architecture for securing AI agents

    AILovable joined AWS, Google Cloud, Databricks, Salesforce, and other firms as a founding member of the Blueprint Alliance, a coalition developing an open reference architecture for securing and governing enterprise AI agents. The blueprint covers registering agents as identities with accountable owners, scoping their access to tasks, enforcing policies through gateways, and responding to incidents by revoking tokens or quarantining agents.

  12. Tencent HyOfficialAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

  13. OpenBMBOfficialAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

Sep 21

Sep 21Mon
  1. Kimi.aiOfficialAI score34

    Kimi K3 now available on Amazon Bedrock

    AIMoonshot AI's Kimi K3 is now available on Amazon Bedrock for coding, document analysis, and extended agent workflows. Bedrock provides access, encryption, and auditing controls, and explicit prompt caching is supported.

    Image from @Kimi_Moonshot's post
  2. xAI News (Grok)OfficialAI score46

    How SpaceXAI uses Grok Bot to scale customer support without new hires

    AISpaceXAI says its combined support team handled a 175% rise in tickets without hiring, crediting Grok Bot, which it says would otherwise have required about 200 additional staff. The company reports resolving tickets for $0.20 to $0.30 each, versus the $1 to $4 per resolution it attributes to traditional AI support tools. Grok Bot is also reported to resolve 99% of refund requests without human intervention.

  3. vLLM BlogOfficialAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    AIvllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    Why it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  4. Amp NewsOfficialAI score36

    Amp Runners Add Git Worktree Creation and Secret Injection

    AIAmp runners can now create Git worktrees from the directory picker, with each new worktree placed as a sibling folder on a new branch off the current HEAD while uncommitted changes stay in the original checkout. Runners can also opt in with --amp-env to inject Secrets & Env Vars configured on ampcode.com into thread shell commands, MCP servers, and plugins, with changes applied to the next thread without a restart.

  5. Latent.SpaceXAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

    Video from @latentspacepod's post
  6. Andrew NgXAI score40

    Andrew Ng says AI extinction fears are overhyped and not rising.

    AIAndrew Ng argues that recent AI danger fears are driven by hype and a PR campaign rather than any new dangerous turn in the technology. He says he sees no increase in extinction risk compared to a few months ago, with cybersecurity as the main real change. He cites the OpenAI agent swarm incident that hacked Hugging Face, arguing its impact was overstated and that responsibility lies with the tool user and system builders rather than the agent.

  7. Xiaomi MiMoOfficialAI score44

    MiMo-V2.6 builds and interacts with 3D worlds from text, images, or video

    AIXiaomi's MiMo-V2.6 combines 3D spatial reasoning, multimodal perception, and computer use to turn text, images, or video into playable 3D worlds. The model coordinates agents to build scenes, write interaction logic, and refine results, and can create Blender objects for animation, 3D printing, and games. It also controls a Franka Panda arm in simulation via visual feedback and uses desktop tools to process data, inspecting results to adjust its next actions.

    Video from @XiaomiMiMo's post
  8. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post
  9. RadixArkOfficialAI score25

    RadixArk's Miles adds async rollout buffer as swappable RL primitive

    AIRadixArk says its Miles framework uses an async rollout buffer that can change which sample groups reach training and which prompts get retried, while reusing the rollout worker and trainer. The post argues that stable, granular extension points let contributors modify one part of an RL system without disrupting its neighbors.

  10. Jeff DeanXAI score30

    Jeff Dean thanks Dawn Song after discussing AI's future

    AIJeff Dean, who recently left Google after 27 years, thanked Dawn Song for a discussion covering foundational ideas, recursive self-improvement, automated scientific discovery, and AI safety. The post is a brief acknowledgment of that conversation, which Song promoted as Dean's first public talk since leaving Google.

  11. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

  12. NVIDIAOfficialAI score38

    Grok 4.7 launches as xAI's most capable model for coding

    AISpaceXAI has released Grok 4.7, which it describes as its most capable model yet for coding and knowledge work, with NVIDIA supporting the launch through accelerated computing. Elon Musk characterized the model as combining strong intelligence, speed, and low cost.

  13. Xiaomi MiMo · new models on Hugging FaceOfficialAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  14. Xiaomi MiMo · new models on Hugging FaceOfficialAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

  15. Tim DettmersBlogAI score62

    Tim Dettmers argues academia can lead AI research with open-source local tools

    AITim Dettmers argues that agents make single research projects cheap, so the ecosystem, not the paper, becomes the unit of research. He says his lab's open-source week will release two projects and four papers together, including a harness that runs frontier-scale models on local hardware. He also argues that AI job fears are overstated and that university labs can compete on creativity and cheap, valuable problems.

  16. ModelScopeOfficialAI score36

    Qwen Launches RecreationBench for Hybrid Computer-Use Agent Evaluation

    AIQwen introduced RecreationBench, a benchmark of 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web. Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass programmatic tests plus VLM-based visual evaluation. The dataset is available on ModelScope.

    Image from @ModelScope2022's post

Sep 20

Sep 20Sun
  1. xAI News (Grok)OfficialAI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

  2. Philipp SchmidXAI score14

    Schmid praises Jev's fast, cheap multimodal AI demos beyond coding

    AIPhilipp Schmid, who works on Google's Gemini, praised a series of demos, replications, and projects by Jev, which he says show what becomes possible with ultra-fast, multimodal models cheap enough for free use. He adds that the focus on coding agents makes it easy to overlook the broader range of tasks AI can handle.

  3. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post