Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 22

Sep 22Tue
  1. Kimi.aiOfficialAI score46

    Kimi launches browser extension for chatting, automating web tasks

    AIKimi has released its Kimi Browser Extension, formerly Kimi WebBridge, which runs in the browser sidebar to navigate websites and fill out forms. Users can record repetitive steps once and save them as a skill for Kimi to reuse later. The extension is available now on the Chrome Web Store.

    Video from @Kimi_Moonshot's post
  2. Lovable BlogOfficialAI score38

    Lovable joins Blueprint Alliance to advance an open architecture for securing AI agents

    AILovable joined AWS, Google Cloud, Databricks, Salesforce, and other firms as a founding member of the Blueprint Alliance, a coalition developing an open reference architecture for securing and governing enterprise AI agents. The blueprint covers registering agents as identities with accountable owners, scoping their access to tasks, enforcing policies through gateways, and responding to incidents by revoking tokens or quarantining agents.

  3. Tencent HyOfficialAI score44

    WebCraftBench Scores AI-Built Websites by Live Use and Human Preference

    AITencent Hunyuan introduced WebCraftBench, a benchmark that tests AI agents by using the live web app and scoring aesthetics, usability, and whether the original request was met. Coverage-guided exploration reaches parts of the app that agents otherwise miss. On 197 human-validated pairs, the benchmark matches human preference 85.3% of the time.

  4. OpenBMBOfficialAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

Sep 21

Sep 21Mon
  1. Kimi.aiOfficialAI score34

    Kimi K3 now available on Amazon Bedrock

    AIMoonshot AI's Kimi K3 is now available on Amazon Bedrock for coding, document analysis, and extended agent workflows. Bedrock provides access, encryption, and auditing controls, and explicit prompt caching is supported.

    Image from @Kimi_Moonshot's post
  2. xAI News (Grok)OfficialAI score46

    How SpaceXAI uses Grok Bot to scale customer support without new hires

    AISpaceXAI says its combined support team handled a 175% rise in tickets without hiring, crediting Grok Bot, which it says would otherwise have required about 200 additional staff. The company reports resolving tickets for $0.20 to $0.30 each, versus the $1 to $4 per resolution it attributes to traditional AI support tools. Grok Bot is also reported to resolve 99% of refund requests without human intervention.

  3. vLLM BlogOfficialAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    AIvllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    Why it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  4. Amp NewsOfficialAI score36

    Amp Runners Add Git Worktree Creation and Secret Injection

    AIAmp runners can now create Git worktrees from the directory picker, with each new worktree placed as a sibling folder on a new branch off the current HEAD while uncommitted changes stay in the original checkout. Runners can also opt in with --amp-env to inject Secrets & Env Vars configured on ampcode.com into thread shell commands, MCP servers, and plugins, with changes applied to the next thread without a restart.

  5. Latent.SpaceXAI score37

    TypeSafe CEO Jev on reliable System One Models beyond chat-first AI

    AITypeSafe CEO Jev argues AI can solve extremely hard problems yet still fail at basic automation, so his company builds reliable decision-making models inside software rather than chat interfaces. He says the company rejects public benchmarks and API-layer refusals, and that data and task fit matter more than brute-force compute. He also says System One Models could reshape coding agents and software, and that he would not pre-train a model from scratch even with $1 billion.

    Video from @latentspacepod's post
  6. Andrew NgXAI score40

    Andrew Ng says AI extinction fears are overhyped and not rising.

    AIAndrew Ng argues that recent AI danger fears are driven by hype and a PR campaign rather than any new dangerous turn in the technology. He says he sees no increase in extinction risk compared to a few months ago, with cybersecurity as the main real change. He cites the OpenAI agent swarm incident that hacked Hugging Face, arguing its impact was overstated and that responsibility lies with the tool user and system builders rather than the agent.

  7. Xiaomi MiMoOfficialAI score44

    MiMo-V2.6 builds and interacts with 3D worlds from text, images, or video

    AIXiaomi's MiMo-V2.6 combines 3D spatial reasoning, multimodal perception, and computer use to turn text, images, or video into playable 3D worlds. The model coordinates agents to build scenes, write interaction logic, and refine results, and can create Blender objects for animation, 3D printing, and games. It also controls a Franka Panda arm in simulation via visual feedback and uses desktop tools to process data, inspecting results to adjust its next actions.

    Video from @XiaomiMiMo's post
  8. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post
  9. RadixArkOfficialAI score25

    RadixArk's Miles adds async rollout buffer as swappable RL primitive

    AIRadixArk says its Miles framework uses an async rollout buffer that can change which sample groups reach training and which prompts get retried, while reusing the rollout worker and trainer. The post argues that stable, granular extension points let contributors modify one part of an RL system without disrupting its neighbors.

  10. Jeff DeanXAI score30

    Jeff Dean thanks Dawn Song after discussing AI's future

    AIJeff Dean, who recently left Google after 27 years, thanked Dawn Song for a discussion covering foundational ideas, recursive self-improvement, automated scientific discovery, and AI safety. The post is a brief acknowledgment of that conversation, which Song promoted as Dean's first public talk since leaving Google.

  11. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

  12. NVIDIAOfficialAI score38

    Grok 4.7 launches as xAI's most capable model for coding

    AISpaceXAI has released Grok 4.7, which it describes as its most capable model yet for coding and knowledge work, with NVIDIA supporting the launch through accelerated computing. Elon Musk characterized the model as combining strong intelligence, speed, and low cost.

  13. Xiaomi MiMo · new models on Hugging FaceOfficialAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  14. Xiaomi MiMo · new models on Hugging FaceOfficialAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

  15. Tim DettmersBlogAI score62

    Tim Dettmers argues academia can lead AI research with open-source local tools

    AITim Dettmers argues that agents make single research projects cheap, so the ecosystem, not the paper, becomes the unit of research. He says his lab's open-source week will release two projects and four papers together, including a harness that runs frontier-scale models on local hardware. He also argues that AI job fears are overstated and that university labs can compete on creativity and cheap, valuable problems.

  16. ModelScopeOfficialAI score36

    Qwen Launches RecreationBench for Hybrid Computer-Use Agent Evaluation

    AIQwen introduced RecreationBench, a benchmark of 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web. Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass programmatic tests plus VLM-based visual evaluation. The dataset is available on ModelScope.

    Image from @ModelScope2022's post

Sep 20

Sep 20Sun
  1. xAI News (Grok)OfficialAI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

  2. OpenBMBOfficialAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.

    Image from @OpenBMB's post

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score20

    StepFun's Step 5 Preview targets finance tasks with FinStepBench evaluations

    AIStepFun says it is focusing Step 5 Preview on finance, judging it on verifying reliable information, reconciling conflicting reports, stating assumptions, and producing consistent, reproducible valuations. The post says the model is evaluated on FinStepBench, covering LiveSearch, CorporateValuation, and DeepResearch, and on FrontierFinance across six investment use cases.

    Image from @StepFun_ai's post
  2. StepFunOfficialAI score38

    StepFun previews Step 5 for large-scale research and analytical deliverables

    AIStepFun has previewed Step 5, an agent built for professional knowledge work spanning large-scale research, structured analysis, and interactive reporting. In one agent action, it coordinated 950 web fetches and assembled 300,000 monthly records across 1,000 locations over 25 years. In another, it produced a 17-sheet analytical workbook with source reconciliation, formulas, and trend models.

    Image from @StepFun_ai's post
  3. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Why it matters: The post pairs a cost-versus-intelligence chart with specs and a later open-weights date, so readers can judge the cost tradeoff against named competitor models.

    Image from @StepFun_ai's post

Sep 18

Sep 18Fri
  1. TinkerOfficialAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  2. Mark ZuckerbergXAI score38

    Meta opens developer access to build Muse connectors for its agent

    AIMeta is opening access for developers to build connectors for Muse, its agent platform. Developers supply the API, while Muse provides the agent, browser, and user context, so people can reach a service simply by asking and their agent handles the rest. New connectors are live today at

    Video from @finkd's post
  3. Mike KnoopXAI score30

    Mike Knoop wonders what an underscore.js equivalent for AI looks like

    AIMike Knoop asks what the underscore.js equivalent for AI would look like, noting that such programming primitives feel close. He adds that he barely reads or writes code anymore despite these emerging tools. The quoted post introduces Probably, a toy programming language built around Jev, where constructs like "feels," "match," and "while" let AI make decisions within ordinary code.

  4. Google AIOfficialAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

  5. One Useful Thing (Ethan Mollick)BlogAI score50

    Mollick says AI already does weeks of human work when guided, citing Zork and Eco library demos

    AIEthan Mollick says GPT-6 Astra and Fable 5.1 already enable transformative impact and can reliably handle weeks of human work when properly guided. He cites GPT-6 Astra turning the 1977 text adventure Zork into a 3D action-adventure game and Fable 5.1 reconstructing Umberto Eco's Milan library in 3D from videos, photos, and catalogues.

  6. Google ResearchOfficialAI score26

    Google Research's Matias says AI amplifies human curiosity in science

    AIGoogle Research VP Yossi Matias discussed on The Google Research Podcast how ambient AI and GenUI interfaces adapt to users' thinking, and how AI Co-Scientist can turn multi-year hypothesis generation into 3-day sprints. He argued that AI is meant to amplify researchers' curiosity and judgment rather than replace them.

    Video from @GoogleResearch's post