Skip to contentSkip to stories

Updated

#Agent

Sep 8

Sep 8Tue
  1. AI at MetaAI score67

    Meta introduces Muse, a personal agent powered by Muse Spark 1.3

    AIMeta announced Muse, a personal AI agent designed to get things done for users across many parts of life. The product is powered by Muse Spark 1.3, and the post links to an app download and a page describing how Muse was built.

    Why it matters: The announcement names Muse Spark 1.3 as the underlying model, giving readers a concrete product and model pairing to track.

  2. Mckay WrigleyAI score80

    OpenAI shares agent-produced proof of Navier-Stokes Millennium Prize problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem. The problem asks whether smooth three-dimensional fluid motion described by the Navier-Stokes equations can break down, and it has remained unresolved for roughly 90 years. The author, Mckay Wrigley, reposted the claim with his own remark about roughly 10k agents working in a datacenter.

    Why it matters: The quoted OpenAI post makes a major mathematical claim about the Navier-Stokes problem, so readers should weigh it against the proof's verification status.

  3. Werner VogelsAI score50

    Werner Vogels highlights Kiro Crew's memory system drawing on brain evolution

    AIWerner Vogels says that after spending time with Kiro Crew since its launch, its memory system stands out for deciding what to keep, compress, and let go. He notes that Amazon engineers, starting from engineering constraints, arrived at an approach resembling the brain's evolved architecture. Per the referenced post, Kiro Crew is a persistent workspace that retains project context across sessions and runs scheduled jobs.

  4. Xiaomi MiMoAI score52

    Xiaomi MiMo Desktop enters invite-only beta as a desktop agent

    AIXiaomi MiMo has launched MiMo Desktop in invite-only beta, a desktop agent that turns Office files, images, video, audio, and zips into finished, editable output. Invitees also get limited access to next-gen MiMo models, and the post lists features including live previews, region-based editing with versioned rollback, automatic model routing, and browser and computer use with record and replay.

  5. NVIDIA · new models on Hugging FaceAI score46

    NVIDIA Releases NV-Reason-CT, a 3D Vision-Language Model for Chest and Abdominal CT

    AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.

Sep 7

Sep 7Mon
  1. Baidu Inc.AI score22

    Baidu launches AI, Evolving podcast on AI in scientific discovery

    AIBaidu has launched AI, Evolving, a new podcast series, with its first episode examining AI's growing role in scientific discovery through Famou's work on pine wilt disease. The post frames this as part of a broader trend in which AI takes on more of the research process itself. It asks whether research agents could become part of the infrastructure of discovery.

  2. Ian JohnsonAI score38

    Ian Johnson: knowing what to ask AI for matters most for value

    AIOrbital is building an operating system that lets non-CS domains like science and mechanical engineering use its team's computer science expertise to build complex apps and research tools. The author argues that clearly specifying what you want from AI is the key lever for getting value, and that robust results are possible without a CS degree if the right pieces are in place. A quoted post on an ETH Zurich study of 100 developers suggests computer science background predicts vibe coding success more strongly than writing skill.

Sep 6

Sep 6Sun
  1. Satya NadellaAI score34

    Copilot Autopilots now complete long-running multi-step work tasks

    AISatya Nadella says Microsoft is bringing new models into Copilot to handle increasingly complex work, from quick questions to delegated tasks and complete long-running jobs via Autopilots. As an example, an Opal-powered Autopilot on a secure Windows 365 Cloud PC sorts a month of trail cam footage, extracts species sightings, and builds a highlight reel, spreadsheet, PowerPoint, and Teams share.

Sep 5

Sep 5Sat
  1. AI at MetaAI score46

    AIRA₃ cuts GPU kernel latency 27% and reaches Kaggle gold level

    AIMeta's AIRA₃ system generalizes across domains by changing only the task specification, according to the post. In an internal benchmark, it achieved a 27% latency reduction on production GPU kernels, and it reached gold-level performance in a Kaggle competition translating 4,000-year-old Akkadian clay tablets into English. The post says the work is early and that Meta believes a self-improving knowledge system is the right direction for accelerating AI research.

  2. AI at MetaAI score43

    AIRA₃ coordinates long-running agents through a shared forum and filesystem

    AIMeta's AIRA₃ replaces a central controller with many long-running agents, each pairing a model with a coding harness in its own isolated environment. The agents coordinate asynchronously through a shared forum for hypotheses and findings and a shared filesystem for solution artifacts. According to the post, performance gains compound over time as agents build on each other's discoveries.

  3. AI at MetaAI score38

    AIRA₃ ensemble places 8th with gold-medal results in live competition

    AIMeta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

Sep 4

Sep 4Fri
  1. Andrew NgAI score42

    Andrew Ng maps key skills for using AI coding agents effectively

    AIAndrew Ng presented an AI Engineering Skills Map for using coding agents such as Claude Code, Codex, Cursor, OpenCode, and Pi. The workflow he describes covers planning, execution, and deployment with monitoring, and he identifies five key skills: directing the workflow, enabling agent autonomy, reviewing the work, customizing the agent and its environment, and coding agent foundations. The source says these skills matter more as the agents evolve quickly.

  2. Lewis TunstallAI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  3. Lewis TunstallAI score46

    Meta paper uses research preference models to guide AI agents' experiments

    AILewis Tunstall praises a new Meta paper on research preference models (RPMs), which instill "research taste" in agents by treating experiments as tree nodes. An RPM acts as an LLM judge that selects the most promising candidate experiment before it is run, reducing wasted compute. Tunstall notes the resulting trajectories could train domain-specific RPMs, which would be valuable in hard fields such as the natural sciences.

Sep 3

Sep 3Thu
  1. Jim FanAI score48

    Jim Fan says OpenAI's 2016 Universe ambitions now reincarnated as Astra

    AIJim Fan recalls that OpenAI's 2016 Universe project tried to have an agent learn computer use from screen pixels, mouse, and keystrokes, which he now calls doomed. He argues the solution is first training a Specialized Generalist across many general tasks, then specializing back to screen-level control, and he congratulates GPT-6 for reliably booking a United flight.

  2. Google Developers BlogAI score23

    Google's Gemini Enterprise DevEx sprint fixes governance setup friction for agents

    AIGoogle's Gemini Enterprise developer experience team tested agent governance workflows without internal shortcuts and fixed friction points across its agent governance products. Fixes included documentation stating that enabling the Identity-Aware Proxy API is a hard requirement, auto-allowing essential Google-managed platform APIs in the Agent Gateway, and adding Private Service Connect and Cloud DNS setup guidance for Semantic Governance. The team also published ready-made Logs Explorer queries for monitoring Agent Gateways and Content Security.

  3. xAI News (Grok)AI score44

    xAI's Haggle Bot finds over $100,000 in procurement savings across SaaS and supplies

    AIxAI built a Grok-powered procurement agent, Haggle Bot, that audits vendor spend, flags unused SaaS seats, and prepares renewal negotiations. The Bot has identified more than $100,000 in direct savings, including $14,220 from 43 unused seats of one SaaS product and $85,662 a year in unused SKUs from another. A person still approves any spending, contract terms, or messages sent to vendors.

  4. Michael TruellAI score47

    Grok Bot for enterprise launches today, free for two weeks to Grok and Cursor customers

    AIxAI has released Grok Bot for enterprise today, and it is free for all Grok and Cursor enterprise customers for the next two weeks. Cursor's Michael Truell says deployments have felt like onboarding thousands of capable teammates, calling it the most internally adopted and most powerful AI product the company has seen so far.

  5. Benedict EvansAI score36

    Benedict Evans on why AI won't simply replace enterprise software

    AIBenedict Evans argues that cheaper tool-building with AI will not automatically sweep away large companies' sprawling software, because people often don't see the tasks they could automate. He says the hard parts are knowing a tool is needed, deciding what it should do, and getting many departments and systems to adopt it. Companies typically move improvised, bottom-up workarounds into institutionalized software once they carry revenue and risk.

  6. Mark ChenAI score80

    Mark Chen announces GPT-6 Astra with computer use and agent oversight

    AIOpenAI researcher Mark Chen announced GPT-6 Astra, which he described as the company's most capable and aligned model yet. He said it can build and test software, work across apps on a computer, and help with open scientific problems. The post also highlights improved computer use compared with Operator and stronger monitoring that can stop potentially unauthorized agent actions.

    Why it matters: The post links a named model release to specific capabilities like computer use and aligned agent behavior, giving readers concrete claims to check against the model.

  7. Thomas DohmkeAI score38

    Copilot's new search tool understands codebase intent and decision context

    AIMicrosoft and Copilot's Thomas Dohmke announced a search tool that understands a codebase beyond literal phrases, returning results based on intent, semantic reasoning, and the context behind key decisions. The post's quoted Entire context describes Agentic Search, an API across accessible repos that returns the code, session, transcript, and prompt behind a change.

  8. Matei ZahariaAI score34

    Databricks uses Unity AI Gateway traces to cut AI waste fast

    AIDatabricks used Unity AI Gateway tracing and Genie One to find seven small MCP-server bugs and eliminate an estimated $1.2M in annual wasted AI spend and lost productivity within an hour. The bugs drove about $499K per year in wasted tokens, roughly 12,000 engineering hours per year in agent wait time, and 1,409 tool errors in a single 24-hour window. Matei Zaharia argues that analyzing tracing data for AI workloads will become a routine form of operational data analysis across companies, much like finance and security.

  9. BAAI · new models on Hugging FaceAI score25

    BAAI Releases Recon2Reason-Reasoning-4B, a Spatial Reasoning Vision-Language Model

    AIBAAI released Recon2Reason-Reasoning-4B, a 4,437,815,808-parameter vision-language model fine-tuned from Qwen3-VL-4B-Instruct for indoor spatial reasoning. The model handles metric distance, relative position, and object-relation questions from single or multiple images, and loads with the standard Qwen3VLForConditionalGeneration interface without trust_remote_code. The checkpoint is released under Apache-2.0 with BF16 Safetensors weights, and the retrieval-augmented scene-reconstruction extension ships separately.

Sep 2

Sep 2Wed
  1. Daniel HanAI score34

    Stanford's Modern Software Developer course adds AI-native engineering curriculum

    AIMihail Eric announced the 2026 edition of his Stanford course "The Modern Software Developer," with 85% of the Fall 2025 material replaced by AI-native topics such as agent skills, context engineering, and agentic code review. Students will ship pull requests to real open-source AI repositories, with partners including Browserbase, HeyGen, and CopilotKit offering mentorship.

  2. ARC PrizeAI score77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    AIOpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  3. Google AI StudioAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.