Skip to content
The AI news worth your attention

#Anthropic

Oct 8

  1. Claude BlogAI score67

    Block describes using Claude Fable to orchestrate thousands of pull requests

    Block's AI capabilities lead describes using Claude Fable to plan large code migrations and direct smaller models like Opus and Sonnet on individual tasks. He says Block routes frontier and smaller models by task and keeps merges and production deploys behind human dual approval.

    AIWhy it matters: Block's engineering lead describes how frontier models orchestrate large migrations and how access, effort levels, and safeguards are managed across an organization.

  2. Anthropic ResearchAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    Johns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    AIWhy it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

  3. Anthropic NewsroomAI score62

    Anthropic launches Cyber Mission with infrastructure defense and free OSS Scanner

    Anthropic has launched the Anthropic Cyber Mission, which starts with the Critical Infrastructure Defense Program for operational technology and OSS Scanner for open-source projects. The defense program brings frontier Claude models, on-site engineers and threat research to trusted providers such as Accenture, CrowdStrike and Palo Alto Networks. OSS Scanner gives enrolled open-source projects periodic free scans from its strongest models, with reports sent without human review and an expected true-positive rate above 90%.

    AIWhy it matters: The announcement shows how a frontier AI lab is packaging cyber defense around critical infrastructure and open-source maintainers, including the program's partners and access routes.

  4. Claude BlogAI score67

    Claude adds live dashboards and animated explainers, Docs and Slides leave beta

    Claude now turns company data into dashboards that stay current, and it can build animated explainers from a prompt. Dashboards connect to BigQuery, Databricks, Snowflake, and Salesforce in beta on paid plans, while Motion is in beta on Team and Enterprise. Docs, Slides, and Design are out of beta and available on every plan, including Free.

    AIWhy it matters: The post specifies which data platforms connect, which features move out of beta, and where admins control access, clarifying what changes for enterprise workflows.

Oct 7

  1. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    AIWhy it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  2. Claude BlogAI score66

    Claude skill commands build evals and hillclimb them against overfitting

    Anthropic added build-eval and hillclimb commands to its claude-api skill for designing evaluations and iteratively improving applications against them. The article covers eval design principles, including production-representative tasks, headroom and low variance, and guards against overfitting through train/test splits. Two examples report results: a customer support benchmark where cost fell to under half while accuracy rose, and a claude-api skill eval that rose from 66% to 88%.

    AIWhy it matters: The article gives a concrete workflow for designing evals and hillclimbing without overfitting, with two worked cost and performance examples that show the tradeoffs.

  3. Artificial Analysis ArticlesAI score60

    Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index

    Anthropic released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku release. Pricing is $0.10/$0.50 per 1M input/output tokens up to 100k tokens, rising to $0.50/$2.50 above that, but at max effort it uses about 162k output tokens per Intelligence Index task, roughly 3x GPT-6 Luna.

    AIWhy it matters: The benchmark shows Haiku 5.5 scores well but uses far more output tokens than GPT-6 Luna, so cost per task matters beyond list price.

  4. Claude BlogAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    Anthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

Oct 6

  1. Epoch AIAI score60

    Epoch AI finds frontier models fall short of an end-to-end AI research task

    Epoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation. GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope. The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.

    AIWhy it matters: The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.

  2. Claude Apps Release NotesAI score60

    Claude Haiku 5.5 launches as a fast, low-cost small model, and Max and Team plans gain monthly API credits

    Anthropic launched Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released, aimed at high-volume, cost-sensitive tasks. Max and Team plans now include monthly API credits for running their own apps and agents on the Claude Platform, rolling out over a few days. Users claim the credits by linking a Claude Console organization in Settings > Billing for Max or Organization settings > Billing for Team.

    AIWhy it matters: The notes name a new small model and a credit change for Max and Team plans, with the claim path, which matters for teams budgeting API use.

  3. Claude BlogAI score62

    Claude now works inside Google Docs, Sheets, and Slides in public beta

    Claude for Google Workspace is in public beta on all paid Claude plans, adding a sidebar to Google Docs, Sheets, and Slides. It can read the open file, edit text, build formulas, pivot tables, charts, and slides, and it asks for approval before changes unless the user chooses "Accept all edits." New Docs, Sheets, and Slides connectors in beta let Claude create and edit Google files from the chat, with access matching existing Google sharing permissions.

    AIWhy it matters: The source specifies how Claude edits Docs, Sheets, and Slides in place and where users keep control, which clarifies the practical workflow change.

  4. Claude BlogAI score62

    Comcast and Booz Allen use Claude Mythos to find exploit chains in codebases

    Comcast and Booz Allen used Claude Mythos Preview to find vulnerabilities that arise from interactions across code, configuration, and deployment rather than single-file bugs. Comcast identified a critical authentication flaw across 258 systems and about 170 million lines of code before any exploitation was observed. Booz Allen reported that one analyst reviewed eight production systems across 138 repositories in twelve days, a review its team estimated would have taken several months without the model.

    AIWhy it matters: The case studies show how security teams validate and remediate model-found exploit chains, a workflow relevant to anyone managing large codebases.

  5. Anthropic NewsroomAI score75

    Anthropic expands Cyber Verification Program into three tiered access levels

    Anthropic is launching an expanded Cyber Verification Program with three access tiers for qualifying security professionals, giving each tier different cyber capabilities and reduced blocking classifiers. On CyScenarioBench, Claude Opus 5.5 was blocked on 46 of 50 trials in the Defense Access tier, while the Red Team Access tier had no blocks and completed 34 of 50 tasks. Existing Project Glasswing members will move to the Specialized Access tier, and data retention is required for enrolled organizations.

    AIWhy it matters: The program lays out three verified access tiers with different cyber blocks, and its CyScenarioBench figures show how safeguards change what defenders can do.

Oct 3

  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    AIWhy it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Oct 2

  1. Epoch AI · The Epoch BriefAI score62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    Epoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    AIWhy it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

Oct 1

  1. Epoch AIAI score62

    Epoch AI estimates how many concurrent AI agents 2025–27 memory shipments could run

    Epoch AI estimates that high-bandwidth memory shipped in 2025–27 could eventually support about 30–170 million concurrent frontier-model agents once fully deployed and allocated. Using DeepSeek V4 Pro serving benchmarks, the estimate rises to about 1.9 billion concurrent agents. The authors compare the implied API-equivalent spending of $2.6–5.3 trillion per year with projected developer revenue of roughly $1 trillion by end-2027, suggesting demand may lag supply.

    AIWhy it matters: The analysis converts HBM shipment data into concurrent agent capacity and compares it with projected API revenue, showing where compute buildout may outpace demand.

  2. JetBrains AI BlogAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    JetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    AIWhy it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  3. Anthropic ResearchAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    Physicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    AIWhy it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

Sep 30

  1. Cloudflare Blog · AIAI score72

    Cloudflare launches Auto Router in AI Gateway to cut AI token spend

    Cloudflare has released Auto Router in public beta through AI Gateway, where setting the model to cloudflare/auto routes each request to a model judged capable enough for the task. Internal tests showed up to 30% cost savings against frontier models, and on a 97-task internal benchmark cloudflare/auto scored 86.6% at $0.0084 per success versus 96.6% at $0.0210 for Claude Opus 5.5. The router is free during beta.

    AIWhy it matters: The source gives a benchmark table of success rates and costs per trial, showing how routing trades quality against price for a gateway deployment.

  2. Artificial Analysis ArticlesAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    Artificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    AIWhy it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.