Skip to contentSkip to stories
Updated

#Expert opinion

Oct 9

TodayOct 9Fri
  1. The Verge · AINewsAI score72

    Mathematicians say OpenAI's mass release of AI-generated results will take years to digest

    AIOpenAI released nearly 400 AI-generated results spread across more than 700 manuscripts in several branches of mathematics. Mathematicians told The Verge that only 300 of 719 manuscripts had been formalized in Lean, and that verification and understanding could take years. Several researchers said some results may warrant top-tier publication, while others raised concerns about paper quality, attribution, and disruption to early-career researchers.

    Why it matters: The article records how mathematicians assessed the volume, verification gaps, and disruption of OpenAI's mass release of AI-generated math results, useful for understanding the research community's reaction.

Oct 8

Oct 8Thu
  1. The DecoderNewsAI score80

    Mathematicians call for OpenAI boycott after AI-generated proofs flood the field

    AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.

    Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.

  2. Lewis Tunstall @ COLM 🌉XAI score60

    Physicist credits GPT-6 Astra for a chiral fermion proof in the Standard Model

    AILewis Tunstall reposts a post by Kyle Cranmer describing a paper by Nate, currently on leave at OpenAI, on non-perturbative simulation of chiral fermions in the Standard Model. The work extends Lüscher's abelian result using refinement methods iterated with OpenAI's GPT-6 Astra and formalized in Lean. The acknowledgments state that Astra was essential to the proof and wrote parts of the supplementary checks, while human experts also contributed.

    Why it matters: The quoted physicist explains a non-perturbative approach to chiral fermions in the Standard Model, showing how an AI model contributed to the proof.

    Image from @_lewtun's post
  3. Claude BlogOfficialAI score67

    Block describes using Claude Fable to orchestrate thousands of pull requests

    AIBlock's AI capabilities lead describes using Claude Fable to plan large code migrations and direct smaller models like Opus and Sonnet on individual tasks. He says Block routes frontier and smaller models by task and keeps merges and production deploys behind human dual approval.

    Why it matters: Block's engineering lead describes how frontier models orchestrate large migrations and how access, effort levels, and safeguards are managed across an organization.

Oct 7

Oct 7Wed
  1. Claude BlogOfficialAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

Sep 29

Sep 29Tue
  1. Microsoft ResearchOfficialAI score75

    Microsoft Research introduces Quine, a multimodal biology world model and research harness

    AIMicrosoft Research introduced Quine, an experimental research system combining a multimodal world model of biology with an interactive harness that connects models, scientific tools, literature, and researchers. In a pancreatic cancer study with the Broad Institute, Quine prioritized compounds that shifted tumor cell states, and several top-ranked candidates were validated in wet-lab assays. Access is initially limited to the Quine Fellows program and select collaborations, and the system is intended for research use only, not clinical use.

    Why it matters: The post shows how a multimodal biology world model is wired into a harness, grounded in one wet-lab cancer example and a limited fellows-program access path.

Sep 25

Sep 25Fri
  1. AnthropicOfficialAI score78

    Claude solves a nine-loop scattering amplitude problem beyond the eight-loop record

    AIAnthropic reports that Claude solved a nine-loop scattering amplitude problem in planar N=4 super-Yang-Mills, surpassing the previous eight-loop record set by SLAC's Lance Dixon and collaborators. Working largely unsupervised for days from a single prompt, at a total cost of a few thousand dollars, Claude used methods developed by Dixon's group, and Dixon independently verified the result.

    Why it matters: The post shows Claude solving a nine-loop physics calculation beyond the previous eight-loop record, verified independently, which bears on AI use in theoretical physics research.

Sep 23

Sep 23Wed
  1. Anthropic · YouTubeOfficialAI score65

    Anthropic launches a molecular biology lab where Claude hunts for unusual proteins

    AIAnthropic is introducing a molecular biology research group and lab to test whether Claude can help scientists find unusual proteins. Claude combs through large DNA datasets, flags uncharacterized proteins, and passes its most promising ideas to scientists, who test them at the bench. In one early program, Claude discovered a novel enzyme system with CRISPR-like repeats.

    Why it matters: The source shows Claude being used in a wet-lab workflow, from scanning DNA datasets to flagging proteins for scientists to test at the bench.

Sep 22

Sep 22Tue
  1. METR BlogOfficialAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

That’s everything