Skip to contentSkip to stories
Updated

#Expert opinion

Oct 8

Oct 8Thu
  1. The DecoderNewsAI score80

    Mathematicians call for OpenAI boycott after AI-generated proofs flood the field

    AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.

    Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.

  2. Lewis Tunstall @ COLM 🌉XAI score60

    Physicist credits GPT-6 Astra for a chiral fermion proof in the Standard Model

    AILewis Tunstall reposts a post by Kyle Cranmer describing a paper by Nate, currently on leave at OpenAI, on non-perturbative simulation of chiral fermions in the Standard Model. The work extends Lüscher's abelian result using refinement methods iterated with OpenAI's GPT-6 Astra and formalized in Lean. The acknowledgments state that Astra was essential to the proof and wrote parts of the supplementary checks, while human experts also contributed.

    Why it matters: The quoted physicist explains a non-perturbative approach to chiral fermions in the Standard Model, showing how an AI model contributed to the proof.

    Image from @_lewtun's post
  3. Claude BlogOfficialAI score67

    Block describes using Claude Fable to orchestrate thousands of pull requests

    AIBlock's AI capabilities lead describes using Claude Fable to plan large code migrations and direct smaller models like Opus and Sonnet on individual tasks. He says Block routes frontier and smaller models by task and keeps merges and production deploys behind human dual approval.

    Why it matters: Block's engineering lead describes how frontier models orchestrate large migrations and how access, effort levels, and safeguards are managed across an organization.

Oct 7

Oct 7Wed
  1. Exponential ViewBlogAI score72

    OpenAI's 722 machine-generated math results may split mathematics into two layers

    AIOpenAI released 722 mathematical manuscripts in 372 families, produced by an unreleased frontier model, with the average result taking the equivalent of three hours of ChatGPT Pro thinking. The author notes many results are verified in Lean but not all, and suggests mathematics could divide into vast machine-verified work and a compressed human 'effective theory' that people can actually understand.

    Why it matters: The piece links OpenAI's batch of machine-generated proofs to a possible split between machine-verified mathematics and the smaller human-understandable layer.

  2. Claude BlogOfficialAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

Oct 6

Oct 6Tue
  1. clem 🤗XAI score62

    Mistral Large 4 announced with API access today and open weights due end of October

    AIMistral announced Mistral Large 4, a natively multimodal model with 1T parameters and 49B active parameters. It is available via API today, with open weights planned for the end of October. Clément Delangue, Hugging Face's CEO, reacted by noting that the model cannot be the best open-weight model until its weights are actually released.

    Why it matters: The quoted announcement gives the parameter scale, active count, and availability path, which help readers compare it with other open-weight releases.

  2. Yuchen JinXAI score72

    Mistral Large 4 launches as a 1T-parameter multimodal model with open weights due end of October

    AIMistral announced Mistral Large 4, a natively multimodal model with 1T parameters and 49B active, available via API today. Mistral claims it is the best open weights model from the US or Europe on aggregated benchmarks, with open weights set for release at the end of October. The author quotes this claim and comments that it appears to beat GLM-5.3.

    Why it matters: The quoted announcement gives specific size, multimodal, and deployment claims, with open weights promised later, useful for comparing it against other open models.

Sep 30

Sep 30Wed
  1. Jensen HuangXAI score62

    Industry leaders sign White House Accord on Super Intelligence safety commitments

    AIJensen Huang says leaders across the industry signed the White House Accord on Super Intelligence at the White House. The accompanying document says each company should run internal controls, an independent external auditor, and board-level oversight for its frontier models.

    Why it matters: The source gives the accord's four layers of controls and audits, showing how the signing parties plan to verify frontier model safety in practice.

    Image from @JensenHuang's post

Sep 29

Sep 29Tue
  1. Microsoft ResearchOfficialAI score75

    Microsoft Research introduces Quine, a multimodal biology world model and research harness

    AIMicrosoft Research introduced Quine, an experimental research system combining a multimodal world model of biology with an interactive harness that connects models, scientific tools, literature, and researchers. In a pancreatic cancer study with the Broad Institute, Quine prioritized compounds that shifted tumor cell states, and several top-ranked candidates were validated in wet-lab assays. Access is initially limited to the Quine Fellows program and select collaborations, and the system is intended for research use only, not clinical use.

    Why it matters: The post shows how a multimodal biology world model is wired into a harness, grounded in one wet-lab cancer example and a limited fellows-program access path.

Sep 28

Sep 28Mon
  1. Alex AlbertXAI score62

    Claude Sonnet 5.5 Is Faster and Cheaper Than Sonnet 5, Per Anthropic

    AIAnthropic has introduced Claude Sonnet 5.5, the second model in the Claude 5.5 family, as a clear upgrade over Sonnet 5. The announcement says it runs more than 30% faster and costs up to 30% less for most work. Alex Albert, quoting the announcement, says the model writes clearly, is very fast, and makes a major capabilities jump over Sonnet 5.

    Why it matters: The quoted announcement gives concrete speed and cost figures for Sonnet 5.5, which helps readers weigh it against Sonnet 5 for everyday work.

Sep 25

Sep 25Fri
  1. AnthropicOfficialAI score78

    Claude solves a nine-loop scattering amplitude problem beyond the eight-loop record

    AIAnthropic reports that Claude solved a nine-loop scattering amplitude problem in planar N=4 super-Yang-Mills, surpassing the previous eight-loop record set by SLAC's Lance Dixon and collaborators. Working largely unsupervised for days from a single prompt, at a total cost of a few thousand dollars, Claude used methods developed by Dixon's group, and Dixon independently verified the result.

    Why it matters: The post shows Claude solving a nine-loop physics calculation beyond the previous eight-loop record, verified independently, which bears on AI use in theoretical physics research.

Sep 23

Sep 23Wed
  1. Anthropic · YouTubeOfficialAI score65

    Anthropic launches a molecular biology lab where Claude hunts for unusual proteins

    AIAnthropic is introducing a molecular biology research group and lab to test whether Claude can help scientists find unusual proteins. Claude combs through large DNA datasets, flags uncharacterized proteins, and passes its most promising ideas to scientists, who test them at the bench. In one early program, Claude discovered a novel enzyme system with CRISPR-like repeats.

    Why it matters: The source shows Claude being used in a wet-lab workflow, from scanning DNA datasets to flagging proteins for scientists to test at the bench.

Sep 22

Sep 22Tue
  1. Boris ChernyXAI score62

    Claude Opus 5.5 ports HAProxy to Rust faster and cheaper than Fable 5.1

    AIAnthropic introduced Claude Opus 5.5 as the first model in its Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 for most tasks at 40% lower run cost than Opus 5. Boris Cherny reports that Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of its tests, with Opus 5.5 finishing in 9.5 hours versus 12 hours and at 51% less cost.

    Why it matters: The author reports a same-task comparison in which Opus 5.5 finished a HAProxy C-to-Rust port faster and cheaper than Fable 5.1, offering a concrete cost and time benchmark.

  2. Alex AlbertXAI score62

    Anthropic introduces Claude Opus 5.5 as first model in Claude 5.5 family

    AIAnthropic has introduced Claude Opus 5.5, the first model in its new Claude 5.5 family. The quoted announcement says it performs at the level of Claude Fable 5.1 for most tasks and costs 40% less to run than Opus 5. Alex Albert's post praises the model as smart, clear, fast, and cheaper, but offers no independent test results.

    Why it matters: The quoted announcement gives a concrete cost comparison against Opus 5, which helps readers weigh the model's value beyond the author's praise.

  3. METR BlogOfficialAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

Sep 12

Sep 12Sat
  1. Demis HassabisXAI score62

    Demis Hassabis backs Dario Amodei's essay calling for AI industry to slow down

    AIDemis Hassabis says Dario Amodei's essay, which argues the AI industry should slow down, points toward the right path, though the details still need working through. He also points to Google DeepMind's recent proposal for an industry-wide standards body for frontier AI. The quoted essay describes a three-part plan, and Anthropic is committing to give third-party evaluators permanent, employee-level access to its systems.

    Why it matters: The post endorses Dario Amodei's essay on slowing AI frontier development, offering a short signal of where a major lab leader stands on industry pacing.

  2. John SchulmanXAI score62

    John Schulman Welcomes Third-Party Evaluator Access Commitments from OpenAI and Anthropic

    AIJohn Schulman praises embedding third-party evaluators as a big positive development and says OpenAI agreed to do it as well. He notes that the idea of pacing the frontier has spread quickly, following Dario Amodei's essay on slowing down AI development.

    Why it matters: The post reacts to a specific commitment that third-party evaluators get employee-level access, a concrete step within the broader pacing debate about frontier AI.

  3. Jakub PachockiXAI score62

    Dario Amodei essay calls for AI industry to pace the frontier

    AIDario Amodei has written an essay arguing that the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first step by giving third-party evaluators permanent, employee-level access to its systems. The evaluators can verify adherence to safety measures, report incidents, and assess model alignment during training.

    Why it matters: The quoted essay proposes a three-part slowdown plan, and Anthropic's pledge of permanent third-party evaluator access is the concrete step to examine.

That’s everything