Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. Sam AltmanXAI score4

    Sam Altman thanks machines and reality for deeper understanding

    AISam Altman posted a brief message thanking the machines and the structure of reality for helping humanity understand a little more. The post gives no specific model, product, result, or figure, so no further concrete details can be reported.

  2. Claude BlogOfficialAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

Oct 6

Oct 6Tue
  1. Miles BrundageXAI score14

    Leap panel finds US strict liability for AI beats slowdown or authorization rules

    AIMiles Brundage called the result a "Weild" finding, referring to Gabriel Weil's work, and it appears to match a Forecasting Research Institute Leap panel's conclusion. According to Weil's quoted post, the panelists judged a US-only strict liability regime for AI to outperform a US-only slowdown or pre-release authorization regime, and to be competitive with globally coordinated versions of those policies.

  2. Yuchen JinXAI score12

    AI now solves hard math problems that GPT-4o once failed

    AIYuchen Jin notes that in 2024 GPT-4o famously got "Is 9.9 > 9.11?" wrong, while AI now appears poised to solve the hardest math problems. He describes the pace of progress as a wild time to be living through.

  3. Matt ShumerXAI score62

    OpenAI releases broad new math results from an internal frontier model

    AIOpenAI says it is releasing a broad range of new mathematical results produced by an internal frontier model. The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release them. The results are linked from a GitHub repository at

  4. Lewis Tunstall @ COLM 🌉XAI score12

    Lewis Tunstall doubts AI will soon crack unified field theory

    AILewis Tunstall says candidate theories for unifying physics already exist, so the real challenge is experimentally discriminating among them. He considers an AI-driven breakthrough in fundamental physics extremely unlikely, though he would welcome being proven wrong.

  5. Andrew CurranXAI score13

    Andrew Curran says AI reasoning generalizes broadly and keeps scaling

    AIAndrew Curran argues that the approach generalizes to everything and continues to scale. The post builds on Christian Szegedy's claim that mathematical reasoning will transfer to other complex, reasoning-heavy domains, filling data gaps with high sample efficiency.

  6. meng shaoXAI score30

    MIT 6.S950 Lecture 4 Explores Programming's Abstraction Ladder in the AI Era

    AIMIT's 6.S950 "Agency with AI" course has released Lecture 4, "The Abstraction Ladder (of Programming)," which compares today's prompt-driven coding with the 1957 FORTRAN paper by Backus et al. The lecture argues that the objections to vibe coding echo the arguments once raised against compilers, but natural-language "compilation" differs because the same prompt can yield different programs each time, unlike deterministic translation.

    Image from @shao__meng's post
  7. Joshua AchiamXAI score6

    Joshua Achiam says the end of ignorance is approaching

    AIJoshua Achiam posted that humanity is approaching the end of ignorance if it chooses to pursue that goal. The post is a brief, forward-looking remark that responds to a question about whether a unified field theory or quantum gravity might arrive within 12 months.

  8. Lewis Tunstall @ COLM 🌉XAI score25

    Beam leads open models in token efficiency, Chinese models lag

    AILewis Tunstall says Chinese open models are strong but token-inefficient, citing a plot from the Beam release at IMO. The background post from @reflection_ai says Beam is 3-4x more efficient than GLM 5.2 and over 4x more efficient than leading Western open models in inference. He hopes future open models will compete on this efficiency axis.

  9. Ethan MollickXAI score22

    AI may re-judge all published science and speed novel discoveries

    AIEthan Mollick argues that AI is likely to bring two revolutions after a brief period of low-quality "slop" science that eroded institutions. The first is that all previously published work will be re-read and re-judged in ways human scientists never anticipated. The second is that novel discoveries will start arriving quickly.

  10. Jerry LiuXAI score30

    Jerry Liu argues agentic OCR beats legacy systems on accuracy and cost

    AIJerry Liu argues that OCR, long dominated by brittle legacy systems, can be solved accurately and cheaply by applying agentic intelligence. He says a properly tuned agentic OCR dynamically allocates extra compute to complex elements, reviews and corrects failures, and builds semantic meaning across the page. He contends frontier models are overengineered for this task in cost and latency yet still struggle with complex edge cases.

    Image from @jerryjliu0's post
  11. Mike KnoopXAI score40

    AI now automates conceptual search and verification for new science

    AIMike Knoop argues AI can now automate conceptual search, transformation, and verification toward new science. He says AI can tell whether an open problem needs new ideas or whether the answer is already latent in existing knowledge. He calls this the most significant change in the philosophy of science since writing was invented about 6,000 years ago.

  12. Nathan LambertXAI score40

    OpenAI releases math results from an internal frontier model on GitHub

    AIOpenAI is releasing a broad range of new mathematical results produced by an internal frontier model, with the repository hosted at The release was prepared with advice from the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study. The main post itself only comments on the humor of the repository's name.

  13. whXAI score58

    OpenAI's Math Results Are About 20% Disproofs and Counterexamples

    AIA breakdown of OpenAI's released internal-model math results shows about 73 disproofs and counterexamples, roughly 20% of the total. The author argues this counters claims that recent math breakthroughs are concentrated in counterexamples because models are only good at brute-force search.

    Image from @nrehiew_'s post
  14. 👩‍💻 Paige BaileyXAI score20

    Paige Bailey shares a brief note on AI progress

    AIPaige Bailey's post says only "slowly, slowly, then all at once," with no model names, figures, or specific claims. It quotes Will DePue, who says he asked GPT 6 Pro and Fable 5.1 to rank discoveries from the last three years and reports that 81% of them were released today.

  15. will depueXAI score35

    Will Depue Surprised AI Labs' Math Results Have Held Up So Far

    AIWill Depue, an OpenAI-affiliated account, says he is surprised that AI lab math results have so far contained no profound errors or real bugs, which he notes is unlike typical human work. He expects at least a couple of today's results will not survive scrutiny.

  16. François CholletXAI score38

    Chollet asks if AI's jagged frontier is driven by math, code, and RLVR

    AIFrançois Chollet asks whether the jagged frontier of AI capability is mainly math and code, which can be pushed far with RLVR. He questions whether steady gains in non-verifiable areas come from higher generalization driven by RLVR or only from continued injection of new human data.

  17. Sam AltmanXAI score30

    OpenAI shares AI progress in mathematics discovery

    AIOpenAI has published a post on sharing its AI progress in mathematics, which Sam Altman says marks the start of a new era of discovery. The post text provides no further details on specific results, models, or benchmarks.

  18. PlatformerBlogAI score49

    Anthropic and OpenAI Leaders Weigh Hard Caps on AI Intelligence

    AISpeakers at The Curve, a Berkeley AI conference, discussed limiting how intelligent large language models can become, amid concerns over recursive self-improvement. Proposed approaches include Anthropic's responsible scaling policy, limits on compute and model copies, and restrictions on using frontier models for AI research. The column notes such enforcement tools do not yet exist and that the Trump administration opposes such restrictions.

  19. Tomasz TunguzBlogAI score46

    OpenAI's Price Cuts Signal AI Models Are Becoming Commodities

    AIOpenAI cut Luna prices by more than 80% to win share, while frontier models' share of tokens slipped from 53% in August into the mid-40s as buyers shifted to cheaper tiers. The author argues that in this commoditizing market, value accrues to platforms that control distribution and aggregate usage rather than to labs with marginal benchmark leads.

  20. TechRadar · AINewsAI score50

    AWS warns that 100 proposed data center bans could harm the US for generations

    AIAWS CEO Matt Garman warned that the more than 100 American communities considering moratoriums on new data centers could leave the US paying for the decision for decades. A Brookings report estimates US data center and AI infrastructure investment could total $10.3 trillion from 2025 to 2032, and Amazon announced a $1 billion-plus Built Together community program over five years.

  21. Epoch AIOfficialAI score47

    GPT-6 Astra Hit 100% on EBR-bench Using a Card That Bypassed Its Time Limits

    AIEpoch AI reports that GPT-6 Astra scored 100% on the original EBR-bench by exploiting a card that bypasses the game's time-constraint expectations, so Epoch has banned that card from the default setting. Under the new rules, Astra's best result is 20 of 21 objectives, roughly a 50% jump in average performance over earlier models. Epoch will report revised scores only for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future models.

  22. Alexander DoriaXAI score16

    Alexander Doria jokes that post-Deep Blue, humans are now chess players

    AIAlexander Doria says that after Deep Blue, humans are effectively chess players now, likening their position to that of machines in chess. The post quotes a reply that questions how a claimed sub-quadratic algorithm running in n^1.9992 time could exist, despite any proof.

  23. Dongxi NLPXAI score22

    OpenAI releases Openai/math, suggesting verifiable problems are being solved

    AIOpenAI has published a repository called Openai/math, which the author reads as a sign that math problems, or any verifiable problems, are being solved. The author says OpenAI's tools exhausted their Pro token allowance on subagent tests unrelated to their main task, concluding that the work was aimed at verification for its own sake.

    Image from @dongxi_nlp's post
  24. will depueXAI score12

    Will DePue asks where AI will be in five years

    AIOpenAI-affiliated researcher Will DePue asked where AI will stand five years from now, without offering a specific prediction. The post was a brief prompt, and the quoted context notes that OpenAI released its grade school math dataset five years ago, a benchmark that AI systems then struggled with.

  25. Thomas WolfXAI score22

    Ben Affleck jokes about convolutions and his AI background

    AIThomas Wolf's post is a short, playful reply: "how do you like them convolutions," apparently referencing Ben Affleck's comments on convolutional neural networks. The quoted context reports Affleck describing his Python scripting, understanding of CNNs and tensors, GPU work, and private looks at Google and OpenAI's video models.

  26. Joshua AchiamXAI score14

    Joshua Achiam praises a thoughtful essay on AI and human agency

    AIJoshua Achiam recommends a deep, carefully considered piece on some of the thorniest problems of our time, regardless of whether readers agree with its prescription. The quoted post, from @satpugnet, presents Phase Lock, a six-month manifesto on how brain-computer interfaces could help align AI and preserve human agency.