Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. Sophia YangXAI score7

    Mistral 4 claims strong intelligence per GPU versus rival models

    AIMistral's Sophia Yang shares a post comparing GPU counts for recent models, noting Mistral 4 used 3,800 Grace Blackwell GPUs. The post cites GPT-6 Astra at 100,000+ Grace Blackwell and estimates for Grok 4.7 and Claude Opus 5.5. The post frames this as evidence of Mistral's efficiency in intelligence per GPU.

  2. laurenXAI score42

    Lauren Tan proposes "time to rewrite" as a heuristic for agent-readiness

    AILauren Tan (@poteto) proposes "time to (fully automated, hands-off) rewrite" (TTR) as a rough thought-experiment heuristic for how well a codebase is set up for agents. She suggests asking how long a single engineer would need to rewrite the code in another language, framework, or architecture, since the answer surfaces gaps like missing verification that agents can use to confirm user-visible behavior matches. The post also raises questions about whether a rewrite would improve, maintain, or regress performance and maintainability over time.

  3. Boris ChernyXAI score47

    Anthropic's Claude Haiku 5.5 offers 100k context at 10x lower cost

    AIAnthropic's Claude Haiku 5.5 is described as a good Haiku, with a 100k token context window and roughly 10x lower cost than Claude Haiku 4.5. The quoted Claude announcement says it is the cheapest, fastest, and most capable small model Anthropic has released, costing about 75% less to run on average than Haiku 4.5.

  4. MuseOfficialAI score10

    Muse fills out a city trash-can replacement request for a user

    AIKevin Xu says Muse, asked this morning how to replace a trash can knocked over overnight, filled out the city government's request form. Muse then sent a confirmation number to his inbox, which he describes as feeling like AGI. A quoted post from Muse's account adds only light trash-talk about the exchange.

  5. Amazon ScienceOfficialAI score10

    Matthew Lease explores harnessing AI for scientific discovery and risks

    AIAmazon Scholar and University of Texas at Austin professor Matthew Lease will give an Expo Talk at COLM on Thursday at 1pm PT. The talk explores how to harness AI for scientific discovery while assessing potential risks, drawing on work from UT's Good Systems and the Cosmic AI Institute.

    Video from @AmazonScience's post
  6. Epoch AIOfficialAI score18

    AI models independently invent a technique resembling SDPO

    AIThe post argues that AI models, lacking any information about SDPO, had to either independently invent something similar or devise another technique with comparable benefits under the same constraints. The post does not identify the specific models or experiment involved.

  7. Ethan MollickXAI score12

    Ethan Mollick predicts many small singularities across fields, not one

    AIEthan Mollick argues that the original meaning of "singularity" is a point where the assumption that the future resembles the past breaks down. He is unsure whether one big Singularity will occur, but he considers a million smaller singularities across fields inevitable.

  8. Lucas Beyer (bl16)XAI score38

    Robotics progress accelerates, partly driven by coding model advances

    AILucas Beyer says robotics is accelerating in the physical world, not just in AI research. He attributes this partly, though not only, to progress in coding models over the past year. The post cites a related thread on scaling a UMI data collection operation from 5 to 90 operators and reaching over 1M unique tasks.

  9. Allie K. MillerXAI score14

    Allie K. Miller urges giving AI systems big North Star goals

    AIAllie K. Miller argues that users rarely give their AI systems long-term North Star goals, distinct from task-specific instructions. She says that if AI is to act as a proactive support system, it should be steered toward the user's larger aspirations, such as owning a dog within a year.

  10. Gergely OroszXAI score31

    Samuel Newman on why LLMs aren't world models and lack causality

    AISam Newman argues the tech world misunderstands LLMs because they have no concept of causality, so "if I do A, B happens" reasoning is absent. He contends LLMs are not world models, unlike older world-model approaches that could in principle track cause and effect. He adds that people overestimate LLM capabilities because they seem smart, and that guardrails are unlikely to be the right long-term fix.

    Video from @GergelyOrosz's post
  11. elvisXAI score14

    Companies Turn to RL and Specialized Models Over AGI

    AIThe post argues that companies are increasingly recognizing business opportunities from reinforcement learning, since many real-world tasks need specialized models, harnesses, and data flywheels rather than AGI. It predicts a new post-training era led by full-stack AI companies, though it provides no concrete figures or raw data to support the claim.

  12. Marcus on AIBlogAI score62

    Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality

    AIGary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.

  13. elvisXAI score22

    Elvis Saravia describes building personal multi-agent teams with Opus 5.5

    AIElvis Saravia reports that agent-to-agent communication with a personal agent, built on models like Opus 5.5, is already coordinating work faster and at higher quality than he can match. He describes progressing from individual Claude Code sessions to subagents, then a persistent team of eight specialized bots with his own orchestrator. He argues everyone should build a personalized agent orchestrator and says most apps like Code and Claude Desktop are behind.

    Image from @omarsar0's post
  14. Mark ChenXAI score46

    OpenAI's Navier-Stokes progress marks a decade of math advances in a week

    AIMark Chen says the Navier-Stokes achievement matters more for the figure it shows than for the problem itself, representing a decade of mathematical progress in a single week. He says he is eager to apply these tools to life sciences, the building of OpenAI's next models, and alignment research.

    Image from @markchen90's post
  15. DeedyXAI score46

    OpenAI's math results spark claims of AGI and Millennium Prize progress

    AIDeedy argues LLMs have made substantial progress on four of the seven Millennium Prize problems, including a claimed Navier-Stokes result, conditional on verification. He says OpenAI's results averaged only 3 hours of thinking compute on unreleased models. He concludes that by most definitions of AGI, we have already achieved it.

  16. Andrew CurranXAI score42

    Andrew Curran says OpenAI's 722 math results omit cryptography breakthroughs

    AIAndrew Curran notes OpenAI's 722 published mathematical results show striking under-representation of cryptographic breakthroughs, and he says he has personally witnessed US government censorship of academic quantum cryptanalysis results. He calls backroom government interventionism his base case and says rumors suggest yesterday's OpenAI math release was only the first of three batches.

    Image from @AndrewCurran_'s post
  17. Fast Company · AINewsAI score14

    IMF Chief Urges Global Action on AI Regulation and Debt

    AIIMF Managing Director Kristalina Georgieva called on nations to urgently address AI regulation and debt curbing ahead of the IMF-World Bank meetings in Bangkok. The source provides no further detail on specific measures or timelines.

  18. dexXAI score12

    Dex Horthy's talk on the state of the software factory

    AIDex Horthy (@dexhorthy) shared a recording of his talk "state of the software factory," delivered at the Agentic AI Foundation event in Amsterdam last week. The post links to the video and urges viewers not to be the last to watch it.

  19. indigoXAI score34

    Grok Bot acts as a model router, using Gemini and Opus together

    AIThe poster says they already use Grok Bot as a model router, citing last weekend's personal agent livestream. In the demo, Gemini produced an infographic inside Grok Bot, and Claude Opus then checked the content. This follows Elon Musk's announcement that Grok Bot will use the best backend model for each task, including Claude Opus 5.5, MidJourney, and Suno.

    Video from @indigox's post
  20. ARC PrizeOfficialAI score10

    ARC Prize announces Alexandre Bouayad as 2026 Research Summit speaker

    AIARC Prize has announced Alexandre Bouayad as a speaker for its ARC Prize Research Summit 2026. Bouayad is a quantitative researcher at G-Research and author of "Weave of Formal Thought," which applies formal structure to how language models learn and generate code.

    Image from @arcprize's post
  21. jasonXAI score10

    Jason Liu's demo of Dots shows AI-generated travel schedule planning

    AIJason Liu was embarrassed when a demo of Dots turned into a request for his week's tasks, and it returned a plan to book 10 plane tickets over 50 days and rent three cars. He said he had asked it for his travel schedule and call times across two production shoots and two talks in four cities.

  22. Simon WillisonBlogAI score14

    Michael Lynch lists anti-patterns in software blogging, from meandering intros to overly formal prose

    AIMichael Lynch warns software bloggers against meandering intros, misjudging reader knowledge, assuming readers have read earlier posts, excessive formality, and overreliance on links instead of explaining terminology. He advises that an article should still make sense even if readers click no links. Simon Willison endorses the advice and argues that writing in one's own voice matters as more developers delegate writing to AI.

  23. Aravind SrinivasXAI score20

    Perplexity's Aravind Srinivas Celebrates AI-Generated Animated Scene Workflow

    AIAravind Srinivas posted that "we're living in incredible times," highlighting a fully computer-made animated scene produced from concept art to final edit. The quoted post says the scene was planned in Blender, with reference images generated by Nano Banana and animation and score produced by Seedance 2.5.

  24. Jim FanXAI score12

    Jim Fan predicts AI will solve Riemann hypothesis before physical Turing test

    AIJim Fan argues that humanity will likely solve the Riemann hypothesis before robots pass the Physical Turing Test, where a person cannot tell whether a human or robot cleaned a home after a party. He suggests people underestimate Moravec's paradox, the principle that physical tasks humans find easy are hard for machines.

  25. Latent.SpaceXAI score34

    Stacklok's Kubernetes creators aim to move agent harnesses fully to cloud

    AIStacklok, founded by two Kubernetes creators, Craig McLuckie and Joe Beda, is pursuing a "cloud-native harness" to bring AI agent harnesses fully into the cloud. The post argues that cloud-based agent harnesses from OpenAI and Anthropic are not yet fully solved, and points to a Latent Space interview with the founders.

    Image from @latentspacepod's post
  26. Exponential ViewBlogAI score72

    OpenAI's 722 machine-generated math results may split mathematics into two layers

    AIOpenAI released 722 mathematical manuscripts in 372 families, produced by an unreleased frontier model, with the average result taking the equivalent of three hours of ChatGPT Pro thinking. The author notes many results are verified in Lean but not all, and suggests mathematics could divide into vast machine-verified work and a compressed human 'effective theory' that people can actually understand.

  27. 👩‍💻 Paige BaileyXAI score10

    Paige Bailey urges values-aligned AI model evaluation nonprofits

    AIPaige Bailey agrees with calls for a Christian METR, and suggests creating nonprofits that evaluate models for values or philosophical alignment. She notes existing benchmarks such as Gloo's flourishing AI initiative, VirtueBench, and FaithGPT.