Skip to contentSkip to stories

Updated

#Expert opinion

Items with an AI score under 20 are hidden. Show low-relevance items

Jul 7

Jul 7Tue

Jul 6

Jul 6Mon

Jul 3

Jul 3Fri
  1. Lil'Log (Lilian Weng)AI score62

    Lilian Weng surveys harness engineering as a path to recursive self-improvement

    AIThe post argues that the system surrounding a base model, called the harness, increasingly determines how well AI agents deploy and improve. It reviews research where harness components such as workflows, context, and code are optimized automatically through evolutionary search and meta-agent loops. The author concludes that evaluators, memory management, and human oversight remain open bottlenecks.

  2. Arthur MenschAI score34

    Mistral argues enterprises need open models and their own data for AI growth

    AIMistral CEO Arthur Mensch says enterprises should use open-source models because closed providers that force data retention gain leverage over their business. He argues companies should store data in open systems, control AI access rules, and build continuous training loops to shrink costs and create hard-to-copy systems. Mistral offers its Studio control plane and Forge training platform, deployed on customer infrastructure or through zero-data-retention hosting.

Jul 2

Jul 2Thu

Jun 30

Jun 30Tue
  1. John SchulmanAI score38

    Bridgewater fine-tuning with expert data beats prompting-only approaches

    AIJohn Schulman argues that fine-tuning with the right data, such as expert judgments, can substantially outperform prompting-only approaches even as general-purpose models improve. He cites Bridgewater's work, where an expert-labeled dataset and on-policy distillation were used to fine-tune a model to triage financial documents reliably and cheaply.

  2. One Useful Thing (Ethan Mollick)AI score62

    Ethan Mollick argues AI is shifting from chatbots to long-running agents

    AIMollick argues AI capability is improving at a better-than-exponential rate, citing METR, GDPval, Epoch, and his own tests showing models working autonomously for hours. He says work is shifting from co-working with chatbots to assigning tasks to agents, with OpenAI workers managing multiple agents and experts getting the most from them. He adds that open-weights Chinese models trail the American frontier by roughly 6-12 months.

  3. Werner VogelsAI score22

    Werner Vogels says two-pizza teams are about ownership, not food

    AIAmazon CTO Werner Vogels argues that the "two-pizza" team concept was never about feeding engineers but about ownership, speed, and avoiding bureaucracy. He says working backwards from the customer and writing documents to force clarity remain core practices. He adds that the industry is changing and it is time to reconsider how products are brought to life.

  4. Tri DaoAI score53

    Tri Dao Praises Etched's Fast Inference Chip Design for LLM Serving

    AITri Dao says Etched designed and produced its chips within two years by hardcoding attention into silicon and reaching high MFU. He expects hardware built for LLM inference to cut the cost of intelligence by 10x. The quoted Etched post says it has built its first racks after an A0 tapeout, raised $800m, holds $1B+ in customer contracts, and plans to ship the racks this summer.

Jun 29

Jun 29Mon
  1. Hamel HusainAI score54

    Why Hard-to-Eval AI Products Need Designs That Support Verification

    AIHamel Husain argues that an AI product whose output is hard to verify is a product design problem, not just an evaluation problem. He shows before-and-after sketches for an AI data agent, a PE lesson planner, and a workers' compensation report tool, each adding provenance, scoped edits, and checkable evidence. He notes that designing for verification also makes evals easier to build and grade.

Jun 26

Jun 26Fri
  1. HyperdimensionalAI score62

    Dean W. Ball proposes private audits and certification for frontier AI labs

    AIDean W. Ball argues that the current government restrictions on frontier model releases amount to a de facto preapproval regime without a known safety standard. He proposes that independent verification organizations audit labs against their own safety frameworks, with government certifying or licensing the auditors. The post also argues that broad distribution of frontier AI is needed to learn what good safety practice looks like.

Jun 23

Jun 23Tue

Jun 19

Jun 19Fri
  1. Andrew NgAI score72

    Andrew Ng says Anthropic and U.S. export controls on Fable expose AI access risks

    AIAndrew Ng argues that Anthropic's restrictions on building competing LLMs and a U.S. Commerce Department license requirement for foreign nationals led Anthropic to disable Fable access worldwide. He says this shows governments and providers can quickly cut off access to frontier AI, which may push nations and businesses toward sovereignty efforts and open-source alternatives, though training frontier models remains difficult.

    Image from @AndrewYNg's post

Jun 18

Jun 18Thu
  1. HyperdimensionalAI score49

    Dean Ball joins OpenAI as Head of Strategic Futures to shape frontier AI policy

    AIDean Ball will join OpenAI on July 6 as Head of Strategic Futures, a new small team reporting to Chief Strategy Officer Jason Kwon that will shape frontier AI policy on catastrophic risk, recursive self-improvement, labor market impact, and government relations. Ball says he will keep writing independently at Hyperdimensional, with no OpenAI preapproval or editorial discretion over his work.

Jun 17

Jun 17Wed
  1. John SchulmanAI score40

    PPO's LLM-era revival and the unexpected reasons behind it

    AIJohn Schulman says PPO gained a second wave in the LLM era for reasons not anticipated in the original paper. He points to the importance-ratio objective, which corrects biases from numeric error, asynchronous training, and forward-pass noise, and to the clipping objective, whose effect on entropy was unknown at publication, citing DAPO's arXiv paper.

Jun 16

Jun 16Tue
  1. HyperdimensionalAI score63

    Dean Ball argues the Anthropic Fable dispute shows frontier AI needs a governance framework

    AIDean Ball analyzes the Trump Administration's export controls on Anthropic's Fable and Mythos models after a jailbreak and a refused de-deployment request. He argues that the episode shows the need for a technocratic framework that separates political judgments about fairness from technical judgments about threats, in place of ad hoc executive action.

  2. Arthur MenschAI score44

    Mistral says its upcoming models will all be open-weight

    AIMistral states that this model and upcoming ones will be open-weight. The company argues that open weights are critical for customer confidence and for research and developer communities. It contends that systems reachable only through someone else's interface cannot be owned, inspected, audited, or improved, especially if data recording can no longer be turned off.

Jun 15

Jun 15Mon
  1. BAAIAI score22

    Turing Award winners Diffie and Barto keynote BAAI Conference on AI security and RL

    AITuring Award winners Whitfield Diffie and Andrew Barto delivered keynotes at the BAAI Conference on AI security and reinforcement learning. Diffie argued that today's feedback-based approach only patches programs after they fail, and that formal methods offer a path to substantially more reliable intended behavior. Barto framed reinforcement learning around control, search, and associative memory, describing its core insight as caching search results rather than searching continuously.

    Image from @BAAIBeijing's post

Jun 13

Jun 13Sat

Jun 12

Jun 12Fri
  1. Jeremy HowardAI score72

    US export directive forces Anthropic to disable Fable 5 and Mythos 5 for customers

    AIThe US government issued an export control directive suspending access to Fable 5 and Mythos 5 for all foreign nationals, inside or outside the United States. Anthropic says the order forces it to disable both models for all customers, while other Claude models are unaffected. Anthropic calls the directive a misunderstanding and says it is working to restore access as soon as possible. The author disagrees with the decision and questions why Anthropic did not anticipate it, given its claim that only it can safely handle these models.

Jun 10

Jun 10Wed
  1. AI Snake OilAI score70

    Why AI hasn't replaced software engineers, and why it likely won't

    AIThe essay argues that AI compresses the execution layer of software work while decision-making and accountability remain human, so AI is not yet replacing software engineers. It cites AI-attributed layoffs at Block, Snap, and Intuit that the authors say were not driven by AI, and WARN Act filings in which only one company checked an AI box. A Federal Reserve analysis is cited as finding software engineer employment growing about 3 percentage points per year more slowly after ChatGPT than a no-AI counterfactual.

Jun 9

Jun 9Tue
  1. Andrej KarpathyAI score65

    Karpathy Calls Claude Fable 5 a Major Step Forward for Long Tasks

    AIAndrej Karpathy says Claude Fable 5 is the same underlying model as Mythos with added safeguards, and that it leads on nearly all benchmarks. He describes it as a step change, especially for long, difficult problem-solving sessions where it handles more ambitious tasks without close supervision. He notes that its safeguards are set a bit too aggressively at launch and may be tuned over time.

  2. One Useful Thing (Ethan Mollick)AI score72

    Ethan Mollick tests Claude 5 Fable and finds it runs long projects with little user input

    AIEthan Mollick, who had early access to Claude 5 Fable, reports that it outperformed other public models in his tests, including an isochrone travel-time map and a nine-and-a-half-hour software build called Concord. He says the model delegated work to other agents and made many design choices he could not see or weigh in on, leaving him closer to a client than a hands-on operator. He also notes high token usage, frequent fallback to Claude 4.8 Opus under security guardrails, and persistent quirks in its writing style.

Jun 5

Jun 5Fri

Jun 4

Jun 4Thu
  1. One Useful Thing (Ethan Mollick)AI score44

    Ethan Mollick Announces Co-Existence, a Sequel Book on Working Alongside AI

    AIEthan Mollick is releasing Co-Existence on October 20, a new book about working with AI systems that are sometimes, but not always, better than humans. The book follows his 2024 title Co-Intelligence, which he says was written about an era of chatbots rather than autonomous agents. Mollick also reports writing every chapter draft himself while using AI readers and fact-checkers, and building the book's website with Claude Code using Opus 4.8.

Jun 3

Jun 3Wed
  1. Mark ChenAI score25

    Mark Chen says OpenAI's models could match Mythos on cyber vulnerabilities

    AIOpenAI's Mark Chen said that after Mythos showed AI models can prove 80-year-old theorems, he expected them to also find cyber vulnerabilities, and they did. He added that researchers in math may now be thinking the same idea in reverse, applying cybersecurity-style capability to mathematics. The post offers no specific models, benchmarks, or figures.

Jun 2

Jun 2Tue

May 28

May 28Thu
  1. Sam BowmanAI score38

    Anthropic highlights AI for transparency in Claude Opus 4.8 system card

    AISam Bowman says he is excited about alignment assessments in the recent system card for Claude Opus 4.8, crediting @MaskedTorah. He argues AI systems have considerable underexplored potential for transparency and coordination. The quoted Claude announcement describes Opus 4.8 as improving on Opus 4.7 with sharper judgment and more honest self-assessment of progress.

    Image from @sleepinyourhat's post
  2. HyperdimensionalAI score34

    A Cascade of Conscientiousness: Foundation for American Innovation Launches Physical Intelligence Team

    AIThe Foundation for American Innovation launched a Physical Intelligence team to address regulatory and legal barriers to deploying autonomous robotics and other physical-world AI in the United States. The team plans to focus on regulatory climate, technical trajectory, industrial strategy, and liability and cybersecurity frameworks. The article argues that physical AI will matter most where human on-site labor drives costs, such as construction.