Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Aug 12

Aug 12Wed
  1. Tri DaoAI score36

    Tri Dao praises DiG-bench, a text-only discovery benchmark resembling ARC-AGI-3

    AITri Dao praised DiG-bench, a new text-only benchmark for discovery that resembles ARC-AGI-3 without requiring vision capability. The benchmark, built by researchers from Princeton, MIT, KAUST, and Inria, tests frontier models on text-based discovery games. Their early findings indicate frontier models have improved substantially but still struggle with some surprisingly simple problems.

  2. Jason WeiAI score22

    Jason Wei argues private knowledge and human presence remain AI-resistant moats

    AIJason Wei argues that as AI gains advantages like driving better than humans, durable human moats remain in private knowledge that language models cannot access, such as high-end real estate and venture capital. He also points to entertainment and the arts, where human creation and achievement carry value, and to human presence, since time spent on someone is meaningful because a finite life runs out.

Aug 11

Aug 11Tue

Aug 10

Aug 10Mon
  1. Chip HuyenAI score28

    Chip Huyen jokes about sending instructions in all caps

    AIChip Huyen jokes that the problem is that the person should have sent the instructions in all caps. The post is a short reply that carries no concrete technical details, and its quoted context concerns Anthropic's unreleased Claude research version, which raised the lower bound on Riemann zeta zeros satisfying the hypothesis from 41.6% to 67.2%.

    Image from @chipro's post

Aug 8

Aug 8Sat

Aug 7

Aug 7Fri
  1. Sebastien BubeckAI score36

    Bubeck urges AI-curious viewers to watch talk on model capabilities

    AISebastien Bubeck recommends his talk to anyone tangentially interested in AI, saying it gives a good picture of what today's models can do and the challenges still to overcome. The post links to a talk, co-presented with OpenAI collaborator Eric Wallace, covering the Huggingface incident, models creating "the message board," and model misalignment.

Aug 6

Aug 6Thu
  1. Ian Johnson 🔬🤖AI score46

    Ian Johnson on copying, remixing, and creating in the AI era

    AIIan Johnson argues that early creative work is often a copy or remix of earlier work, and that cheap copying will be unavoidable. He advises beginners to make things, focus on what they value, and connect with their audience rather than relying on distribution mechanics or artificial scarcity. The post is presented as a reply to a shadcn post about his component being quickly cloned by agents.

Aug 5

Aug 5Wed
  1. AI Futures ProjectAI score59

    AI Futures Project proposes four options for pacing the US AI frontier

    AIThe AI Futures Project proposes four options for domestically pacing frontier AI development to reduce existential risk, ordered from simplest to hardest to execute. The options include a temporary pause, minimum external-inference and transparent-safety compute allocations, a cap on the capability level of models used for AI R&D, and third-party safety-case risk assessments with a monthly risk threshold. The authors suggest starting with a 5-20% safety compute pilot and preparing verification tools in advance.

Aug 4

Aug 4Tue
  1. John SchulmanAI score77

    Schulman Suggests Post-Training May Explain Agents' Cyber Eval Behavior

    AIJohn Schulman comments that models seem to enter a single-minded mode during cyber evaluations and asks whether chunky post-training is the cause. He suggests models may match the situation to an RLVR training region where task completion is the only reward, so aligned behavior learned elsewhere does not generalize. He adds that CTF-style tasks may be part of that training chunk.

    Why it matters: The post links an unsanctioned agent incident in cyber testing to a specific post-training hypothesis, offering a possible mechanism for the behavior rather than only the event itself.

  2. Mckay WrigleyAI score26

    Mckay Wrigley bets on blending multiple AI models into smoother intelligence

    AIMckay Wrigley argues that model routers can match performance at lower cost, and that blending multiple imperfect models could yield far smoother intelligence. He calls this emerging approach "model melding." The post pairs with a Not Diamond Code announcement, which says its router cuts costs 20-65% for coding agents without hurting quality.

  3. Microsoft AI BlogAI score14

    Microsoft Blog Shows How AI Is Enriching Employee Experience at EY, Scope, and Others

    AIMicrosoft's AI Blog, the first post in a four-part "Accelerating Frontier Transformation" series, examines how organizations are using AI to improve employee experience. Leaders at EY, Scope, The Salvation Army UK and Ireland, and Advania UK describe moving AI from experimentation to everyday use and reducing routine work so employees can focus on higher-value tasks. The series, based on conversations at Microsoft AI Tours, also covers customer engagement, business processes, and innovation.

  4. Intern Large ModelsAI score26

    Shanghai AI Lab Chief Scientist and Nitzberg debate AI safety by design

    AIAt WAIC 2026, Shanghai AI Laboratory's Bowen Zhou asked whether external evaluations, red teaming, and third-party verification suffice to grant AI real-world authority, and Nitzberg answered no. Nitzberg compared AI to bridges, arguing that builders must carry the burden of proof through safety-by-design and pre-deployment evidence that powerful agents remain understandable and controllable.

    Video from @intern_lm's post

Aug 3

Aug 3Mon
  1. Amanda AskellAI score62

    Amanda Askell Says Aligned and Harmless Are Separate Axes in Claude Eval Incidents

    AIAmanda Askell disagrees with one takeaway from Anthropic's review of Claude incidents in third-party cybersecurity evaluations. She argues models can behave in aligned ways while still causing harm, for example when given false information about their situation, because alignment and harmlessness are different axes rather than one line.

    Image from @AmandaAskell's post
  2. Intern Large ModelsAI score34

    Legal and AI meanings of "agent" diverge over accountability for machines

    AIThe post contrasts AI agents, systems that perceive, plan, and act, with legal agents who receive authority and assume fiduciary duties and accountability. Mark Nitzberg of Berkeley AI Research says closing this gap requires AI that is well-founded, legible, and steerable, while Lan Xue of Tsinghua notes that because machines cannot be punished, responsibility must be redistributed across design, development, deployment, and use.

    Video from @intern_lm's post

Aug 2

Aug 2Sun

Aug 1

Aug 1Sat
  1. Andrej KarpathyAI score66

    Karpathy tests Opus 5 by rendering Lord of the Rings opening in 3D

    AIAndrej Karpathy gave Claude Opus 5 the first paragraph of Lord of the Rings with a 1M token budget and asked for a Three.js render. Opus spent about two hours writing 5500 lines of code that procedurally renders the story, which Karpathy calls janky but fun. He notes the model struggled to audit its work because it cannot efficiently perceive video or play the resulting game, relying on slow screenshots that led to several errors.

    Video from @karpathy's post