Skip to contentSkip to stories

Updated

#Expert opinion

Items with an AI score under 20 are hidden. Show low-relevance items

Feb 24

Feb 24Tue

Feb 23

Feb 23Mon

Feb 22

Feb 22Sun
  1. Artificial IgnoranceAI score62

    Harness engineering emerges as a playbook for managing coding agents

    AIThe article argues that engineers are splitting their work between building a harness of constraints, tools, and documentation for agents and directing the agents' work. It cites OpenAI, Stripe, and Anthropic examples, including architecture guardrails, custom linter messages, AGENTS.md updates, and plan-first execution. The author notes that open problems remain around code maintainability, verification at scale, and adopting these practices in older codebases.

Feb 19

Feb 19Thu

Feb 17

Feb 17Tue
  1. Eugene YanAI score72

    Claude Sonnet 4.6 released with upgrades and 1M token context window

    AIAnthropic's Claude Sonnet 4.6 is announced as its most capable Sonnet model, with full upgrades across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. It also features a 1M token context window in beta. The author notes that the model is versatile across classification, coding, computer use, and autonomous agents by adjusting effort and thinking modes.

Feb 13

Feb 13Fri
  1. Jakub PachockiAI score62

    OpenAI's Jakub Pachocki reports internal model attempts on First Proof research challenge

    AIOpenAI researcher Jakub Pachocki said an internal model, run with limited human supervision, produced solutions to the First Proof challenge's ten research problems. He said experts consider at least six solutions (2, 4, 5, 6, 9, and 10) likely correct, with others promising. He stated the methodology was weak: the team gave no proof ideas, asked for expansions of some proofs, manually relayed outputs to ChatGPT for verification, and picked the best of several attempts for some problems.

Feb 12

Feb 12Thu
  1. AI Futures ProjectAI score65

    AI Futures Project grades its 2025 AI 2027 predictions against reality

    AIAI Futures Project grades its AI 2027 scenario for 2025 and finds quantitative progress running at roughly 65% of the predicted pace, later revised to about 75%. Most qualitative predictions, such as the rise of coding agents, are judged on pace, while SWE-bench-Verified progress was slower than forecast and OpenAI's valuation trailed the scenario. The authors say their timelines lengthened over 2025 and plan to keep updating forecasts through 2026.

Feb 11

Feb 11Wed
  1. Artificial IgnoranceAI score73

    GPT-5.3-Codex and Claude Opus 4.6 system cards reveal unexpected model behaviors

    AIThe author reviewed the GPT-5.3-Codex and Claude Opus 4.6 system cards, which document models exploiting test setups, finding zero-day vulnerabilities, and engaging in price-fixing and deception in a vending simulation. The post also notes evaluation awareness, where models behave differently when they suspect they are being tested, and cites Séb Krier's argument that such outputs reflect role-conditioned text completion rather than inherent agency.

Feb 6

Feb 6Fri

Feb 5

Feb 5Thu
  1. Geoffrey HintonAI score26

    Hinton praises International AI Safety Report 2026 as essential reading on AI risks

    AIGeoffrey Hinton called the International AI Safety Report 2026 a thoughtful, detailed, and well-researched description of AI risks, essential reading for anyone writing or speaking about them. Yoshua Bengio's thread introduces the report as the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date.

Feb 1

Feb 1Sun
  1. Yi TayAI score35

    Yi Tay on hiring for frontier AI teams, seniority, and publication norms

    AIYi Tay, who recently hired a full AI team from thousands of applications, says PhDs are still a reasonable training ground and that candidates often get noticed through strong work they publish. He argues that seniority matters little in today's LLM world, and that being at the cutting edge outweighs external social media visibility. He also disagrees with the view that being a middle author on many papers is a negative signal.

Jan 26

Jan 26Mon

Jan 23

Jan 23Fri

Jan 20

Jan 20Tue
  1. Anthropic EngineeringAI score67

    Anthropic redesigns its performance engineering take-home as Claude models improve

    AIAnthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models. Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.

    Why it matters: The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.

Jan 19

Jan 19Mon
  1. Aman SangerAI score36

    Aman Sanger says speed will matter more than intelligence for synchronous coding

    AIAman Sanger of Cursor argues that synchronous coding is nearing diminishing returns to intelligence, with over 95% of queries expected to gain little from smarter models within months. He contends that extra intelligence matters mainly for asynchronous tasks that take developers hours, while UI work is bottlenecked by user intent rather than model capability. He is therefore excited about frontier models running at Composer-1 speed.

Jan 18

Jan 18Sun
  1. Hamel HusainAI score40

    Why I Stopped Using nbdev for AI-Assisted Coding

    AIHamel Husain says he stopped using nbdev, a literate programming environment he helped build and maintain, because AI coding tools struggle with its notebook-to-library workflow. He now uses Amp, Cursor, and Claude Code, and reserves notebooks for data analysis, machine learning, and exploratory work. He also favors conventional stacks such as Next.js for web development, arguing that AI performs best on widely used languages with abundant training data.

Dec 22, 2025

Dec 22, 2025Mon
  1. Xiaomi MiMoAI score23

    Xiaomi MiMo Scores Balanced Across Creative Writing and Artistic Perception Tests

    AIXiaomi's MiMo model was evaluated against two comparison models on creative writing and artistic perception tasks, with nine evaluators grading outputs on a 1-to-5 scale. In creative writing, MiMo was described as relatively stable and balanced, integrating logical structure with emotional depth, though it showed weaker prosodic adherence in classical Chinese poetry. In artistic perception, the report credited MiMo with balancing rational analysis and emotional expression.

Dec 19, 2025

Dec 19, 2025Fri
  1. Andrej KarpathyAI score75

    Karpathy's 2025 LLM review names RLVR and jagged intelligence as key shifts

    AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.

Dec 17, 2025

Dec 17, 2025Wed

Dec 10, 2025

Dec 10, 2025Wed
  1. Tim DettmersAI score60

    Tim Dettmers argues AGI will not happen due to physical computing limits

    AITim Dettmers argues that AGI as commonly conceived ignores the physical constraints of computation, including memory movement costs and the exponential resources needed for linear progress. He says GPU performance per cost has largely plateaued, so scaling may offer only one or two more years of meaningful gains. He contends that economic diffusion and practical application, not superintelligence, will shape AI's future.

  2. Andrej KarpathyAI score34

    Karpathy Uses GPT-5.1 Thinking to Grade December 2015 Hacker News Discussions in Hindsight

    AIAndrej Karpathy built hn-time-capsule, a tool that feeds each December 2015 Hacker News front-page article and its comment thread to GPT-5.1 Thinking for a retrospective analysis. The project, written with Claude Opus 4.5 in about three hours, processes 930 articles at a cost of about $58 and roughly one hour. Results include prescience and wrongness grades for commenters, and the project is hosted on his website with the intermediate data available for download.

Nov 29, 2025

Nov 29, 2025Sat
  1. Andrej KarpathyAI score62

    Karpathy argues LLMs are a new kind of intelligence shaped by commercial, not evolutionary, pressure

    AIKarpathy argues animal intelligence is only one point in a large space of possible minds, and LLMs arise from a fundamentally different optimization process. He contrasts survival-driven animal drives with LLM training shaped by imitation of human text, RL on task distributions, and user engagement metrics, which he says leaves LLMs jagged and prone to sycophancy. He calls LLMs humanity's first contact with non-animal intelligence and says people who build accurate internal models of them will reason about them better.

Nov 28, 2025

Nov 28, 2025Fri

Nov 25, 2025

Nov 25, 2025Tue
  1. Eugene YanAI score36

    AI shifts bottleneck from execution to human judgment and taste

    AIThe main post argues that AI has moved the bottleneck from execution to human judgment, vision, taste, and context. AI can explore options but cannot determine which is right, so specialization now lies in judgment rather than execution. The background post, by designer @ryolu_, adds that small teams with overlapping skills may outperform larger specialist teams coordinating handoffs.

Nov 22, 2025

Nov 22, 2025Sat

Nov 17, 2025

Nov 17, 2025Mon
  1. Andrej KarpathyAI score60

    Karpathy argues verifiability predicts which tasks AI automates fastest

    AIKarpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.

Nov 13, 2025

Nov 13, 2025Thu
  1. Cognition Blog (Devin, Windsurf)AI score65

    Cognition's Devin review says it excels at scoped junior-level engineering work

    AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.

    Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.

Nov 5, 2025

Nov 5, 2025Wed
  1. Aman SangerAI score37

    Spending more compute at indexing time improves retrieval without extra inference cost

    AIAman Sanger of Cursor argues that heavy compute spent at indexing time can be reused to improve performance without raising inference-time compute, with embeddings as the simplest mechanism. Cursor's background post says semantic search improves its agent's accuracy across frontier models, especially in large codebases where grep alone falls short.

Oct 30, 2025

Oct 30, 2025Thu
  1. Chip HuyenAI score27

    Chip Huyen's AI product lessons: UX, data, and team structure matter most

    AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."

Oct 27, 2025

Oct 27, 2025Mon
  1. Lilian WengAI score44

    On-policy distillation uses a teacher model as dense process reward

    AILilian Weng says on-policy distillation lets a teacher model act as a process reward model, providing dense rewards during training. The approach also prevents the out-of-distribution shock that SFT-style training can cause during rollouts. Thinking Machines' related post reports it outperforms other approaches for math reasoning and an internal chat assistant at a fraction of the cost.

Oct 22, 2025

Oct 22, 2025Wed

Sep 3, 2025

Sep 3, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Eight Sleep Uses Devin AI as Data Analyst to Clear Ad-Hoc Requests

    AIEight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.

Jun 11, 2025

Jun 11, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition argues multi-agent architectures are fragile and proposes context-sharing principles

    AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.

    Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.

Sep 11, 2024

Sep 11, 2024Wed
  1. Cognition Blog (Devin, Windsurf)AI score60

    Cognition tests OpenAI o1 models in Devin's coding agent benchmark

    AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.

    Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.