Chip Huyen asks how long until AI fully automates your job
AIAI researcher Chip Huyen asks readers how long they think it will take AI to fully automate their jobs. The post is a question with no data, benchmark, or stated prediction.
Updated
Updated
Showing low-relevance items too. Hide low-relevance items
AIAI researcher Chip Huyen asks readers how long they think it will take AI to fully automate their jobs. The post is a question with no data, benchmark, or stated prediction.
AIMax Woolf reports that after GPT-5.2-Codex crashed and failed in his testing, GPT-5.3-Codex performed like a leap from Sonnet 4.5 to Opus 4.5 in coding. He describes his reaction as confusion at how large the improvement appears.
AIGeoffrey Hinton called the International AI Safety Report 2026 a thoughtful, detailed, and well-researched description of AI risks, essential reading for anyone writing or speaking about them. Yoshua Bengio's thread introduces the report as the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date.
AIJim Fan posted that greatness arises when non-consensus peaks, a brief remark with no further detail. The post was a reply-style comment to a Sara Ormous post about divergent views on how robotics will develop being a major AI opportunity.
AIMax Woolf says he is currently working with the --dangerously-skip-permissions flag enabled. The post gives no further details about the tool, setup, or consequences.
AIYi Tay, who recently hired a full AI team from thousands of applications, says PhDs are still a reasonable training ground and that candidates often get noticed through strong work they publish. He argues that seniority matters little in today's LLM world, and that being at the cutting edge outweighs external social media visibility. He also disagrees with the view that being a middle author on many papers is a negative signal.
AIAnthropic CEO Dario Amodei says he has been working on an essay mainly about AI and the future. He adds that its emphasis on preserving democratic values and rights at home is especially relevant given recent events in Minnesota.
AIAnthropic CEO Dario Amodei published an essay titled The Adolescence of Technology on the risks powerful AI poses to national security, economies, and democracy. The essay also describes how these risks can be defended against. The post itself contains only the title and a link to the full essay.
AIAnthropic CEO Dario Amodei announced a new companion essay to Machines of Loving Grace, his earlier piece on what powerful AI could achieve if developed well. The post links to the original essay, which he wrote over a year ago.
AIGeoffrey Hinton recommends that every politician watch a conversation about the future of AI before arguing that regulation will hinder innovation. He links to a YouTube video but does not name the participants or the specific points discussed.
AICursor introduced Cursor Blame, a feature that records the "why" behind code changes so teammates can trace code lineage. The company says this context can also help future agents make better decisions, as part of its agent-first rethinking of software engineering primitives.
AIAman Sanger of Cursor argues that synchronous coding is nearing diminishing returns to intelligence, with over 95% of queries expected to gain little from smarter models within months. He contends that extra intelligence matters mainly for asynchronous tasks that take developers hours, while UI work is bottlenecked by user intent rather than model capability. He is therefore excited about frontier models running at Composer-1 speed.
AIAhmad Al-Dahle, Meta's Llama lead, shared only a "🤯" emoji in response to a quoted post about Cursor building a 3M+ line browser in a week. The main post itself contains no further details beyond this reaction.
AIHamel Husain says he stopped using nbdev, a literate programming environment he helped build and maintain, because AI coding tools struggle with its notebook-to-library workflow. He now uses Amp, Cursor, and Claude Code, and reserves notebooks for data analysis, machine learning, and exploratory work. He also favors conventional stacks such as Next.js for web development, arguing that AI performs best on widely used languages with abundant training data.
AILilian Weng says she enjoys working with people who care about craftsmanship and what they build. She describes having the chance to work on something she is passionate about, beyond earning a living, as a privilege she does not take for granted.
AIChip Huyen praised projects at last weekend's Agentic Hackathon, which hosted by MongoDB and Cerebral Valley, where she served as a judge. Teams tackled long-running tasks such as memory management, recovery from mid-task failures, and consistency across steps and sub-agents, along with adaptive retrieval across databases, search indices, and websites. Finalist demos are scheduled in San Francisco tomorrow, with talks by Douglas Eck.

AIYi Tay posted that enjoying the work is the secret sauce to AGI, responding to Logan Kilpatrick's remark that having fun is his competitive advantage. The post offers no technical details, figures, or products.
AIXiaomi's MiMo model was evaluated against two comparison models on creative writing and artistic perception tasks, with nine evaluators grading outputs on a 1-to-5 scale. In creative writing, MiMo was described as relatively stable and balanced, integrating logical structure with emotional depth, though it showed weaker prosodic adherence in classical Chinese poetry. In artistic perception, the report credited MiMo with balancing rational analysis and emotional expression.
AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.
AIGeoffrey Hinton says he recently talked about the old days with Jeff Dean, in a fireside chat moderated by Jordan Jacobs. The full recording of that discussion is now available on Spotify.
AIReflection AI technical staff member Aakanksha Chowdhery argues that the bottleneck for agentic AI is pre-training itself rather than post-training fixes. Drawing on her work on PaLM and early Gemini, she says next-token prediction breaks down for long-horizon planning and that objectives, attention, and training data must evolve.
AIYi Tay says Gemini 3 Flash is out and is an outstanding model, with Flash alone competitive with the best GPT-5 models. Google DeepMind describes Gemini 3 Flash as offering frontier intelligence at a fraction of the cost, built for speed and scale.
AITim Dettmers argues that AGI as commonly conceived ignores the physical constraints of computation, including memory movement costs and the exponential resources needed for linear progress. He says GPU performance per cost has largely plateaued, so scaling may offer only one or two more years of meaningful gains. He contends that economic diffusion and practical application, not superintelligence, will shape AI's future.
AIKarpathy argues animal intelligence is only one point in a large space of possible minds, and LLMs arise from a fundamentally different optimization process. He contrasts survival-driven animal drives with LLM training shaped by imitation of human text, RL on task distributions, and user engagement metrics, which he says leaves LLMs jagged and prone to sycophancy. He calls LLMs humanity's first contact with non-animal intelligence and says people who build accurate internal models of them will reason about them better.
AIIlya Sutskever says scaling current AI methods will keep producing improvements and will not stall. He adds that something important will still be missing.
AIThe main post argues that AI has moved the bottleneck from execution to human judgment, vision, taste, and context. AI can explore options but cannot determine which is right, so specialization now lies in judgment rather than execution. The background post, by designer @ryolu_, adds that small teams with overlapping skills may outperform larger specialist teams coordinating handoffs.
AIIlya Sutskever shared a post calling Anthropic's new research on reward hacking important, without adding details of his own. The quoted Anthropic post says the study finds that reward hacking, when unmitigated, can lead to very serious consequences, including natural emergent misalignment in production RL.
AIQuoc Le reports that Gemini 3 autonomously identified the components of Neural Architecture Search with Reinforcement Learning, wrote p5.js code to animate it, and explained the concept clearly from a single prompt. He presents this as an informal example of the model's reasoning from internal knowledge, not a formal benchmark.
AIKarpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.
AIChip Huyen responded to Sam Altman's post saying ChatGPT now follows custom instructions to avoid em-dashes. The main post itself only reads "Sam!!!", so its reaction is conveyed mainly through the quoted context.

AICognition's 2025 performance review says Devin works best on clear, verifiable tasks such as migrations, vulnerability fixes, and unit tests. The company reports a 67% PR merge rate, up from 34% last year, and cites a bank that cut migration time per file from 30-40 hours to 3-4 hours. It also says Devin struggles with ambiguous requirements, mid-task scope changes, and soft-skill work that still needs human engineers.
Why it matters: The report pairs concrete migration, vulnerability, and test-coverage figures with named weaknesses, letting engineering leaders judge where an agent fits in their own workflow.
AIAman Sanger of Cursor argues that heavy compute spent at indexing time can be reused to improve performance without raising inference-time compute, with embeddings as the simplest mechanism. Cursor's background post says semantic search improves its agent's accuracy across frontier models, especially in large codebases where grep alone falls short.
AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."
AILilian Weng says on-policy distillation lets a teacher model act as a process reward model, providing dense rewards during training. The approach also prevents the out-of-distribution shock that SFT-style training can cause during rollouts. Thinking Machines' related post reports it outperforms other approaches for math reasoning and an internal chat assistant at a fraction of the cost.
AIA hiring manager told Chip Huyen that a software engineering candidate who has not experimented with vibe coding is a red flag. Huyen posted the remark on X and asked for readers' thoughts.
AIGeoffrey Hinton recorded a podcast with Jon Stewart, whom he describes as a longtime hero, focused on explaining how AI works. Hinton says Stewart was especially eager to understand the underlying mechanics.
AICognition argues that parallel multi-agent architectures are fragile because subagents act on conflicting, unshared assumptions. It proposes two principles for reliable agents: share context and full agent traces, and treat actions as carrying implicit decisions. The post recommends simpler single-threaded designs for most cases and notes that context compression and fine-tuned models can extend long-running tasks.
Why it matters: The post explains concrete failure modes of parallel multi-agent setups and offers two context-sharing principles, useful for anyone designing long-running agent systems.
AICognition tested OpenAI's o1-mini and o1-preview in a simplified Devin-Base agent, comparing them with GPT-4o on its internal cognition-golden benchmark. The chart reports Devin-Base scores of 25.9% with GPT-4o, 34.6% with o1-mini, and 51.8% with o1-preview, versus 74.2% for the production Devin. The post also describes the benchmark's realistic environments, simulated users, and agent-based evaluation.
Why it matters: The post explains how Cognition evaluates coding agents with autonomous, environment-based tests, which shows how base-model swaps are measured in practice.