This is a great report that provides a thoughtful, detailed and very well researched description of the risks of AI.
AIIt is essential reading for anyone who wants to write or talk about AI risks.
Updated
Updated
AIIt is essential reading for anyone who wants to write or talk about AI risks.
AIYi Tay, who recently hired a full AI team from thousands of applications, says PhDs are still a reasonable training ground and that candidates often get noticed through strong work they publish. He argues that seniority matters little in today's LLM world, and that being at the cutting edge outweighs external social media visibility. He also disagrees with the view that being a middle author on many papers is a negative signal.
AIChip Huyen built GoodAIList.com, which tracks 14K open-source AI repositories with contributions from over 145K developers. Each day it searches for new repos using 123 keywords and topics, surfaces those gaining traction, and categorizes them with AI-generated annotations that she notes are not highly accurate. The site also maps contributor locations, which she uses to find people doing interesting AI work when she travels.
AIEvery politician should watch it before they join the lemmings saying that regulation of AI will interfere with innovation.
AIMaster: PhD:
AI…Embodied Intelligence, AI Agents, and AI for Science. 🚀🚀🚀 👉 Explore our open roles: 📧 Or email your CV to: Zstar@baai.ac.cn (Subject: “AAAI2026”) #AIResearch #AIJobs #AAAI2026
AIChip Huyen praised projects at last weekend's Agentic Hackathon, which hosted by MongoDB and Cerebral Valley, where she served as a judge. Teams tackled long-running tasks such as memory management, recovery from mid-task failures, and consistency across steps and sub-agents, along with adaptive retrieval across databases, search indices, and websites. Finalist demos are scheduled in San Francisco tomorrow, with talks by Douglas Eck.
AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.
AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.
Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.
AIYi Tay announced that Google DeepMind's Gemini team is starting a new research team in Singapore focused on advanced reasoning, LLM/RL, and improving frontier models such as Gemini and Gemini Deep Think. The team is led by Tay and reports to Quoc Le's broader team in Mountain View, which recently contributed to IMO and ICPC gold medal results with Gemini Deep Think. The team is starting small and is recruiting exceptionally capable engineers and researchers from the region and beyond.
AIIn particular, it won’t stall. - But something important will continue to be missing.
AIKarpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.
AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."
AIThe energy of a fast-growing company is magical. At ByteDance, I've had so much room to work with tech in creative ways and be on the frontier of modern music and technology."