Skip to contentSkip to stories

Updated

#Trend

Showing low-relevance items too. Hide low-relevance items

Mar 1

Mar 1Sun
  1. Artificial IgnoranceAI score46

    Build Your Own Benchmark: Why Public AI Evals Are Saturating and What Replaces Them

    AIPublic AI benchmarks such as MMLU, SWE-bench Verified, and GPQA Diamond are saturating or showing contamination, prompting OpenAI to call SWE-bench Verified "no longer suitable" in late February and recommend SWE-bench Pro. OpenAI's audit found 59.4% of the problems its best model failed had flawed test cases, and GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash could reproduce original fixes from memory. The article argues that behavioral tests, such as Vending-Bench's simulated vending machine business, may be more useful for everyday model choice.

Feb 27

Feb 27Fri

Feb 25

Feb 25Wed
  1. Jim FanAI score75

    EgoScale trains a 22-DoF humanoid mostly on 20,000 hours of human video

    AIResearchers trained a humanoid with 22-DoF dexterous hands mainly on over 20,000 hours of egocentric human video, with no robot in the loop, to perform tasks such as assembling model cars and folding shirts. They report a log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and state that this loss predicts real-robot success rate. The recipe, called EgoScale, pre-trains GR00T N1.5 on the video, adds only 4 hours of robot play data, and reports a 54% gain over training from scratch across five dexterous tasks.

    Video from @DrJimFan's post

Feb 19

Feb 19Thu

Feb 12

Feb 12Thu
  1. AI Futures ProjectAI score65

    AI Futures Project grades its 2025 AI 2027 predictions against reality

    AIAI Futures Project grades its AI 2027 scenario for 2025 and finds quantitative progress running at roughly 65% of the predicted pace, later revised to about 75%. Most qualitative predictions, such as the rise of coding agents, are judged on pace, while SWE-bench-Verified progress was slower than forecast and OpenAI's valuation trailed the scenario. The authors say their timelines lengthened over 2025 and plan to keep updating forecasts through 2026.

Feb 10

Feb 10Tue

Feb 9

Feb 9Mon

Feb 5

Feb 5Thu
  1. Geoffrey HintonAI score26

    Hinton praises International AI Safety Report 2026 as essential reading on AI risks

    AIGeoffrey Hinton called the International AI Safety Report 2026 a thoughtful, detailed, and well-researched description of AI risks, essential reading for anyone writing or speaking about them. Yoshua Bengio's thread introduces the report as the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date.

Feb 1

Feb 1Sun
  1. Yi TayAI score35

    Yi Tay on hiring for frontier AI teams, seniority, and publication norms

    AIYi Tay, who recently hired a full AI team from thousands of applications, says PhDs are still a reasonable training ground and that candidates often get noticed through strong work they publish. He argues that seniority matters little in today's LLM world, and that being at the cutting edge outweighs external social media visibility. He also disagrees with the view that being a middle author on many papers is a negative signal.

Jan 29

Jan 29Thu
  1. Chip HuyenAI score18

    Chip Huyen launches GoodAIList.com to track trending open-source AI repos

    AIChip Huyen built GoodAIList.com, which tracks 14K open-source AI repositories with contributions from over 145K developers. Each day it searches for new repos using 123 keywords and topics, surfaces those gaining traction, and categorizes them with AI-generated annotations that she notes are not highly accurate. The site also maps contributor locations, which she uses to find people doing interesting AI work when she travels.

    Image from @chipro's post

Jan 26

Jan 26Mon

Jan 25

Jan 25Sun

Jan 23

Jan 23Fri

Jan 15

Jan 15Thu

Jan 14

Jan 14Wed
  1. Chip HuyenAI score14

    Agentic Hackathon projects tackle long-running tasks, retrieval, and multimodal agents

    AIChip Huyen praised projects at last weekend's Agentic Hackathon, which hosted by MongoDB and Cerebral Valley, where she served as a judge. Teams tackled long-running tasks such as memory management, recovery from mid-task failures, and consistency across steps and sub-agents, along with adaptive retrieval across databases, search indices, and websites. Finalist demos are scheduled in San Francisco tomorrow, with talks by Douglas Eck.

    Image from @chipro's post

Jan 13

Jan 13Tue

Dec 19, 2025

Dec 19, 2025Fri
  1. Andrej KarpathyAI score75

    Karpathy's 2025 LLM review names RLVR and jagged intelligence as key shifts

    AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.

Dec 11, 2025

Dec 11, 2025Thu

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

  2. Yi TayAI score38

    Google DeepMind's Gemini team launches new reasoning research group in Singapore

    AIYi Tay announced that Google DeepMind's Gemini team is starting a new research team in Singapore focused on advanced reasoning, LLM/RL, and improving frontier models such as Gemini and Gemini Deep Think. The team is led by Tay and reports to Quoc Le's broader team in Mountain View, which recently contributed to IMO and ICPC gold medal results with Gemini Deep Think. The team is starting small and is recruiting exceptionally capable engineers and researchers from the region and beyond.

    Image from @YiTayML's post

Nov 28, 2025

Nov 28, 2025Fri

Nov 26, 2025

Nov 26, 2025Wed

Nov 17, 2025

Nov 17, 2025Mon
  1. Andrej KarpathyAI score60

    Karpathy argues verifiability predicts which tasks AI automates fastest

    AIKarpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.

Nov 14, 2025

Nov 14, 2025Fri

Oct 30, 2025

Oct 30, 2025Thu
  1. Chip HuyenAI score27

    Chip Huyen's AI product lessons: UX, data, and team structure matter most

    AIChip Huyen argues that many AI product failures stem from user experience, data quality, and organizational structure rather than the AI itself. She cites a chatbot whose traction improved after adding pre-populated questions and a voice option for users whose hands were busy, and a lead scoring model that was broken because marketing wasn't asking the right questions. She also notes that senior engineers gain the most from AI coding while resisting it more, and recommends building small tools for daily frustrations to solve the "idea crisis."

Oct 22, 2025

Oct 22, 2025Wed

Oct 14, 2025

Oct 14, 2025Tue

Jul 6, 2021

Jul 6, 2021Tue