Skip to content

Formats · Latest news

Trends

Industry-wide changes in usage, capabilities, social impact, and market structure.

25 top picks all-time · 15 in the past 30 days · chosen from 1,098 items collected all-time

Latest pick

Top picks archive · Page 2

Top picks 21–25 of 25

May 26

May 26Tue
  1. MiniMax BlogOfficialAI score67

    MiniMax Agent Team Adds Parallel Multi-Agent Collaboration for Long Tasks

    AIMiniMax has upgraded its Agent, renamed Mavis, and introduced Agent Teams that run multiple role-based Agents in parallel on desktop. The team uses Leader, Worker, and Verifier roles so complex tasks can be split, checked, and reported at key checkpoints, and it merges TokenPlan and Agent Plan into one subscription with credits shared between Agent and API. The post also discusses the added token, handoff, and retry costs of multi-Agent work, and says the Agent will be open-sourced alongside MiniMax M3.

    Why it matters: The post explains why multi-Agent helps long tasks and where its verification, token, and aggregation costs come from, useful for judging when a team setup beats a single Agent.

May 6

May 6Wed
  1. Nick TurleyXAI score62

    OpenAI rolls out GPT-5.5 Instant to ChatGPT with better factuality

    AIOpenAI has shipped GPT-5.5 Instant to ChatGPT, rolling out to everyone over the next couple of days. The quoted post says the model focuses on factuality, reducing hacks, and improving baseline intelligence, and is significantly less likely to hallucinate.

    Why it matters: The quoted post gives concrete targets for the update, factuality and hallucination reduction, useful for judging whether the default ChatGPT model changed in practice.

Feb 25

Feb 25Wed
  1. Jim FanXAI score75

    EgoScale trains a 22-DoF humanoid mostly on 20,000 hours of human video

    AIResearchers trained a humanoid with 22-DoF dexterous hands mainly on over 20,000 hours of egocentric human video, with no robot in the loop, to perform tasks such as assembling model cars and folding shirts. They report a log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and state that this loss predicts real-robot success rate. The recipe, called EgoScale, pre-trains GR00T N1.5 on the video, adds only 4 hours of robot play data, and reports a 54% gain over training from scratch across five dexterous tasks.

    Why it matters: The source reports a measured scaling law linking human video volume to robot action loss, which bears on how dexterous robot training data might be collected and reused.

    Video from @DrJimFan's post

Dec 19, 2025

Dec 19, 2025Fri
  1. Andrej KarpathyBlogAI score75

    Karpathy's 2025 LLM review names RLVR and jagged intelligence as key shifts

    AIAndrej Karpathy's year-in-review lists the LLM paradigm changes he found most notable in 2025. He highlights Reinforcement Learning from Verifiable Rewards (RLVR), which drove most capability gains as labs ran longer RL training, and describes LLM intelligence as jagged, strong in verifiable domains and weak elsewhere. He also covers Cursor-style LLM apps, Claude Code running on the user's computer, vibe coding, and the case for a visual LLM GUI.

    Why it matters: Karpathy ties the year's shifts to RLVR, jagged capability, and local agents, giving readers a framework for judging how LLM progress is changing.

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeOfficialAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.