Skip to contentSkip to stories

Updated

#Google

Showing low-relevance items too. Hide low-relevance items

Dec 17, 2025

Dec 17, 2025Wed

Dec 12, 2025

Dec 12, 2025Fri

Dec 11, 2025

Dec 11, 2025Thu

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

  2. Yi TayAI score38

    Google DeepMind's Gemini team launches new reasoning research group in Singapore

    AIYi Tay announced that Google DeepMind's Gemini team is starting a new research team in Singapore focused on advanced reasoning, LLM/RL, and improving frontier models such as Gemini and Gemini Deep Think. The team is led by Tay and reports to Quoc Le's broader team in Mountain View, which recently contributed to IMO and ICPC gold medal results with Gemini Deep Think. The team is starting small and is recruiting exceptionally capable engineers and researchers from the region and beyond.

    Image from @YiTayML's post

Nov 25, 2025

Nov 25, 2025Tue

Nov 20, 2025

Nov 20, 2025Thu

Nov 18, 2025

Nov 18, 2025Tue

Nov 13, 2025

Nov 13, 2025Thu

Nov 3, 2025

Nov 3, 2025Mon
  1. ARC PrizeAI score38

    ARC Prize Launches Verified Program to Certify ARC-AGI Benchmark Scores

    AIARC Prize Foundation announced ARC Prize Verified, a program that certifies frontier model scores on the ARC-AGI benchmark using hidden test sets and adds a third-party academic panel to audit and open-source its testing process. Five AI labs, including Google and xAI, are sponsoring ARC-AGI-3 development, and the foundation says donations do not influence verification scoring. Models that pass verification will appear on the official leaderboard with a verification badge.