Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. François CholletAI score62

    Chollet Argues Reasoning Models Differ from Base LLMs by Inductive Program Prediction

    AIFrançois Chollet argues the key difference between base LLMs and modern LRMs is a shift from transductive answer prediction to inductive prediction of the program or reasoning chain behind an answer. He says this enables test-time induction and substantial fluid intelligence in LRMs, which he claims base LLMs largely lack. He cites ARC 1 results: base LLMs remain around 10-15%, while LRMs of the same size or smaller saturated the benchmark in 2025.

  2. Guillermo RauchAI score38

    Guillermo Rauch says verification engineering is the future of software

    AIGuillermo Rauch argues that the future is verification engineering, spanning proofs, end-to-end tests, benchmarks, and linters. He expects some of these tests to be deterministic and others agentic, and he says the approach looks great. The quoted post introduces e2e, an open-source agentic testing framework that mixes deterministic and agentic APIs and runs locally or in CI.

  3. Google GemmaAI score54

    Google Gemma credits StudentBench study comparing AI and expert human GRE tutors

    AIGoogle Gemma relays a StudentBench study reporting that AI tutors matched expert human tutors on immediate GRE learning gains. The author reports 2,383 students and a cost of 7 cents per AI tutor hour versus $75 for an expert human hour. The post also says the top AI tutor beat expert human tutors on average in 5 of 7 academic topics, and that the data and paper are publicly available.

  4. Prime IntellectAI score34

    Qwen3.6 reward rises 2.8x via GRPO on Hosted Training

    AIPrime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families. Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.

    Image from @PrimeIntellect's post
  5. Yellowbrick InvestingAI score18

    Yellowbrick 2.0 launches with leaderboards, API access, and custom feeds

    AIYellowbrick 2.0 is live, tracking 35,000+ stock pitches from 4,000+ authors and adding 300+ new pitches weekly. The rebuilt platform adds author leaderboards, custom feeds and alerts, API access, and paid research partner discounts. Premium subscribers get a 30% discount on Koyfin, which the company says covers the cost of Yellowbrick Premium.

  6. Mustafa SuleymanAI score40

    Microsoft AI launches MAI-Transcribe-2-Streaming, claiming top real-time transcription accuracy

    AIMicrosoft AI launched MAI-Transcribe-2-Streaming, which Artificial Analysis ranks #1 of 38 models for final transcript accuracy at 2.5% WER, returned 0.13s after end of speech. Artificial Analysis lists its streaming price at $0.54 per hour of audio, at the higher end among leading streaming models. Microsoft's post claims the model is 55% faster and 60% cheaper than ElevenLabs and invites developers to build agents on its platform.

  7. Jerry LiuAI score42

    LlamaIndex launches Extract v2.5 document extraction agents with improved accuracy

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction across cost-effective, agentic, and agentic plus tiers. The company reports the agents outperform Opus 5.5 and GPT-6 Sol while costing 30% to 4x less, with accuracy gains on long lists (86.1% to 95.5%), multi-page records (85.5% to 96.5%), and scanned forms (90.9% to 95.7%) on its agentic tier. The release adds advanced citations with bounding boxes and structural reasoning, and the agents are available on LlamaParse.

    Video from @jerryjliu0's post
  8. Lewis Tunstall @ COLM 🌉AI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  9. Cloudflare Blog · AIAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.

  10. LangChain BlogAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.

Sep 30

Sep 30Wed
  1. indigoAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  2. Apple Machine Learning ResearchAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  3. whAI score67

    Gemini 4 Argon previewed with frontier coding and cyber defense claims

    AIThe post quotes Google's Sundar Pichai introducing Gemini 4 Argon as an early look at the next model. It claims frontier performance in complex workflows, cyber defense, and software engineering, and says Google teams are using it for tasks from coding to quantum computing. The author adds that on FrontierSWE the model is very self-critical and often says "Eureka!", a personality they describe as a large improvement over previous Gemini models.

    Image from @nrehiew_'s post
  4. Google DeepMindAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  5. Google · Gemini appAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.