Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 2

Oct 2Fri
  1. Baseten BlogAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  2. Design ArenaAI score40

    GPT-6 Astra hedges far more than Claude Opus 5.5 in reasoning summaries

    AIDesign Arena analyzed 324 thinking summaries and found OpenAI's GPT-6 Astra uses hedging words like "maybe," "might," and "it seems" about 20 times as often as Anthropic's Claude Opus 5.5. Opus usually weighs a few options and commits early, in about 4 out of 5 summaries versus 1 in 4 for Astra, which the post says works more like a designer while Opus works more like a builder.

    Video from @DesignArena's post
  3. Aravind SrinivasAI score62

    Perplexity open-sources models, an inference engine, and security tools

    AIPerplexity has released several open source projects, including the pplx-decider-v1-27b multimodal decision model, the pplx-embed-v2-context-9b-preview contextual embeddings model, and the Lily local inference engine for Apple silicon. The post also lists the 0.6B on-device PII-Tracer classifier with its PII-TRACE benchmark, the WANDR research agent benchmark, and the Numbat and Bumblebee security tools, and says more open source releases are coming soon.

  4. SGLangAI score39

    SGLang adds a scoring API and multi-item scoring for decision models

    AISGLang's update adds a /v1/score endpoint that returns scores for requested labels such as Yes/No or A/B/C, avoiding the label loss of generate with top-k logprobs. Its multi-item scoring computes shared context once and keeps each candidate isolated, with 16-candidate p95 on Qwen3-8B dropping from 54.1 ms (Generate) to 20.6 ms.

  5. Kilo (acq. by Anaconda)AI score36

    Ling 3.1 Flash is free in Kilo Code until October 13

    AIKilo Code is offering Ling 3.1 Flash for free until October 13, with the model served by Novita Labs. Ant Ling's background post describes the model as roughly 560B total parameters with about 25B active per token and a context window of up to 1M tokens. Ant Ling says it scores 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional, and plans to open-source it soon.

  6. Jerry LiuAI score44

    LlamaIndex Extract v2.5 hits 93–96% on dense table extraction benchmarks

    AILlamaIndex released Extract v2.5, a set of document extraction agents that it says reach 93%–96%+ accuracy on long-list extraction, including records spanning pages. The post claims the agents outperform frontier VLMs, which it says stop early, miss repeated records, and struggle to attribute values to sources, while LlamaIndex attributes every extracted value to its source. The agents are available through LlamaParse.

    Video from @jerryjliu0's post
  7. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  8. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  9. Lucas Beyer (bl16)AI score45

    Lucas Beyer praises new coding benchmark for finding bugs in repos

    AILucas Beyer calls SWE-sweep a useful new benchmark, where agents must find and fix bugs in a repo checked out at an earlier commit, scored against unit tests from real later bugfixes. He notes two limitations: a model may find valid bugs that don't match the tested ones, and the construction makes training on the test set easy. He advises not overemphasizing small ranking differences once models score highly.

  10. Hugging Face BlogAI score62

    AutoSynthData generates targeted training data for enterprise agents from failures

    AIServiceNow CoreAI introduced AutoSynthData, which uses a target model's failures and a stronger teacher's successes to generate and validate new agent training tasks. In EnterpriseOps Gym experiments, the Hybrid domain produced 2,000 samples and raised Gemma-4-26B-A4B-it mean Pass@1 by 7.2 percentage points, while the ITSM domain produced 1,994 samples and raised it from 18.77% to 27.18%.

    Why it matters: The post shows how failure analysis, teacher demonstrations, and verifier checks combine into a repeatable pipeline for generating targeted agent training data.

Oct 1

Oct 1Thu
  1. NVIDIA AIAI score44

    CoreWeave RL rollouts reload model weights 15× faster with Dynamo

    AICoreWeave's new RL rollouts service uses ModelExpress and Router in NVIDIA Dynamo to speed up model weight reloads during RL post-training with minimal downtime. Working with NVIDIA and You.com, CoreWeave achieved 15× faster model reloads than its baseline while post-training Nemotron 3.5 Lightning. The speedup addresses GPUs sitting idle while inference workers wait to load updated weights between training iterations.

  2. Apple Machine Learning ResearchAI score34

    Limits of Confidence-Based Sampling in Discrete Diffusion Models

    AIApple Machine Learning Research reports that discrete diffusion steps match the training distribution only when simultaneously written token positions are conditionally independent given already-fixed tokens. The authors show that per-position distributions cannot determine such dependence, and on the synthetic ScanAndAdd task, confidence-ranked groups of two or more positions were dependent and produced a generated distribution 29 times the sampling-noise floor in total variation.

  3. François CholletAI score62

    Chollet Argues Reasoning Models Differ from Base LLMs by Inductive Program Prediction

    AIFrançois Chollet argues the key difference between base LLMs and modern LRMs is a shift from transductive answer prediction to inductive prediction of the program or reasoning chain behind an answer. He says this enables test-time induction and substantial fluid intelligence in LRMs, which he claims base LLMs largely lack. He cites ARC 1 results: base LLMs remain around 10-15%, while LRMs of the same size or smaller saturated the benchmark in 2025.

  4. Guillermo RauchAI score38

    Guillermo Rauch says verification engineering is the future of software

    AIGuillermo Rauch argues that the future is verification engineering, spanning proofs, end-to-end tests, benchmarks, and linters. He expects some of these tests to be deterministic and others agentic, and he says the approach looks great. The quoted post introduces e2e, an open-source agentic testing framework that mixes deterministic and agentic APIs and runs locally or in CI.

  5. Google GemmaAI score54

    Google Gemma credits StudentBench study comparing AI and expert human GRE tutors

    AIGoogle Gemma relays a StudentBench study reporting that AI tutors matched expert human tutors on immediate GRE learning gains. The author reports 2,383 students and a cost of 7 cents per AI tutor hour versus $75 for an expert human hour. The post also says the top AI tutor beat expert human tutors on average in 5 of 7 academic topics, and that the data and paper are publicly available.

  6. Prime IntellectAI score34

    Qwen3.6 reward rises 2.8x via GRPO on Hosted Training

    AIPrime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families. Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.

    Image from @PrimeIntellect's post
  7. Mustafa SuleymanAI score40

    Microsoft AI launches MAI-Transcribe-2-Streaming, claiming top real-time transcription accuracy

    AIMicrosoft AI launched MAI-Transcribe-2-Streaming, which Artificial Analysis ranks #1 of 38 models for final transcript accuracy at 2.5% WER, returned 0.13s after end of speech. Artificial Analysis lists its streaming price at $0.54 per hour of audio, at the higher end among leading streaming models. Microsoft's post claims the model is 55% faster and 60% cheaper than ElevenLabs and invites developers to build agents on its platform.

  8. Jerry LiuAI score42

    LlamaIndex launches Extract v2.5 document extraction agents with improved accuracy

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction across cost-effective, agentic, and agentic plus tiers. The company reports the agents outperform Opus 5.5 and GPT-6 Sol while costing 30% to 4x less, with accuracy gains on long lists (86.1% to 95.5%), multi-page records (85.5% to 96.5%), and scanned forms (90.9% to 95.7%) on its agentic tier. The release adds advanced citations with bounding boxes and structural reasoning, and the agents are available on LlamaParse.

    Video from @jerryjliu0's post
  9. Lewis Tunstall @ COLM 🌉AI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  10. Cloudflare Blog · AIAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.