Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 4

Oct 4Sun
  1. Jerry LiuAI score34

    LlamaIndex launches Extract v2.5 document extraction agents, cutting errors on scanned forms

    AILlamaIndex introduced Extract v2.5, a series of agents tuned for document extraction, including cost-effective, agentic, and agentic plus tiers, available in LlamaParse. The company says the agents reduce error rates by 2x or more compared with frontier models at a small fraction of the price, and they handle handwritten and drawn annotations on scanned documents while grounding values in the source text.

    Video from @jerryjliu0's post

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

  2. IndexTeam (Bilibili) · new models on Hugging FaceAI score27

    Index-Echo-S2ST-2B FP4 Quantized Speech-to-Speech Translation Model Released on Hugging Face

    AIIndexTeam released Index-Echo-S2ST-2B-FP4, an NVFP4 (W4A4) quantized version of the Index-Echo-S2ST-2B speech-to-speech translation model, with only the text LLM backbone quantized and the audio components kept in BF16. On a fixed corpus, perplexity rose from 5.9332 to 6.4980 (+9.52%), while zh->en and en->zh generations matched the original. Full FP4 acceleration requires an NVIDIA Blackwell GPU, and the model loads via compressed-tensors in vLLM or transformers.

  3. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-2B speech translation model

    AIIndexTeam has published an official NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-2B speech-to-text translation model on Hugging Face. Only the text LLM backbone is quantized, while the audio tower, connector, and speech-synthesis components remain in BF16. Perplexity rises 5.80%, from 4.8772 to 5.1599, on a fixed corpus, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  4. Latent SpaceAI score52

    Latent Space daily roundup covers GPT-6.1 Sol, Sonnet 5.5, agent harnesses, and eval integrity debates

    AIThis Latent Space AINews roundup compiles a weekend's AI news from Twitter and Reddit rather than a single announcement. It covers OpenAI's GPT-6.1 Sol pricing and Agent Arena placement, Anthropic's Sonnet 5.5 debut, Meta's open-sourced Muse hardware firmware, and several research and benchmark items, many reported with unverified claims.

Oct 2

Oct 2Fri
  1. Jerry LiuAI score34

    LlamaIndex's Extract v2.5 agents reason over tables spanning multiple pages

    AILlamaIndex introduced Extract v2.5, a set of document extraction agents that can reconstruct records split across pages and assemble them with thousands of other cells into structured tabular output. The post says the agents handle real-world documents like insurance claims, regulatory filings, and legal schedules, where a record may start on one page and finish on the next. The accompanying background post claims record-spanning-page accuracy rose from 85.5% to 96.5%, and that the agentic tier outperforms Opus 5.5 and GPT-6 Sol at 30% to 4x lower cost.

    Video from @jerryjliu0's post
  2. Prime IntellectAI score43

    CMU's SMDD-Bench adds 502 drug design tasks for RL training

    AICMU researchers released SMDD-Bench, a benchmark of 502 small-molecule drug design tasks that use RDKit, ADMET-AI, and Boltz-2 as feedback loops. The authors argue that long-horizon planning, exploration, and learning from imperfect feedback remain open problems beyond math and coding, and the benchmark is available in Prime Intellect's Environments Hub for training with prime-rl.

  3. Baseten BlogAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  4. Design ArenaAI score40

    GPT-6 Astra hedges far more than Claude Opus 5.5 in reasoning summaries

    AIDesign Arena analyzed 324 thinking summaries and found OpenAI's GPT-6 Astra uses hedging words like "maybe," "might," and "it seems" about 20 times as often as Anthropic's Claude Opus 5.5. Opus usually weighs a few options and commits early, in about 4 out of 5 summaries versus 1 in 4 for Astra, which the post says works more like a designer while Opus works more like a builder.

    Video from @DesignArena's post
  5. Aravind SrinivasAI score62

    Perplexity open-sources models, an inference engine, and security tools

    AIPerplexity has released several open source projects, including the pplx-decider-v1-27b multimodal decision model, the pplx-embed-v2-context-9b-preview contextual embeddings model, and the Lily local inference engine for Apple silicon. The post also lists the 0.6B on-device PII-Tracer classifier with its PII-TRACE benchmark, the WANDR research agent benchmark, and the Numbat and Bumblebee security tools, and says more open source releases are coming soon.

  6. SGLangAI score39

    SGLang adds a scoring API and multi-item scoring for decision models

    AISGLang's update adds a /v1/score endpoint that returns scores for requested labels such as Yes/No or A/B/C, avoiding the label loss of generate with top-k logprobs. Its multi-item scoring computes shared context once and keeps each candidate isolated, with 16-candidate p95 on Qwen3-8B dropping from 54.1 ms (Generate) to 20.6 ms.

  7. Kilo (acq. by Anaconda)AI score36

    Ling 3.1 Flash is free in Kilo Code until October 13

    AIKilo Code is offering Ling 3.1 Flash for free until October 13, with the model served by Novita Labs. Ant Ling's background post describes the model as roughly 560B total parameters with about 25B active per token and a context window of up to 1M tokens. Ant Ling says it scores 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional, and plans to open-source it soon.

  8. Jerry LiuAI score44

    LlamaIndex Extract v2.5 hits 93–96% on dense table extraction benchmarks

    AILlamaIndex released Extract v2.5, a set of document extraction agents that it says reach 93%–96%+ accuracy on long-list extraction, including records spanning pages. The post claims the agents outperform frontier VLMs, which it says stop early, miss repeated records, and struggle to attribute values to sources, while LlamaIndex attributes every extracted value to its source. The agents are available through LlamaParse.

    Video from @jerryjliu0's post
  9. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  10. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  11. Lucas Beyer (bl16)AI score45

    Lucas Beyer praises new coding benchmark for finding bugs in repos

    AILucas Beyer calls SWE-sweep a useful new benchmark, where agents must find and fix bugs in a repo checked out at an earlier commit, scored against unit tests from real later bugfixes. He notes two limitations: a model may find valid bugs that don't match the tested ones, and the construction makes training on the test set easy. He advises not overemphasizing small ranking differences once models score highly.

  12. Hugging Face BlogAI score62

    AutoSynthData generates targeted training data for enterprise agents from failures

    AIServiceNow CoreAI introduced AutoSynthData, which uses a target model's failures and a stronger teacher's successes to generate and validate new agent training tasks. In EnterpriseOps Gym experiments, the Hybrid domain produced 2,000 samples and raised Gemma-4-26B-A4B-it mean Pass@1 by 7.2 percentage points, while the ITSM domain produced 1,994 samples and raised it from 18.77% to 27.18%.

    Why it matters: The post shows how failure analysis, teacher demonstrations, and verifier checks combine into a repeatable pipeline for generating targeted agent training data.

Oct 1

Oct 1Thu
  1. NVIDIA AIAI score44

    CoreWeave RL rollouts reload model weights 15× faster with Dynamo

    AICoreWeave's new RL rollouts service uses ModelExpress and Router in NVIDIA Dynamo to speed up model weight reloads during RL post-training with minimal downtime. Working with NVIDIA and You.com, CoreWeave achieved 15× faster model reloads than its baseline while post-training Nemotron 3.5 Lightning. The speedup addresses GPUs sitting idle while inference workers wait to load updated weights between training iterations.

  2. Apple Machine Learning ResearchAI score34

    Limits of Confidence-Based Sampling in Discrete Diffusion Models

    AIApple Machine Learning Research reports that discrete diffusion steps match the training distribution only when simultaneously written token positions are conditionally independent given already-fixed tokens. The authors show that per-position distributions cannot determine such dependence, and on the synthetic ScanAndAdd task, confidence-ranked groups of two or more positions were dependent and produced a generated distribution 29 times the sampling-noise floor in total variation.