Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Sep 4

Sep 4Fri
  1. Lewis Tunstall @ COLM 🌉AI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

  2. Tencent · new models on Hugging FaceAI score36

    Tencent Releases EVIE-8B Open-Source Visual Document Retrieval Model

    AITencent has open-sourced EVIE-8B, an 8.4B-parameter visual document retriever that scores 66.75 nDCG@10 on ViDoRe V3 and ranks first on that leaderboard's mean task score of 66.24. The model uses 4096D per-token multi-vector embeddings with MaxSim late-interaction scoring and bidirectional attention, and it serves as the teacher for the lightweight EVIE-4.5B model. Model weights, inference pipelines, and evaluation suites are available, while the formal research paper is promised for a future release.

  3. Tencent · new models on Hugging FaceAI score36

    Tencent Open-Sources EVIE-4.5B Visual Document Retrieval Model With Elastic Embeddings

    AITencent released EVIE-4.5B, a 4.5B-parameter visual document retrieval model, with weights, training pipelines, HAC token compression, and evaluation suites open-sourced on Hugging Face. It scores 66.02 on ViDoRe V3 and ranks second on that leaderboard behind the 8.4B EVIE-8B, which scores 66.24. Its Prefix-MRL head lets a single 2048D projection be truncated to 64–2048 dimensions at runtime without separate models.

Sep 3

Sep 3Thu
  1. TinkerAI score25

    Tinker highlights training objectives for legible chain-of-thought and interpretability evals

    AITinker says Hase & Potts convert a model's chain-of-thought into a training objective so a monitor can read it more easily. Karvonen et al. use tested counterfactual outputs to build an interpretability eval. The post notes that counterfactuals do not explain the underlying mechanism, but their predictability is a useful foundation.

  2. TinkerAI score51

    Bespoke Labs post-trains Inkling on one code repo and reports broader coding gains

    AIBespoke Labs post-trained the Inkling base model on a single GitHub repository using supervised fine-tuning and GRPO reinforcement learning. The post reports a 57-point improvement on the held-out fontTools evaluation over the base model, along with gains on Terminal-Bench 2.1 and SWE-bench Lite. It also says the post-trained model uses about 40% fewer tokens.

    Image from @tinkerapi's post
  3. Google DeepMind · The KeywordAI score72

    Google DeepMind releases WeatherNext 3, a global weather model with hourly satellite-based forecasts

    AIGoogle DeepMind and Google Research introduced WeatherNext 3, which generates hourly global forecasts at up to 5-kilometer resolution using live geostationary satellite data. The company reports that precipitation forecasts improved by up to 60% against IMERG in medium-range evaluations, and that longer-range precipitation forecasts are up to 50% more accurate. The model is now available across Search, Gemini, Google Maps, Google Maps Platform Weather API, Google Earth Engine, BigQuery, and Google Cloud Storage.

    Why it matters: The post explains how training on live satellite data and station observations changes resolution and update frequency, with precipitation accuracy gains reported against named baselines.

Sep 2

Sep 2Wed
  1. TinkerAI score44

    Lightning Rod's new work shows scoring rules reshape LLM forecaster profiles

    AILightning Rod, working with Philip Tetlock and Ville Satopää, post-trained five versions of the same LLM that differed only in the scoring rule used as the RL reward. The versions reached similar aggregate scores but had very different bias, information, and noise (BIN) profiles, so a good Brier score alone does not show whether a forecaster can distinguish likely from unlikely events.

  2. ARC PrizeAI score77

    OpenAI's GPT-6 Astra scores 62.7% on ARC-AGI-3 Semi-Private

    AIOpenAI's GPT-6 Astra (max) scores 62.7% on ARC-AGI-3 Semi-Private for $26K under the Standard harness, and 99.9% for $19K under the Provider Adapter harness. The authors say Astra used fewer actions than the human baseline on 96.0% of levels, and they note it is not claimed to be AGI.

    Why it matters: The report pairs benchmark scores with replays of the model's notation and tool use, showing how it solved unfamiliar environments rather than only that it did.

  3. The Register · AIAI score39

    AI Models Misidentify Mushrooms in Test, Sometimes Calling Deadly Species Edible

    AIPiotr Migdał tested 16 AI models on 1,040 mushroom photos covering 55 species, and the best, Gemini-3.8-flash, was correct on its first guess only 65 percent of the time. Dangerous mistakes were common, with the death cap called edible 16 percent of the time, and Qwen3.8-27b wrongly labeled poisonous mushrooms edible 36 percent of the time. Migdał warns users not to eat any mushroom because an AI says it is safe.

  4. NVIDIA · new models on Hugging FaceAI score67

    NVIDIA releases Nemotron-3-Labs-Ultra-Math-RL for mathematical proof reasoning

    AINVIDIA has published Nemotron-3-Labs-Ultra-Math-RL on Hugging Face, a 550B total, 55B active parameter model for solving difficult math problems and identifying proof mistakes. The model is part of an ensemble that reached gold-medal level at the International Mathematical Olympiad 2026, and it is available for commercial and non-commercial use under the OpenMDW-1.1 license. Deployment is designed for NVIDIA Blackwell or Hopper GPUs, with a recommended minimum of 8× B200 on a single node and a context length of up to 1M tokens.

    Why it matters: The release details the model's math-proof role, its 550B total and 55B active parameters, and its vLLM deployment requirements for teams weighing adoption.

  5. Google AI StudioAI score78

    Google releases Gemini 3.8 Flash and restricted 3.8 Flash Cyber model

    AIGoogle introduces Gemini 3.8 Flash for coding, agentic tasks, and multi-step reasoning, priced at $0.75 per million input tokens and $3.75 per million output tokens during the introductory period. Gemini 3.8 Flash Cyber targets vulnerability detection and automated patching and is available only to trusted defenders through the new Fairwind Program. The introductory price expires December 31, 2026, after which $1.50 and $7.50 per million tokens apply.

    Why it matters: The post separates a general coding and agent model from a restricted cyber variant, showing how one shared core is deployed under different access and safety tiers.

  6. Sundar PichaiAI score62

    Google introduces Gemini 3.8 Flash Cyber, a cybersecurity model for vulnerability work

    AIGoogle introduces Gemini 3.8 Flash Cyber, which it describes as its most capable cybersecurity model. The company reports 86.2% on CyberGym, 47.2% on CWE-Bench for patching, and a 70%+ success rate in discovering vulnerabilities across 20 programming languages on its internal benchmark. Google says the model offers frontier-level performance at Flash-level speed and pricing.

    Image from @sundarpichai's post
  7. koray kavukcuogluAI score62

    Gemini 3.8 Flash claims stronger engineering results at lower cost than larger models

    AIGoogle's Koray Kavukcuoglu says Gemini 3.8 Flash is a major step up from Gemini 3.7 Flash and outperforms most larger frontier models on complex engineering problems at a fraction of the cost. The attached DeepSWE V1.1 chart, sourced to Datacurve AI, plots average cost per task against score for Gemini 3.8 Flash and other models. A link to Google's blog post with more details is included.

    Image from @koraykv's post

Sep 1

Sep 1Tue
  1. Anthropic · YouTubeAI score78

    Anthropic releases Claude Fable 5.1, an upgrade to its most capable model class

    AIAnthropic has released Claude Fable 5.1, the latest upgrade to its most capable class of models, and it is available everywhere today. The company says it handles complex, long-running, multi-step work and avoids shortcuts when fixing root causes of software issues. At lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost, according to Anthropic's benchmarks.

    Why it matters: The source names the upgraded model class and its cost tradeoff at lower effort levels, which helps readers weigh it against the earlier version for their own workloads.

  2. Anthropic · YouTubeAI score72

    Anthropic releases Claude Fable 5.1 for complex, long-running tasks

    AIAnthropic has released Claude Fable 5.1, an upgrade to its most capable model class, and says it is available everywhere today. The company reports that at lower effort levels, Fable 5.1 can match or beat Fable 5 at a much lower cost. It is described as strong at complex multi-step work, such as long proofs and contracts with hundreds of cross-references, and at fixing root causes in software issues.

    Why it matters: The source reports cost and effort-level tradeoffs for long-running tasks, helping readers judge whether the upgrade changes their workloads or budgets.

  3. Ai2 · new models on Hugging FaceAI score22

    Ai2 Releases Supplemental ACE2S-SHiELD+ Ablation Checkpoints on Hugging Face

    AIAi2 has published supplemental checkpoints for its ACE2S-SHiELD+ climate model on Hugging Face, covering four ablation configurations that test random CO2 data and energy conservation. Each configuration includes two random-seed models, and the repository recommends the main ACE2S-SHiELD+ checkpoint for most uses. The checkpoints are licensed under Apache 2.0 for research and educational use.

  4. Microsoft AI BlogAI score34

    Microsoft Publishes 2026 Responsible AI Transparency Report on Governance and Agentic AI Risks

    AIMicrosoft published its 2026 Responsible AI Transparency Report, its third annual edition, detailing updates to its governance and risk management. The company re-engineered its Responsible AI Standard to adapt to evolving technical risks and regulatory requirements, and is extending controls such as agent identities, tool permissions, and action monitoring to agentic AI systems.

  5. Tencent HyAI score58

    Tencent Hy4 preview reports 31.8% throughput gain from self-found bottlenecks

    AITencent Hunyuan says its Hy4 preview model found inference bottlenecks on its own and raised end-to-end throughput by 31.8% through operator fusion and communication optimizations. The post says the gain holds across context lengths and concurrency levels. The release is listed at 770B total parameters with 49B active and a 1M context window, with links to the Hy blog, Hugging Face, and GitHub.

    Image from @TencentHunyuan's post
  6. Ai2 (Allen Institute for AI)AI score56

    Ai2 introduces BenchMIRT to audit what individual LLM benchmark questions measure

    AIAi2 introduces BenchMIRT, a multidimensional item response theory method that audits LLM benchmarks at the level of individual prompts. Trained on results from 100 LLMs across 16 benchmarks, it recovered safety and general reasoning as the two dominant dimensions, and found BBQ aligns more with general reasoning than safety. Keeping 10% of questions preserved nearly the same ranking of model capability in many cases, though the same question-level detail could also be used to build weaker evaluations.

  7. OpenBMB (MiniCPM) · new models on Hugging FaceAI score49

    MiniCPM5-2B-Midtrain: OpenBMB releases mid-training checkpoint of 2B-class model

    AIOpenBMB released MiniCPM5-2B-Midtrain, a BF16 mid-training checkpoint taken before SFT in the MiniCPM5-2B series, on Hugging Face and ModelScope. The series is a 2B dense Transformer with 2,516,756,480 total parameters and a 131,072-token context length, and the final MiniCPM5-2B reports an average score of 53.9 against 51.1 for the best larger comparison model. The release also includes GGUF, MLX, and GPTQ variants, along with the UltraData datasets.

Aug 31

Aug 31Mon
  1. Liquid AI NewsletterAI score46

    Liquid AI launches Pipette, an open-source benchmark for on-device foundation models

    AILiquid AI and Artificial Analysis released Pipette, an open-source benchmark platform for foundation models on edge devices, covering over 1,000 configurations across 30+ models. It measures five on-device metrics, including throughput, latency, context scaling, and memory use, on macOS, Windows, iOS, and Android. Liquid AI also said its updated LFM2.5 Q4_0 checkpoints, trained with Quantization-Aware Distillation, retain roughly 97% of BF16 baseline performance and suffer 73.4% less quality loss than standard post-training Q4_0 quantization.

  2. DeepSeek · new models on Hugging FaceAI score65

    DeepSeek releases V4-Flash-Vision-Exp, an experimental multimodal agent model

    AIDeepSeek introduces DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model in the DeepSeek-V4 family, built on V4-Flash with visual modules. It reports substantial gains over DeepSeek-V4-Flash-0731 on multimodal agent benchmarks, such as ApexBench at 36.5 versus 26.2, while keeping text agent performance comparable. The repository provides tokenizer files, prompt encoding, vLLM and SGLang serving instructions, and is licensed under MIT.

    Why it matters: The source compares the model with its text-only predecessor and Opus-4.8 on agent benchmarks, showing where vision gains occur and where text performance holds.

Aug 30

Aug 30Sun
  1. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP Releases Core-Embed 8B for Compositional Multimodal Retrieval

    AIAlibaba NLP has released core-emb-8b, an MLLM-based multimodal embedding model that distills a reranker's compositional judgments to distinguish attribute-object bindings such as "a white plate and a black chair" versus "a black plate and a white chair." The 8B dense embedding model, built on the Qwen3-VL-based VL-Emb backbone, scores 0.666 total average on compositional benchmarks, 5.7 points above its backbone. It is part of a family that also includes 2B embedding and reranker models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score38

    Alibaba-NLP releases Core-Reranker-8B, a compositional multimodal reranker on Hugging Face

    AIAlibaba-NLP has published Core-Reranker-8B on Hugging Face, an 8B-parameter multimodal reranker fine-tuned from Qwen3-VL-Reranker to better distinguish attribute-object bindings in text and image relevance scoring. On compositional reasoning benchmarks COLA, SugarCrepe++, and NegBench, it reports an 82.7% total average, 10.7 points above Jina-Reranker. The model is part of the Core-Embed family, which also includes 2B and 8B embedding models, with Core-Embed-8B reporting a 0.666 total average.

  3. Alibaba NLP (Tongyi) · new models on Hugging FaceAI score40

    Alibaba NLP releases Core-Embed multimodal embedding models for compositional retrieval

    AIAlibaba NLP has released core-emb-2b and core-emb-8b, multimodal embedding models built on Qwen3-VL that distill reranker judgments to better match attribute-object bindings in text and image retrieval. The Core-Embed-8B model posts the best total average (0.666) among evaluated embedding models on compositional benchmarks, 5.7 points above its VL-Emb-8B backbone. Companion Core-Reranker-2B and 8B models are also available, with the 8B reranker reaching 82.7% total average on the same benchmarks.