Skip to contentSkip to stories

Updated

#Eval/Benchmark

Oct 7

Oct 7Wed
  1. Hugging Face BlogAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    AILiquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  2. LlamaIndexAI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

  3. Elvis SaraviaAI score18

    Viktor, a Slack AI employee, reviews overnight agent eval failures

    AIElvis Saravia describes using Viktor, an AI employee in Slack, to review his nightly agent harness evaluation results. Viktor traces tasks that regressed from passing to failing back to the specific harness change that caused them and suggests reverting it, while the human makes the final decision. The post is a sponsored partnership, offering $100 in free credits with no card required.

  4. Elvis SaraviaAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    AINVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.

  5. Ars Technica · AIAI score63

    Mistral releases Le Chonk, a 1 trillion-parameter open-weight model

    AIMistral has released Mistral Large 4, nicknamed Le Chonk, a 1 trillion-parameter model it says can be used and customized by anyone. It is in preview, with a final version due by the end of the month, and is optimized for coding and cyberdefense as well as manufacturing, finance, and electrical engineering tasks. Mistral claims it is the most capable open-weight model developed outside China and says it was trained from scratch rather than through distillation.

  6. IEEE Spectrum · AIAI score32

    HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object Interaction

    AIHiPHI is a 617.5-hour whole-body human motion dataset captured with optical motion capture at sub-millimeter accuracy, including 245.7 hours of human-object interaction with synchronized object trajectories and meshes. The dataset organizes coverage using FrameNet, a linguistic framework for human action. The white paper also reports results from policies trained on HiPHI and deployed on a physical Unitree G1 humanoid robot.

  7. Hugging Face BlogAI score53

    TII releases Falcon-ASR, a 1.6B speech recognition model focused on Emirati Arabic

    AIThe Technology Innovation Institute introduces Falcon-ASR, a 1.6 billion parameter speech recognition model for Arabic with a focus on the Emirati dialect. On six Arabic test sets it reports an average word error rate of 20.92%, versus 23.17% for the best published leaderboard result it compared against. The model also transcribes English, French, Spanish and Portuguese with the same weights, and a demo Space is available while API access and native apps are planned.

  8. Hugging Face BlogAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    AINVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

  9. Semafor · TechnologyAI score62

    OpenAI's announced math breakthroughs prompt debate over AI's role in proofs

    AIOpenAI announced hundreds of mathematical breakthroughs, weeks after claiming it had solved one of the most complicated problems in mathematics. The findings raised questions about whether the model used creative thinking or only completed the final steps of human work. Experts say AI could be revolutionary for mathematics if it provides proofs, since proof techniques often underpin other breakthroughs.

  10. Latent SpaceAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

  11. Claude BlogAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

  12. Artificial Analysis ArticlesAI score60

    Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index

    AIAnthropic released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku release. Pricing is $0.10/$0.50 per 1M input/output tokens up to 100k tokens, rising to $0.50/$2.50 above that, but at max effort it uses about 162k output tokens per Intelligence Index task, roughly 3x GPT-6 Luna.

    Why it matters: The benchmark shows Haiku 5.5 scores well but uses far more output tokens than GPT-6 Luna, so cost per task matters beyond list price.

  13. Claude BlogAI score66

    Claude skill commands build evals and hillclimb them against overfitting

    AIAnthropic added build-eval and hillclimb commands to its claude-api skill for designing evaluations and iteratively improving applications against them. The article covers eval design principles, including production-representative tasks, headroom and low variance, and guards against overfitting through train/test splits. Two examples report results: a customer support benchmark where cost fell to under half while accuracy rose, and a claude-api skill eval that rose from 66% to 88%.

    Why it matters: The article gives a concrete workflow for designing evals and hillclimbing without overfitting, with two worked cost and performance examples that show the tradeoffs.

Oct 6

Oct 6Tue
  1. meng shaoAI score62

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind released EmbeddingGemma 2, an open 740M-parameter embedding model that maps text, code, images, video, and audio into one 768-dimensional space. Text-only use needs a 270M-parameter footprint, about 191MB active RAM when quantized on a Pixel 11 Pro, while loading all modalities takes about 567MB. The reported MTEB Code NDCG@10 score is 78.68, about 14% above the first generation, and MTEB Multilingual v2 is 61.36, roughly flat.