Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. Boris PowerAI score22

    Frontier AI research taste reportedly doubling every three months since December 2025

    AIResearch by pzeroresearch estimates that frontier models' experimental research taste has doubled roughly every three months since December 2025, with Opus 5.5 now exceeding their expert human baseline. The author of the main post, Boris Power, calls the plot very interesting for recursive self-improvement implications, while noting that the details matter for doing useful work at frontier labs.

  2. elvisAI score41

    Parsewave audit fixes 206 verifier bugs in AutomationBench

    AIParsewave audited all 600 public tasks in Zapier's AutomationBench and human review confirmed 206 real verifier bugs, all of which were fixed in AutomationBench Verified. Replaying 1,235 Kimi K3 runs on the old and fixed verifiers changed 27.9% of grades, with pass rate rising from 18.8% to 43.8% where verifiers were too strict and falling from 60.2% to 49.7% where they were too lenient.

  3. ARC PrizeAI score22

    Grok 4.7 uses more reasoning tokens than Grok 4.6 on ARC-AGI-2

    AIGrok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning levels, raising its cost per task. Per test-pair attempt, medium used 136% more tokens, high 125% more, and xhigh 173% more, while low used 27% fewer. A chart compares the two models at xhigh on the 20 public tasks where Grok 4.7 increased token use the most.

    Image from @arcprize's post
  4. Microsoft ResearchAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  5. Google DeepMind · The KeywordAI score72

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.

    Why it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.

  6. ARC PrizeAI score14

    ARC Prize hosts Frontier AI benchmarking dinner with Snorkel AI during SF Tech Week

    AIARC Prize is hosting its first Frontier AI Benchmarking Ecosystem Dinner with Snorkel AI during SF Tech Week. The event will bring together 40 experts from nonprofit benchmarking organizations, industry, academia, and government to discuss AI benchmark design and building a more scalable, transparent evaluation ecosystem. Attendees named in the post include representatives from METR, Epoch, Vals, Artificial Analysis, Harvard, and Stanford.

  7. ARC PrizeAI score25

    ARC Prize finds DeepSeek V4.1 Flash high reasoning gains no clear edge

    AIOn ARC-AGI-1, DeepSeek V4.1 Flash scored 88.5% at high reasoning versus 90.5% at low, with high using 35% more output tokens without consistently better answers. On ARC-AGI-2, the reported per-task cost of max reasoning ($0.129) appears slightly lower than high ($0.133), but after excluding incomplete tasks caused by API issues, max is about 4.5% more expensive per task.

    Image from @arcprize's post
  8. SemiAnalysisAI score18

    ClusterMAX rates FarmGPU underperform on Slurm and Kubernetes testing

    AISemiAnalysis rated FarmGPU as ClusterMAX Underperform after its Slurm layer failed to advertise GPU resources and Kubernetes exposed no RDMA devices for scale-out networking. The post credits FarmGPU's Grafana monitoring, provisioning notes, and trustworthy technical team, while noting the team may be stretched thin across small clusters.

    Image from @SemiAnalysis_'s post
  9. Simon WillisonAI score36

    Mistral's Pelican SVG Test Passes, Tied to Mistral Large 4 Context

    AISimon Willison reports that Mistral can now generate his pelican SVG test, shared via a Markdown SVG renderer. The post links to a rendered result but gives no benchmark or scoring details. Background from Mistral's own announcement describes Mistral Large 4 as a 1T-parameter, natively multimodal model with 49B active parameters, available via API today and with open weights planned for end of October.

    Image from @simonw's post
  10. Aravind SrinivasAI score42

    Perplexity Computer plays real-time StarCraft against itself, Blue wins 2-5

    AIPerplexity's Computer ran two agents playing StarCraft against each other in real time, with the game never paused while each agent thought. Blue, playing with 41 Dragoons, lost the final match 2-5 to Red, which used High Templar and Psionic Storm after Blue failed to scout Red's build. Each agent received only its own fog-of-war-limited game state, and video input was not provided.

    Video from @AravSrinivas's post
  11. Sophia YangAI score45

    Mistral Large 4 tops benchmarks across cybersecurity, legal, and agentic tasks

    AIMistral Large 4 is a 1T-parameter natively multimodal model with 49B active parameters, which the Mistral account says leads open-weights models from the US or Europe on aggregated benchmarks. The post claims it beats closed frontier models on visual grounding and posts strong results across cybersecurity, legal, and agentic behavior. It is available via API now, with open weights due at the end of October.

    Image from @sophiamyang's post
  12. Guillaume Lample @ NeurIPS 2024AI score40

    Mistral's ML4 hits open-model SOTA across capabilities and cyber benchmarks

    AIMistral says its ML4 model reaches state-of-the-art performance among open models across a wide range of capabilities, and outperforms the best models in visual grounding, legal, and spreadsheet manipulation. The post reports ML4 ranks among the best on the AA Cyber Index, scoring 82% on vulnerability reproduction and patching and 93% on Cybench. It argues that self-hosted, auditable open models are the best defense option for enterprises today, and that they do not refuse to help.

    Image from @GuillaumeLample's post