Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 14

Sep 14Mon
  1. Intern Large ModelsOfficialAI score62

    Intern-S2-397B released in BF16 and FP8 under Apache 2.0

    AIShanghai AI Laboratory's Intern Large Models announced Intern-S2-397B, available in BF16 and FP8 under Apache 2.0. The post reports 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both, and says it was jointly trained across 20+ scientific domains with long-horizon agent RL.

    Why it matters: The post names the benchmark scores and training scope behind Intern-S2-397B, letting readers compare its scientific and agentic claims against the table.

  2. Intern Large ModelsOfficialAI score62

    Intern-S2-397B: Shanghai AI Lab releases open multimodal model for scientific research

    AIIntern Large Models introduces Intern-S2-397B, a multimodal foundation model built for long-horizon scientific research and scientific agents. The post reports leading open-source results on IMO-Proof and AdvancedMathBench, and says the model reaches the level of Gemini 3.1 Pro on those tasks. It is now supported by vLLM and SGLang, with weights on Hugging Face and ModelScope and a chat demo available.

    Why it matters: The post pairs a new open multimodal model with benchmark tables against named Qwen, DeepSeek, Kimi, GLM, GPT, Gemini, and Claude models, letting readers compare scientific and agentic results directly.

    Image from @intern_lm's post
  3. Sherwin WuXAI score40

    GPT Image 2.5 now live in ChatGPT with better editing consistency

    AIGPT Image 2.5 is now available in ChatGPT, and the post says it keeps consistency well while editing images. It also claims state-of-the-art results on all image leaderboards, and suggests users who saw faces shift in GPT Image 2 edits try again.

    Video from @sherwinwu's post

Sep 13

Sep 13Sun
  1. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score36

    SingProbe adds a streaming guardrail to Step-3.7-Flash without a separate safety model

    AIinclusionAI released Step-3.7-Flash-singprobe, an 8.13M-parameter probe that reuses Step-3.7-Flash hidden states to score query intent, response unsafety, and hallucination risk at every generated token. The probe adds less than 0.5% decode-time overhead and reports 0.9858 R-AUC and 0.9295 T-AUC on streaming safety benchmarks. It is supported through SGLang and vLLM integration branches and loads from Hugging Face by checkpoint ID.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score40

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.5-397B-A17B

    AIinclusionAI has released Qwen3.5-397B-A17B-singprobe, an intrinsic streaming guardrail built on Qwen/Qwen3.5-397B-A17B that scores query intent, response unsafety, and hallucination risk at every generated token using the base model's hidden states. The probe has 8.13M parameters, taps layers 18, 38, and 58, and adds less than 0.5% decode-time overhead. Training code is available at inclusionAI/SingProbe, and the probe runs through SGLang or vLLM integration branches.

  3. Fireworks AI BlogOfficialAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.

  4. Sebastian RaschkaXAI score35

    Raschka's Reasoning from Scratch Round 3 Builds a Math Verifier

    AISebastian Raschka's third "Reasoning from Scratch" video covers building a math verifier for evaluating language models and for later reinforcement learning with verifiable rewards (RLVR) training. The walkthrough covers extracting final answers from boxed outputs, normalizing them, checking mathematical equivalence, and running evaluation on the MATH-500 dataset.

    Video from @rasbt's post
  5. Mike KnoopXAI score50

    Mike Knoop argues intelligence is capped at optimal decision-making

    AIMike Knoop argues intelligence can be measured as the ratio of a decision's quality to the optimal decision, capped at 100%. He says Astra is already 80% optimal on ARC v3 speedruns and identifies horizontal data acquisition and efficiency/cost as the most plausible near-term areas for RSI. Background from @mhmazur reports that GPT-6 Astra scored 100% on the 25 ARC-AGI-3 public games using 6,485 actions versus a human baseline of 17,135.

Sep 12

Sep 12Sat
  1. Mike KnoopXAI score46

    Mike Knoop urges keeping AI research open amid slowdown proposals

    AIMike Knoop says he sees a path to an ARC-AGI-4 benchmark focused on open-ended invention, which he calls the gating capability between zero-sum automation and positive-sum innovation. He argues that coordinated slowdown efforts would likely apply to everyone, including open-source work, and cites chain of thought and the transformer as inventions that grew out of open science research. He concludes the research frontier must stay open to keep humanity on a positive-sum path.

  2. Epoch AI · The Epoch BriefOfficialAI score60

    Epoch Brief covers Huawei chips, Nvidia's GDP effect, and GPT-6 Astra benchmarks

    AIEpoch AI's newsletter reports that Huawei is far behind Nvidia and is unlikely to catch up this decade due to export controls. It also finds official US GDP statistics understate growth by about 0.3 percentage points over the past year, and that GPT-6 Astra set new records on Epoch's evaluations, including the Epoch Capabilities Index.

    Why it matters: The newsletter bundles several analyses of AI chips, GDP measurement, and benchmarks, so it helps readers scan the research agenda behind each finding.

  3. The Algorithmic BridgeBlogAI score52

    AI's Math Breakthroughs Could Starve Mathematics of the Hard Problems It Needs

    AIAlberto Romero argues that AI solving Millennium Prize problems in 2026 threatens mathematics through success, not failure. He draws on Terence Tao's view that struggle shapes mathematicians, and that proof abundance without hard problems could leave fields depleted, like overplanted farmland.

Sep 11

Sep 11Fri
  1. InferactOfficialAI score34

    vLLM adds day-0 support for DeepSeek v4.1 Flash across six NVIDIA GPUs

    AIInferact says vLLM now supports DeepSeek v4.1 Flash on day zero across H100, H200, B200, B300, GB200, and GB300 GPUs. SemiAnalysis independently verified the NVIDIA support, while the post notes AMD vLLM still does not work with the model. Serving recipes are available at recipes.vllm.ai.

  2. Baseten BlogOfficialAI score62

    DeepSeek-V4.1-Flash arrives on Baseten with a split prefill architecture

    AIDeepSeek released open weights for V4.1-Flash, which Baseten now offers through its Model APIs. The model has 552B total parameters, 8B active for prefill and 16B for decode, a 1M token context window, and text plus image input. Its Causal Encoder-Decoder design runs only the encoder during prefill and reuses a projected KV cache, and the source reports the global KV cache at a quarter of V4-Flash's memory.

    Why it matters: The post explains how the CED architecture splits prefill and decode compute and cuts KV cache memory, which matters for coding agent costs.

  3. VercelOfficialAI score26

    Tailscale's Aperture model router is built on Vercel AI Gateway

    AITailscale offers instant access to hundreds of models for any user in a secure tailnet through its customer-facing model router, Aperture. Aperture is built on Vercel's AI Gateway and offers zero data retention, zero markup with free BYOK, and cost and usage data on every response.

  4. Redwood Research BlogBlogAI score62

    Prompt tuning lifts CoT controllability scores on open models

    AIRedwood Research reports that better prompt templates raise chain-of-thought controllability scores on the CoTControl eval for open-source reasoning models by roughly 2-3x or more. For example, GPT-OSS-120B rose from 5.5% to 15% in the zero-shot setting. The author concludes that current CoT controllability numbers may underestimate what models can do, though the finding does not significantly undermine the view that current models probably cannot consistently evade CoT monitoring.

  5. BAAIOfficialAI score46

    BAAI unveils AREX, a 122B MoE research agent for hard search

    AIBAAI introduced AREX, a research agent built on a 122B-parameter mixture-of-experts model with 10B active parameters. It drafts candidate answers, checks each constraint, and revisits unresolved points rather than running one long search. The post says AREX performs on hard search benchmarks comparable to GPT-5.4.

    Video from @BAAIBeijing's post

Sep 10

Sep 10Thu
  1. Ai2 · new models on Hugging FaceOfficialAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  2. Greg BrockmanXAI score36

    GPT-6 Astra tops frontier models in antibody developability prediction

    AIOpenAI's GPT-6 Astra is reported as the best-performing frontier model for antibody developability prediction in one benchmark, outperforming other frontier models tested on properties such as aggregation and stability. The post also says Astra built an interactive antibody visualization in about one hour.

  3. Bryan CatanzaroXAI score25

    Bryan Catanzaro praises humansand's Persimmon model for human interaction

    AINVIDIA's Bryan Catanzaro welcomed the idea of models that understand how humans interact and thanked builders using Nemotron Ultra. Humansand introduced Persimmon, which it describes as the first large-scale model designed to realistically simulate how people talk and interact.

  4. Amazon ScienceOfficialAI score40

    Amazon research explains why ML research agents don't overfit benchmarks

    AIAmazon Science researchers propose that machine learning research agents avoid overfitting benchmarks despite years of iteration against the same tests. They attribute this to generalizable strategies being expressed compactly, leaving no room for memorization, while overfitting strategies fail to survive a compression bottleneck.

  5. Cognition Blog (Devin, Windsurf)OfficialAI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

  6. Mustafa SuleymanXAI score38

    Microsoft says five of its last eight AI models debuted at No. 1 on leaderboards

    AIMustafa Suleyman says Microsoft launched eight new models in two months, and five debuted at #1, including MAI-Transcribe-2 and MAI-Image 2.5, 2.6, and 2.6 Flash on the Artificial Analysis leaderboard. He says MAI-Cyber in MDASH ranked #1 on CyberGym while cutting costs by 50% versus other frontier models. He adds that MAI-Code-1.1-Flash, launched in GitHub Copilot four weeks ago, now accounts for a third of small-model traffic there.

  7. Amazon ScienceOfficialAI score62

    Amazon Science explains why ML research agents don't overfit benchmarks

    AIAmazon Science says LLM research agents that repeatedly optimize against a validation set tend not to overfit because successful strategies can be compressed into short prompts. In its experiments, 32-token prompts let a fresh reproducer match the explorer's models on most of eight datasets, and agents forced to overfit failed this compression test. The authors note that the framework assumes no side channel from pretraining memorization and that fully resolving this may require fresh datasets collected after training cutoffs.

    Why it matters: The post explains a testable compression argument for why benchmark-driven research agents avoid overfitting, which helps readers judge when validation gains are likely to transfer.

  8. Tencent HyOfficialAI score60

    Tencent Hunyuan releases open-source AuK audio model for speech generation and editing

    AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.

    Why it matters: The release combines speech generation and editing under one natural-language interface, and its 4-step AuK-Flash variant reports about 4.5 times faster inference under matched conditions.

    Video from @TencentHunyuan's post
  9. Chips and CheeseBlogAI score46

    Geekbench 7 Shows Binary Translation Costs Snapdragon X2 Elite Performance

    AIGeekbench 7 testing on the Snapdragon X2 Elite shows x86-64 binaries running through Windows 11's Prism translator lose substantial performance compared with native aarch64 execution. Binary translation roughly doubles executed instructions when running the x86-64 version, and every tested core, including Qualcomm's, takes a notable penalty. Even with that penalty, the Snapdragon X2 Elite's E-Cores outperform Neoverse N1 and its P-Cores outperform Neoverse N2.

  10. DeepSeekOfficialAI score38

    DeepSeek unveils 552B MoE model with asymmetric encoder-decoder design

    AIDeepSeek has introduced a 552B-parameter MoE model built on a new Causal Encoder–Decoder architecture, activating 8B parameters for input and 16B for output. The company says new pre-training methods and larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

    Image from @deepseek_ai's post
  11. DeepSeek API NewsOfficialAI score72

    DeepSeek releases V4.1-Flash with native multimodal support and API updates

    AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.

    Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.

Sep 9

Sep 9Wed
  1. DeepSeek · new models on Hugging FaceOfficialAI score78

    DeepSeek-V4.1-Flash releases a multimodal MoE model with 1M-token context

    AIDeepSeek released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters and support for contexts up to one million tokens. The technical report says its global KV cache footprint is 890 bytes per token, roughly one quarter of DeepSeek-V4-Flash, and reports 8B activated parameters per token during prefill and 16B during decode.

    Why it matters: The report shows KV cache per token falling to about one quarter of DeepSeek-V4-Flash, a concrete tradeoff between long-context serving cost and benchmark results.

  2. Fireworks AI BlogOfficialAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.

  3. TinkerOfficialAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.

  4. METROfficialAI score31

    METR plans investigation into AI misalignment incidents and propensities

    AIMETR says its planned investigation will cover all questions raised in its recently updated post on how independent researchers could study AI propensities after misalignment incidents. The post defines misalignment incidents as cases where an AI agent autonomously took sophisticated, sustained actions violating human intent.

  5. Cognition Blog (Devin, Windsurf)OfficialAI score82

    Cognition's Devin factors RSA-260 using a GPU lattice siever

    AICognition's Devin agent, directed by Eric Lu, factored the 260-digit RSA-260 number using a new GPU implementation of the general number field sieve built on CADO-NFS. The author estimates the run cost about 13.5 GPU-years, roughly $400k at market prices, and projects RSA-1024 factoring at around $30M, while RSA-2048 is not meaningfully affected.

    Why it matters: The source gives a full cost breakdown and scaling estimates for RSA factoring on GPUs, showing how far the cost of breaking RSA-1024 has fallen.

Sep 8

Sep 8Tue
  1. Google Developers BlogOfficialAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

  2. InferactOfficialAI score42

    Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

    AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

  3. Noam BrownXAI score36

    OpenAI model reportedly delivers a huge step up over today's LLMs

    AINoam Brown says a plot shows the model OpenAI used is a major advance beyond today's LLMs, and that no one relied on Levent's or Tristan's prompts. He was responding to questions about how Navier-Stokes might be achieved and the attention given to those prompts. Background from Sebastien Bubeck describes coordination over Euler and Navier-Stokes results, including a disputed suggestion about Levent's authorship.

    Image from @polynoamial's post
  4. BAAIOfficialAI score34

    Robot models excel at single moves but fail chained tasks

    AIBAAI reports that robot models trained on individual skills such as grasping, placing, pulling, and opening performed poorly when asked to chain them into full tasks without extra practice. The best score was 16.7%, and some models scored zero. The post's example notes a robot may open a drawer yet get stuck on the handle.

    Image from @BAAIBeijing's post
  5. BAAIOfficialAI score46

    Top embodied models hit 98% in sim but drop sharply on hardware

    AIIn simulation, the best embodied AI models complete easy tabletop tasks about 98% of the time. On physical Franka robots, their success rate falls to 24%–72% of simulated performance, and to 13%–60% in a dual-arm setup. Models that look tied in simulation can differ by 30 points on real hardware.

    Image from @BAAIBeijing's post
  6. BAAIOfficialAI score43

    FlagEval-Robo tests 12 open-weight embodied AI models across simulation and real robots

    AIBAAI introduces FlagEval-Robo, an open dual-track evaluation suite linking simulation with real-world execution. The team post-trained and stress-tested 12 leading open-weight embodied AI models under strictly aligned conditions. The post raises whether high benchmark scores reflect physical reality, though it does not yet report specific results.

    Image from @BAAIBeijing's post
  7. Google DeepMindOfficialAI score74

    Google DeepMind launches AlphaGenome Atlas to predict 9 billion DNA variant effects

    AIGoogle DeepMind has introduced AlphaGenome Atlas, a platform with predicted molecular effects for 9 billion single-nucleotide variants in the human genome. It is free for academic research through a web portal, and the AlphaGenome Variant Impact score condenses predictions from AlphaGenome and AlphaMissense into one number for ranking variants. The source says collaborators used it to identify variants in unsolved rare disease cases and to find rare non-coding variants linked to traits.

    Why it matters: The source details how precomputed variant predictions, a single impact score, and linked feature attributions make genome-wide mutation effects searchable for researchers without coding skills.