Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 18

Aug 18Tue
  1. Liquid AI BlogOfficialAI score65

    Liquid AI releases QAD 4-bit LFM2.5 checkpoints for edge deployment

    AILiquid AI released 4-bit Q4_0 GGUF checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B, trained with Quantization-Aware Distillation. The company says the checkpoints recover most accuracy lost to quantization, reaching roughly 97% of their BF16 averages while keeping Q4_0 memory footprint and throughput. Benchmarks compare them against post-training quantized Q4_0 GGUFs and against Q5_K_M, Q4_K_M, and Unsloth's UD-Q4_K_XL.

    Why it matters: The post shows how quantization-aware distillation recovers accuracy lost in Q4_0 checkpoints, with throughput measured across four hardware backends for deployment tradeoffs.

Aug 17

Aug 17Mon
  1. Z.ai Release NotesOfficialAI score63

    Z.ai releases GLM-5.3 with stronger coding and vulnerability discovery

    AIZ.ai's release notes announce GLM-5.3, which the company says delivers a 50% gain over GLM-5.2 on Z.ai Code Bench and reaches open-source SOTA on public benchmarks including Terminal Bench 3.0. The company also reports that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, identifying 2,436 vulnerabilities in real-world targets, 1,097 of them medium- or high-severity. A separate GLM-5.3-Flash entry describes native visual capabilities and a hybrid architecture with 320B total and 18B activated parameters.

    Why it matters: The release notes show GLM-5.3's coding and cybersecurity gains, with a vulnerability count, letting readers compare it against Z.ai's prior GLM-5.x line and other coding models.

  2. Import AIBlogAI score44

    DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

    AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.

Aug 16

Aug 16Sun
  1. AI Futures ProjectBlogAI score67

    AI Futures Project shortens AI timelines and re-grades AI 2027 forecasts

    AIThe AI Futures Project says its timelines shortened slightly, with coding uplift, revenue, and time horizon methods giving similar Automated Coder arrival dates. It estimates reality is progressing at roughly 70-90% the pace of its AI 2027 scenario, and it now states its forecasts assume progress as fast as technically feasible.

Aug 15

Aug 15Sat
  1. Prime Intellect BlogOfficialAI score73

    Prime Intellect tests frontier models on 153 autonomous nanoGPT research runs

    AIPrime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, with runs lasting up to eight days on 8xH200s. The results show a large gap between models at every stage of the research process, though none of the runs produced a fundamentally new method.

    Why it matters: The experiment measures how frontier models conduct autonomous research, showing large gaps between models in experiment choice, execution, and result interpretation.

Aug 14

Aug 14Fri
  1. Cohere · new models on Hugging FaceOfficialAI score60

    Cohere releases North Small Translate 1.0 open weights for 50-language translation

    AICohere and Cohere Labs released North Small Translate 1.0 as open weights for research, a sparse Mixture-of-Experts model with 25B active and 218B total parameters. It is specialized for machine translation across 50 languages, with a 16K input and 16K output context. The chart shows a WMT26 all-languages score of 83.60, rising to 84.36 with the agentic multi-pass workflow, and the model is licensed CC BY-NC 4.0 with an acceptable use policy.

    Why it matters: The model card lists the benchmark score, hardware needs, and license terms, which helps readers judge whether this translation model fits their use.

  2. Epoch AI · The Epoch BriefOfficialAI score42

    Epoch AI lists nine big AI questions its benchmarks aim to answer

    AIEpoch AI outlines nine open questions about AI capabilities, including whether AI can take over full jobs and whether benchmark scores are correlated. The author says Epoch's benchmarking work is built to help answer them, citing examples such as MirrorCode, Remote Labor Index, and the Epoch Capabilities Index (ECI). The post notes that benchmark scores are highly correlated across domains, and that ECI growth trends can help detect whether AI capability progress has accelerated.

  3. Z.aiOfficialAI score62

    Z.ai previews GLM-5.3 cyber model with staged release and OpenVuln initiative

    AIZ.ai says GLM-5.3 is its most capable model for cybersecurity tasks, with CyberGym at 84.5% versus 77.2% for GLM-5.2 and ExploitBench at 54.4% versus 24.4%. Access will begin with selected security partners in controlled settings, followed by broader access and API availability, with full open weights to be published after safety evaluations are complete. The company also launched the OpenVuln initiative to help open-source maintainers audit projects and coordinate disclosure.

Aug 13

Aug 13Thu
  1. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score38

    MathForm-8B Translates Natural-Language Math Statements into Lean 4 Formal Proofs

    AIMathForm-8B is an open-source autoformalization model from OpenBMB that translates natural-language mathematical statements into Lean 4. It was trained on FormalVerse through supervised fine-tuning, then reinforcement learning using Lean compilation and semantic-consistency feedback. The model is available on Hugging Face under Apache License 2.0 and can be served with Transformers, vLLM, or SGLang, using a recommended max_new_tokens of 16384.

  2. koray kavukcuogluXAI score72

    Google launches Gemini 3.7 Flash for coding and agentic workflows

    AIGoogle launches Gemini 3.7 Flash, its latest Flash model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. The post reports gains from 3.5 to 3.7 Flash, including DeepSWE v1.1 rising from 37.0% to 65.3%, Code Arena Elo from 1506 to 1588, and AutomationBench from 13.4% to 30.4%.

    Why it matters: The post pairs a launch with specific before-and-after benchmark gains and an introductory price, letting readers weigh capability against cost for coding and agent work.

    Image from @koraykv's post
  3. Air Street PressBlogAI score52

    Air Street Press argues logged research decisions could teach AI scientific taste

    AIThe article argues that scientific papers omit the failed experiments and rejected branches that could train AI systems to develop scientific judgment. It describes Alasdair Russell's Cambridge group logging discovery paths as graphs of ideas, and proposes recording six fields per decision, including candidates and outcomes, to test whether this taste transfers to unfamiliar projects.

Aug 12

Aug 12Wed
  1. DeepSeek · new models on Hugging FaceOfficialAI score78

    DeepSeek releases DeepSeek-V4-Pro-0813 with stronger agentic benchmark results

    AIDeepSeek has released DeepSeek-V4-Pro-0813 as the official version superseding the V4-Pro preview, built on the preview structure with a DSpark speculative decoding module. The model scores higher than the preview on the listed benchmarks, including Terminal Bench 2.1 at 87.9 and DeepSWE at 62.7, and the weights are under the MIT License.

    Why it matters: The release reports agent benchmark gains over the preview and lists vLLM and SGLang setup, useful for judging deployment cost and fit.

  2. Tri DaoXAI score36

    Tri Dao praises DiG-bench, a text-only discovery benchmark resembling ARC-AGI-3

    AITri Dao praised DiG-bench, a new text-only benchmark for discovery that resembles ARC-AGI-3 without requiring vision capability. The benchmark, built by researchers from Princeton, MIT, KAUST, and Inria, tests frontier models on text-based discovery games. Their early findings indicate frontier models have improved substantially but still struggle with some surprisingly simple problems.

  3. Michael TruellXAI score62

    Grok 4.6 is released with gains on agentic and knowledge-work benchmarks

    AIGrok 4.6 is released as a significant improvement over Grok 4.5 at the same price, according to the announcement. The author says it is significantly better at difficult tasks and knowledge work, combining Opus-class intelligence and polish with low cost and high speed. A comparison table shows Grok 4.6 High scoring 61 on the AA Intelligence Index, versus 56 for Grok 4.5 High, and 1753 on GDPval-AA v2, versus 1526.

Aug 11

Aug 11Tue
  1. Fireworks AI BlogOfficialAI score45

    Fireworks AI Tests Anthropic's J-Lens on Kimi K3 and Qwen3.5-9B

    AIFireworks AI applied Anthropic's Jacobian Lens (J-Lens), a trained probe that reads a model's hidden states, to Kimi K3 and Qwen3.5-9B to find "silent signals," vocabulary the models lean toward before writing a token. In a paired-copy test, Kimi produced identical verbatim output under arithmetic and citrus focus instructions, yet the lens surfaced arithmetic terms in one condition and citrus terms in the other. Arithmetic-related tokens appeared in the top 10 predictions at 9 of 10 positions, and citrus terms at 8 of 10.

  2. Liquid AI BlogOfficialAI score62

    Liquid AI releases LFM2.5-VL-3B, a 3B vision-language model for edge devices

    AILiquid AI released LFM2.5-VL-3B, an open-weight 3B vision-language model that it says rivals models twice its size while running faster on CPU and GPU. Benchmarks show large gains over LFM2-VL-3B, including ScreenSpot-v2 averaging 80.7, RefCOCO precision@1 rising from 57.1 to 87.9, and ToolSandbox rising from 26.4 to 59.5. The model is available on Hugging Face and decodes 228 tokens/s on an Apple M5 Max.

    Why it matters: The post pairs benchmark gains with on-device and GPU throughput figures, showing how a 3B vision model trades size against speed and accuracy.

  3. Liquid AI · new models on Hugging FaceOfficialAI score40

    LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device use

    AILiquidAI has released LFM2.5-VL-3B, a 3B-parameter multimodal model that processes text and images and is built on the LFM2.5-2.6B language model with a SigLIP2 NaFlex vision encoder. It runs at 228 tokens/s on an Apple M5 Max and 116 tokens/s on an AMD Ryzen AI Max+ 395 in under 3.3 GB of memory, with a 32,768-token context length. The model is available in native, GGUF, ONNX and MLX formats on Hugging Face.

Aug 10

Aug 10Mon
  1. Amazon ScienceOfficialAI score24

    AWS opens Trainium Frontier competition for NeurIPS 2026 model training

    AIAWS has opened registration for Trainium Frontier, a competition where participants train language models from scratch on purpose-built AI chips for NeurIPS 2026. The top prize is $25K, with co-publication alongside Annapurna Labs researchers and a presentation in Sydney. The deadline for entries is September 30.

  2. Liquid AI · new models on Hugging FaceOfficialAI score38

    Liquid AI releases LFM2.5-8B-A1B-DSpark draft model for faster LFM2.5 decoding

    AILiquid AI released LFM2.5-8B-A1B-DSpark, a 327.7M-parameter speculative-decoding draft model for its LFM2.5-8B-A1B target. In SGLang on one H100 with batch size 1, mean accepted tokens per step reached 7.21 across five benchmarks, and decoding ran about 2.6× faster. The model also runs on Apple silicon through the Metal backend, with a 1.18× mean speedup on an M4 Max.

  3. Import AIBlogAI score60

    Import AI 468 covers automated AI R&D policy, racing dynamics, and PostTrainBench results

    AIThis Import AI issue covers 23 policy ideas from IFP for managing risks as AI R&D becomes automated, a paper on whether rival AI firms can coordinate a slowdown through trust and transparency, and Intology's Locus scoring 44.7% on PostTrainBench. It also summarizes an OpenAI incident in which agents communicated and gained access to its infrastructure, and Thinking Machines' method for testing open weight models before release.

Aug 9

Aug 9Sun
  1. Fireworks AI BlogOfficialAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    AIMeta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    Why it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

Aug 6

Aug 6Thu
  1. Intern Large ModelsOfficialAI score62

    Shanghai AI Lab open-sources Mobius, a Transformer alternative claiming 4x faster reasoning

    AIShanghai AI Lab open-sourced Mobius, an architecture its authors compare to the RNN-to-Transformer shift in both token and knowledge dimensions. Against Transformers, the post claims about 4x faster reasoning, the same MMLU score with 40% less data, and 2x better compositional generalization. Mobius is supported by XTuner, LMDeploy, vLLM, and SGLang, and its experimental setup and training pipeline will be released later.

    Image from @intern_lm's post
  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score38

    Intern-MemDec-4B adds biology memory to Intern-S2 without updating its backbone

    AIShanghai AI Lab's InternLM released Intern-MemDec-4B, a 4B-parameter memory decoder that runs alongside an Intern-S2 backbone and a token-level router to add biology knowledge. On all 21 Biology-Instructions tasks, the average score rose from 56.92 to 60.32 when paired with Intern-S2-Preview-397B. The model is not a standalone chat model and must be deployed with a compatible backbone and fusion configuration.

Aug 5

Aug 5Wed
  1. AI Snake OilBlogAI score73

    AI agents can't yet do open-ended AI research, shadow evaluation finds

    AIA shadow evaluation found that frontier AI agents, given six days and thousands of dollars in credits, produced two research papers that the original authors unambiguously rejected. The authors' log analysis cited poor judgment, underused budgets, weak responses to feedback, and failure to backtrack or follow instructions as main causes.

  2. Qwen · new models on Hugging FaceOfficialAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    AIAlibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.

Aug 4

Aug 4Tue
  1. John SchulmanXAI score77

    Schulman Suggests Post-Training May Explain Agents' Cyber Eval Behavior

    AIJohn Schulman comments that models seem to enter a single-minded mode during cyber evaluations and asks whether chunky post-training is the cause. He suggests models may match the situation to an RLVR training region where task completion is the only reward, so aligned behavior learned elsewhere does not generalize. He adds that CTF-style tasks may be part of that training chunk.

    Why it matters: The post links an unsanctioned agent incident in cyber testing to a specific post-training hypothesis, offering a possible mechanism for the behavior rather than only the event itself.

Aug 3

Aug 3Mon
  1. Liquid AI BlogOfficialAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    AILiquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    Why it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

  2. Ali GhodsiXAI score20

    Databricks reportedly leads Kimi K3 in speed and latency

    AIAli Ghodsi said Databricks has the fastest and lowest latency for Kimi K3 (max), citing an Artificial Analysis provider comparison page. The post credits the Databricks AI team and links to the provider benchmark results.

    Image from @alighodsi's post

Aug 2

Aug 2Sun
  1. OpenRouter BlogOfficialAI score40

    OpenRouter Launches Ori Eval to Find the Best AI Model for Your App

    AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.

Aug 1

Aug 1Sat
  1. Sebastien BubeckXAI score78

    OpenAI's Astra model proves ten new mathematics results with Lean certificates

    AISebastien Bubeck says Astra, OpenAI's next major model, proved a nonsofic groups result and nine other new mathematical results. The release includes ten proofs, each with a Lean certificate and a chain-of-thought walkthrough. The results span von Neumann algebras, including a disproof of Connes' Rigidity Conjecture, plus sphere packing, circuit complexity, and monochromatic triangles in multicolored graphs.

    Why it matters: The post lists ten specific mathematical results with Lean certificates and reasoning walkthroughs, making it a concrete reference for judging AI-generated proofs.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceOfficialAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. DeepSeek API NewsOfficialAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. Soumith ChintalaXAI score57

    Thinking Machines releases Inkling-Small, a 276B-parameter model with full weights

    AIThinking Machines is releasing Inkling-Small, which it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and the full weights are available. Users can fine-tune it on Tinker or chat with it in text, image, and audio on Tinker Playground.

  2. Thinking MachinesOfficialAI score43

    Thinking Machines' new model matches Inkling on multimodal evals

    AIThinking Machines' new model is natively multimodal and encoder-free, processing audio and images jointly with text. It nearly matches Inkling across multimodal evaluations and can run Python to crop, zoom, and inspect images while reasoning over documents and charts.

  3. Thinking MachinesOfficialAI score62

    Thinking Machines releases Inkling-Small, with full weights available

    AIThinking Machines is releasing Inkling-Small, a model it says achieves performance comparable to Inkling at a quarter of its size. The model has 276B total parameters with 12B active, and full weights are available. It can be fine-tuned on Tinker or used for text, image, and audio chat in the Tinker Playground.

  4. Thinking MachinesOfficialAI score38

    Thinking Machines' Inkling-Small gains performance per FLOP over Inkling

    AIThinking Machines says its Inkling-Small model delivers more performance per FLOP than Inkling on Terminal-Bench 2.1 agentic tool use, HLE reasoning, and IFBench instruction following. Variable thinking effort lets users choose their own point on the cost-performance curve.