Skip to content

#Eval/Benchmark

Oct 8

TodayOct 8Thu113 items
  1. PandailyAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    ByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. Tencent HunyuanAI score63

    Tencent Hunyuan releases ExplorationBench to test AI rule discovery

    Tencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that tests whether AI systems can discover rules through experiments in verifiable alien worlds. Across 10 frontier systems, feedback from experiments raised the best AlienCode score to 89.0% after four rounds, while closed-book runs without feedback stayed at 0.5–11.0%.

  3. IThome · AI (IT之家)AI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    OpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  4. PandailyAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    Shanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  5. Leiphone (雷峰网)AI score46

    IROS 2026 papers show AI reintegrating with classical robotics rather than replacing it

    Of 1,933 IROS 2026 papers, Robot Learning/Embodied AI appears in about 809, while Navigation/Planning covers 564 and Perception/Vision 556. The article argues large models are being embedded into traditional planning, geometry, and control rather than replacing them. Vision-language-action models are shifting toward efficiency, 3D understanding, memory, and system integration.

  6. Leiphone (雷峰网)AI score46

    IROS 2026 Best Paper goes to LT-Mem robot long-term memory study

    At IROS 2026 in Pittsburgh, the Best Paper Award went to Yumin Lee, Hyoseok Ju and Giseop Kim for LT-Mem, a volatility-aware spatio-temporal memory system for lifelong robot scene understanding. The Best Student Paper Award went to Pei-An Hsieh and colleagues for flatness-preserving residual learning enabling real-time tight quadrotor formation flight. Other honors included a humanoid tennis-skills paper and SteadyTray, a humanoid tray-transport study.

  7. QbitAI (量子位)AI score52

    Claude Haiku 5.5 launches with higher benchmark scores and new migration requirements

    Anthropic released Claude Haiku 5.5, which the article says outperforms DeepSeek V4.1 Flash and GLM-5.3-Flash on official benchmarks and matches GPT-6 Luna on price. On OSWorld 2.1, its Low effort tier scores 42.0% at $0.07 per task, versus 15.7% at $1.45 for Haiku 4.5 at Max. Migrating from Haiku 4.5 requires changes to thinking configuration, sampling parameters, assistant prefill, and the computer-use tool version.

  8. Elvis SaraviaAI score55

    HERMES harness lifts GPT-5.6 Sol repository migration from 6.5% to 31.0%

    A paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis. With the same model and effort setting, GPT-5.6 Sol's whole-repository migration score rose from 6.5% to 31.0% when Codex was replaced by HERMES. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4 points on average, and Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.

  9. MarkTechPostAI score48

    Laya Open-Source Decision Engine Tutorial: Zero-Shot Decisions and Calibration

    Laya is a 421-million-parameter non-autoregressive decision engine from Convai Innovations that returns calibrated option probabilities in a single forward pass with zero output tokens. This tutorial tests its zero-shot accuracy, probability calibration, temperature fitting, and abstention gating on the CLINC150 banking intent dataset.

  10. TechCrunch · AIAI score62

    Common Sense Media rates ChatGPT for Teens an unacceptable risk over engagement design

    Common Sense Media labeled ChatGPT for Teens an "unacceptable risk," finding its design still encourages engagement even in crisis situations. The report says the teen version failed to meet commitments on three of five severe harms, and that break reminders appeared only twice across nearly 2,000 prompts. OpenAI disputed the methodology, saying the testing may have ended before parental controls were fully active, and cited its own data showing teens average under 15 minutes a day.

  11. Xiaomi MiMoAI score44

    Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support

    Xiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.

  12. Jerry LiuAI score38

    LightOn OCR-3 now on OpenDocRouter, near Gemini 3.8 Flash at lower cost

    LightOn OCR-3 is now available on OpenDocRouter at $0.28 per 1M input tokens and $1.40 per 1M output tokens, about $3.19 per 1k pages on ParseBench. On ParseBench, the author says it sits on the Pareto frontier for open-weight OCR models, with performance similar to Gemini 3.8 Flash low at roughly 45% lower price. It is described as decent at tables, workable for charts, and quite good at grounding.

  13. Artificial AnalysisAI score18

    Check out Grok Imagine Video 1.5 Lite for yourself on the AA-Video Leaderboards: AA-Video-T2V v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video AA-Video-T2V-Silent v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video?audio-output=false Or vote in the Video Arena: https://artificialanalysis.ai/video/arena

    Check out Grok Imagine Video 1.5 Lite for yourself on the AA-Video Leaderboards: AA-Video-T2V v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video AA-Video-T2V-Silent v2.0: https://artificialanalysis.ai/video/leaderboard/text-to-video?audio-output=false Or vote in the Video Arena: https://artificialanalysis.ai/video/arena

  14. Artificial AnalysisAI score7

    Artificial Analysis publishes AA-Video-T2V v2.0 prompt for snowy cabin scene

    Artificial Analysis shares the second part of an AA-Video-T2V v2.0 prompt describing a four-shot documentary-style handheld video of a glass cabin in falling snow. The shots follow a caretaker sweeping snow from the deck, empty snow-covered windows, an empty interior, and the same caretaker stamping snow off his boots at the door, with hard cuts between shots.

  15. Artificial AnalysisAI score38

    Grok Imagine Video 1.5 Lite leads in architecture, consumer, and knowledge-work use cases

    Artificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier in Architecture & Real Estate, Consumer, and Productivity & Knowledge Work use cases. It sits furthest from the frontier in Live-Action Film and Frontier use cases. Against Grok Imagine Video 1.5, Lite matches it in Social Media & Creator Content and trails it on the other nine use cases.

  16. Artificial AnalysisAI score29

    Grok Imagine Video 1.5 Lite nears frontier on three AA-Video-T2V capabilities

    Artificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier on AA-Video-T2V v2.0 in Multi-Scene & Narrative, Lighting & Materials, and Text Rendering. It is furthest behind in Dialogue & Lip Sync and Human Anatomy. Compared with Grok Imagine Video 1.5, Lite matches it in Physics and trails on the other nine capabilities, by the least in Multi-Scene & Narrative.

  17. Artificial AnalysisAI score31

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

    Grok Imagine Video 1.5 Lite sits on the quality and speed frontier on AA-Video-T2V-Silent v2.0 Among the 12 models on AA-Video-T2V-Silent v2.0 that we benchmark for generation speed, no model is both faster and higher quality than Grok Imagine Video 1.5 Lite. It generates a 10 second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher and takes 94 seconds for a 5 second clip. Vidu Q3 Turbo is 9 seconds faster on a 5 second 720p clip, and scores well below it.

  18. Artificial AnalysisAI score46

    Grok Imagine Video 1.5 Lite ranks ahead of Google's Veo 3.1 on AA-Video-T2V v2.0, at about a third of the price At 1080p with audio, Grok Imagine Video 1.5 Lite costs $0.14 per second, against $0.40 per second for Veo 3.1. It ranks #17 on AA-Video-T2V v2.0, two places above Veo 3.1. Against Grok Imagine Video 1.5, Lite costs 44% less at 1080p and ranks six places lower.

    Grok Imagine Video 1.5 Lite ranks ahead of Google's Veo 3.1 on AA-Video-T2V v2.0, at about a third of the price At 1080p with audio, Grok Imagine Video 1.5 Lite costs $0.14 per second, against $0.40 per second for Veo 3.1. It ranks #17 on AA-Video-T2V v2.0, two places above Veo 3.1. Against Grok Imagine Video 1.5, Lite costs 44% less at 1080p and ranks six places lower.

  19. Artificial AnalysisAI score42

    Grok Imagine Video 1.5 Lite ranks #17 in video arena at lower cost

    SpaceXAI's Grok Imagine Video 1.5 Lite ranks #17 on both AA-Video-T2V v2.0 leaderboards, ahead of Google's Veo 3.1 at about a third of its price. It is the fastest model at its quality level in Artificial Analysis benchmarks, with a median of 60.5 seconds for a 10-second 1080p clip, and it costs $0.14 per second at 1080p, 56% of Grok Imagine Video 1.5's $0.25 per second.

  20. Epoch AIAI score31

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

  21. Sara HookerAI score10

    We just released a nice blog post on adaptive checklists. The idea of checklists isn't new. However, checklists often fail to capture the edge of current capabilities or what domain specific requirements. Our @adaption_ai checklists do both, without requiring supervision.

    We just released a nice blog post on adaptive checklists. The idea of checklists isn't new. However, checklists often fail to capture the edge of current capabilities or what domain specific requirements. Our @adaption_ai checklists do both, without requiring supervision.

  22. SiliconANGLE · AIAI score24

    CoreWeave Pitches Open Full-Stack AI Cloud With Forge Development Platform

    CoreWeave is positioning its AI cloud around an open development loop, connecting training, inference and evaluation through its newly announced CoreWeave Forge platform. Chief marketing officer Jean English said the company wants production learnings to improve models and agents and that the loop should work across different models, frameworks and clouds. She argued that competitive differentiation extends beyond GPUs to partner tooling, infrastructure and APIs.

  23. Artificial AnalysisAI score22

    The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current Alliance members are @CollinearAI, @IBM, @nvidia, and @vercel. Partners contribute expert input on the design and implementation of the Index, and may contribute datasets and external research directly. Organizations interested in joining the Cyber Index Alliance can contact us at cyber@artificialanalysis.ai

    The Artificial Analysis Cyber Index Alliance brings together industry partners to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current Alliance members are @CollinearAI, @IBM, @nvidia, and @vercel. Partners contribute expert input on the design and implementation of the Index, and may contribute datasets and external research directly. Organizations interested in joining the Cyber Index Alliance can contact us at cyber@artificialanalysis.ai

  24. Artificial AnalysisAI score38

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

    GPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model, achieving the #1 spot on the Index. Compared to the publicly available GPT-6 Sol, the Daybreak Blue model has the largest gains on CyberGym-E2E, which is the benchmark where we observe the most safety refusals

  25. Artificial AnalysisAI score46

    GPT-6 Sol (Daybreak Blue, max) takes the top position on the Artificial Analysis Cyber Index at much lower cost than other leading models. At a Cost per Task of $1.77, it is significantly more cost-effective than other leading models, including Grok 4.7 (xhigh) which costs $11.67 per task

    GPT-6 Sol (Daybreak Blue, max) takes the top position on the Artificial Analysis Cyber Index at much lower cost than other leading models. At a Cost per Task of $1.77, it is significantly more cost-effective than other leading models, including Grok 4.7 (xhigh) which costs $11.67 per task

  26. Artificial AnalysisAI score62

    GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index

    Artificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now ranks first. The model is available only through OpenAI's Daybreak program and records no safety blocks across the Index. Its overall score is 32 points higher than the publicly available GPT-6 Sol (max), at a cost of $1.77 per task versus $11.67 for Grok 4.7 (xhigh).