Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Sep 30

Sep 30Wed
  1. Liquid AIOfficialAI score42

    LongevityBench: Liquid AI's compact LFMs beat frontier models on aging tasks

    AILiquid AI and InSilicoMeds released LongevityBench, an aging benchmark with 17 tasks spanning clinical records, DNA methylation, transcriptomics, proteomics, and genetics. On several tasks, Liquid AI's compact LFMs outperformed every frontier model the team evaluated. The team plans to present the work to the longevity research community at ARDD this week.

    Video from @liquidai's post
  2. Tencent HyOfficialAI score62

    Tencent Hunyuan releases ExplorationBench to test how AI systems discover rules

    AIResearchers from Tencent Hy, Fudan University, and Tsinghua University released ExplorationBench, a benchmark that tests whether AI systems can discover hidden rules in executable Alien World sandboxes. Across 10 frontier systems, getting feedback from experiments outperformed thinking alone, with the best run reaching 89.0% after four rounds. The authors note that rankings barely transfer between the two worlds, and the code is listed as coming soon.

    Image from @TencentHunyuan's post
  3. ModelScopeOfficialAI score46

    IndexTeam releases Index-Translate multilingual translation model family

    AIIndexTeam has released Index-Translate, a multilingual family covering text, speech, dubbing, and long-document translation across 150 languages. Its 9B model scores 0.8789 on FLORES, 0.8209 on instTrans, and 0.7387 on MEME, and the 2B and 9B models are released under Apache 2.0.

    Video from @ModelScope2022's post
  4. The SequenceBlogAI score50

    The Sequence Learning Loop: Opus 5.5, DeepSeek Environments, and Claude's DNA Discovery

    AIIssue 942 of The Sequence links Anthropic's Claude Opus 5.5, reported for the week of September 21–27, to DeepSeek's September 19 environments paper and a report of AI-assisted biological discovery. The newsletter argues that progress increasingly depends on the surrounding machinery that governs where a model acts, what it observes, and how its conclusions are checked.

  5. Hamel HusainBlogAI score42

    Hamel Husain Tests Anthropic's Claude Eval Plugin on Leasing Assistant Traces

    AIHamel Husain reviewed Anthropic's new build_eval and hill-climb commands in the claude-api plugin for Claude Code, finding it useful for discovering issues like human handoff, formatting, and voice agent problems. He criticized it for pushing evaluator creation before data review, asking for label validation in Markdown files, and bundling four failure checks into one broad call-transfer evaluator. Husain says he would hold off on using it for now.

  6. ModelScopeOfficialAI score62

    InSpatio-World 1.5 turns images and videos into real-time explorable 4D worlds

    AIInSpatio-World 1.5 from InSpatio_AI turns a single image, four images, a panorama, or a video into a navigable scene with wide viewpoint changes. The 1.3B model scores 68.72 on WorldScore-Dynamic, ranking first among evaluated real-time and interactive methods, with speeds up to 24 FPS. The post says the code is released under Apache 2.0 and that dependencies keep their own licenses.

    Video from @ModelScope2022's post
  7. Artificial Analysis ArticlesOfficialAI score39

    Upstage Releases Solar Mini 4 Reasoning Model, Scoring 24 on Intelligence Index

    AIKorean AI lab Upstage has released Solar Mini 4, a proprietary reasoning model that scores 24 on the Artificial Analysis Intelligence Index with 35B total and 3B active parameters. It is priced at $0.10/$0.40 per 1M input/output tokens and has a 1M-token context window, but averages 7.1 minutes per task due to heavy output token use. Its weights are not released, and its size cannot be independently verified.

  8. Artificial Analysis ArticlesOfficialAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    AIArtificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    Why it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.

Sep 29

Sep 29Tue
  1. Jerry LiuXAI score20

    Jev, a System One model, tops OSS rivals on document tasks

    AIJerry Liu says Jev, a System One model, outperformed other open-source classifiers and document-specific models on orientation detection, language detection, classification, and splitting. The benchmark measured accuracy, cost, and latency across these fast document decisions, with Jev leading most comparisons. The benchmark code is available in the run-llama/jev_vs_oss repository.

    Video from @jerryjliu0's post
  2. Jerry LiuXAI score22

    GPT-6.1 Sol Improves Table Parsing and Reading Order in OCR Benchmarks

    AIJerry Liu benchmarked gpt-6.1 sol on document OCR tasks and found a sizable increase in table parsing and reading order over gpt-6 sol from a week earlier. Its table parsing is similar to gpt-6 astra. He noted frontier models still cost roughly an order of magnitude more than cost-effective document parsing solutions, leaving room to improve the premium end above 1c per page.

    Image from @jerryjliu0's post
  3. Hugging Face BlogOfficialAI score46

    Open TTS Leaderboard ranks multilingual and voice cloning models using objective metrics

    AIHugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.

  4. Apple Machine Learning ResearchOfficialAI score38

    LLM Conditioning Study Finds Steering Methods Trade Fluency for Effectiveness

    AIApple researchers systematically tested LLM conditioning methods and found efficient activation steering often degrades fluency. Steering is far less effective on instruction-tuned models than base models, while prompting and full supervised fine-tuning work for concept injection but are weaker at concept removal. Cheap textual metrics correlate highly with costly LLM-as-judge scores.

  5. Prime IntellectOfficialAI score20

    Prime Intellect to deploy on NVIDIA Vera CPU for agentic workloads

    AIPrime Intellect says it will be among the first to deploy on NVIDIA's new Vera CPU, after earlier access to benchmark sandboxes on Vera in March. The company plans to use Vera's dynamic memory latency for agentic workloads, which it describes as ideal for Prime Sandboxes.

    Image from @PrimeIntellect's post
  6. Liquid AIOfficialAI score32

    Liquid AI launches d1, first decision model, beating Jev on HF index

    AILiquid AI announced d1, its first decision model, which it says is the first to outperform Jev on Hugging Face's Decision Index. The company claims d1 wins on multilingual evals, resists prompt injection better, handles longer inputs more effectively, and is built for fast, structured decision-making in software environments. It is available via the Liquid API at console.liquid.ai, with OpenRouter availability coming soon.

    Image from @liquidai's post
  7. Noam BrownXAI score25

    OpenAI's Noam Brown says AI evals should measure intelligence against cost

    AINoam Brown praised OpenAI for presenting model evaluations as intelligence plotted against cost, arguing that cost should be part of how intelligence is measured. OpenAI's linked context says GPT-6.1 Sol delivers near-Astra intelligence at one-fifth the price and is the most cost-efficient model for its performance available today.

    Image from @polynoamial's post
  8. BAAI · new models on Hugging FaceOfficialAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

  9. CognitionOfficialAI score40

    GPT-6.1 Sol Now Available in Devin at Lower Cost

    AIGPT-6.1 Sol is now available in Devin, scoring 60.4% on FrontierCode 1.1, close to GPT-6 Sol's 60.7%. At medium reasoning effort it costs $0.31 per task, 81% less than GPT-6 Sol at max effort.

    Image from @cognition's post
  10. OpenAIOfficialAI score37

    GPT-6.1 Sol improves alignment and transparency over GPT-6 Sol

    AIOpenAI reports that GPT-6.1 Sol shows major alignment improvements over GPT-6 Sol in its evaluations, moving closer to GPT-6 Astra. The model is more transparent about its limitations and more reliable at respecting user intent and safety constraints.

    Image from @OpenAI's post
  11. Jerry LiuXAI score22

    Jerry Liu and Snorkel's Vincent Sun discuss evals and RL environments

    AIJerry Liu hosted a dinner with Snorkel's Vincent Sun on evals and RL environments, a topic shaped by models rapidly saturating benchmarks. The conversation highlighted that building fair RL environments is hard, since failures are difficult to attribute to input, harness, or reward model, and that long-horizon evals spanning weeks or months remain very difficult. The post also noted that regulated industries still require human-in-the-loop review because 80% accuracy is not sufficient.

    Image from @jerryjliu0's post
  12. Replit BlogOfficialAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

  13. ModelScopeOfficialAI score54

    IQuest-Q1 released as 320B MoE model for long-horizon coding agents

    AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.

    Image from @ModelScope2022's post
  14. ModelScopeOfficialAI score44

    Intern-Decision multimodal models scale structured decisions at 0.8B–4B

    AIShanghai AI Laboratory's Intern-Decision family of 0.8B, 2B, and 4B multimodal models averages 79.38, 84.68, and 90.02 across seven decision benchmarks. Intern-Decision-4B scores 88.74, surpassing Jev while achieving better probability calibration. Reported mean latency is 33.98, 33.28, and 44.16 ms, versus 109.70 ms for Jev in the same local HF setup.

    Image from @ModelScope2022's post
  15. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score40

    InternLM releases AdvancedMathBench-AutoVerifier to grade natural-language math proofs

    AIInternLM's AutoVerifier, built on Qwen3_5MoeForConditionalGeneration with about 68 GiB of weights across 40 safetensors shards, evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench, which accepts a proof only when all eight judgments report -1. The model is a learned grader rather than a formal proof checker and can make errors.

  16. Thomas WolfXAI score29

    Thomas Wolf calls a post simply "impressive"

    AIThomas Wolf, owner of the Hugging Face account, posted the single word "impressive" in response to a quoted post. The quoted post reports a new NanoGPT training record of 39.9s, down 27.7s from the prior 67.6s, achieved through per-flop optimizations such as sampled softmax and sparse updates.

  17. Matei ZahariaXAI score36

    Matei Zaharia says autoresearch results are going into serving stack

    AIMatei Zaharia said autoresearch produced strong results that are being integrated into a model serving stack. The post gives no specific figures, benchmarks, or product names. Background context from a related post says Databricks ranked #1 on NVIDIA's SOL-ExecBench kernel leaderboard across all four tracks using agents.

  18. Artificial Analysis ArticlesOfficialAI score62

    Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents

    AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.

    Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.

  19. Artificial Analysis ArticlesOfficialAI score78

    GPT-6.1 Sol replaces GPT-6 Sol with near-Astra intelligence at lower cost

    AIArtificial Analysis reports that GPT-6.1 Sol replaces GPT-6 Sol after seven days and scores 1 point below GPT-6 Astra on the Intelligence Index. At max effort it costs $0.72 per Intelligence Index task, compared with $3.26 for GPT-6 Astra and $1.05 for GPT-6 Sol. Its pricing matches GPT-6 Sol at $2/$10 per million input/output tokens, but it uses about 10-30% more output tokens.

    Why it matters: The source compares GPT-6.1 Sol against GPT-6 Sol, GPT-5.6 Sol, and GPT-6 Astra on cost per task and token use, helping readers weigh performance against price.

  20. Anthropic ResearchOfficialAI score80

    Anthropic says GLM-5.3 gives attackers cyber capabilities with weak safeguards

    AIAnthropic reports that Zhipu AI's GLM-5.3 can autonomously build end-to-end cyber exploits and is released without meaningful safeguards against misuse. In its simulated tests, attackers bypassed the model's safeguards 64% to 100% of the time using simple techniques, while the same attacks failed against safeguarded Claude models. Anthropic also cites an NIST CAISI assessment calling GLM-5.3 the most cyber-capable open-weight model released to date.

    Why it matters: The report shows how open-weight safeguards fail under simple bypasses, offering concrete test figures for judging misuse risk in released models.

Sep 28

Sep 28Mon
  1. ModelScopeOfficialAI score44

    Audio8 ASR Infinite enables unlimited-length streaming speech transcription with bounded memory

    AIAudio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.

    Video from @ModelScope2022's post
  2. Alexander DoriaXAI score14

    Document parsing favors large models: Astra annotates, Gemma 4 31B finetunes

    AIAlexander Doria says high parameter capacity still matters for harder document processing, running Astra for initial annotation and Gemma 4 31B for finetuning. Yifei Hu reports that gpt-6-sol improved over last week's version on domain-specific document parsing but remains far behind gpt-6-astra, with the benchmark itself built using Astra.

  3. Ali GhodsiXAI score62

    Databricks finds Opus 5.5 cheaper and better, GPT-6 Luna 20x cheaper per task

    AIDatabricks tested recent AI models across 2,400 engineers and found Opus 5.5 offers the highest quality mid-tier performance, with about 20% lower same-task costs than Opus 4.8. The company is now encouraging Opus 5.5 as a default model for coding, and reports that GPT-6 Luna is at least 20 times cheaper per task than Opus 5.5, roughly matching Opus 4.6 on one difficult evaluation suite. The Luna findings are preliminary.

  4. ReplicateOfficialAI score28

    Pruna's P-Video-2-Pro video model now runs on Replicate

    AIReplicate has added P-Video-2-Pro, the latest video model from Pruna AI, which sits on the edge of the preference-speed and preference-price Pareto frontiers. Design Arena ranks its Quality and Speed variants tied for #2 on the Image to Video leaderboard with an Elo of 1325, with the Quality version generating in 8.0 seconds and the Speed version in 4.5 seconds.

  5. Hacker News · Launch HN, YC launches (10+ points)BlogAI score54

    Vespper launches a DOCX MCP for agents editing Word documents

    AIVespper, a Y Combinator F24 startup, launches a DOCX MCP that lets agents edit Word files through HTML that a trained reconciler converts back to OOXML. On its 279-task internal benchmark, the company reports its MCP is 2.7–2.9x cheaper and 2.7–3.5x faster than Anthropic's DOCX skill, with higher pass rates. The post also lists current limitations, including no comment creation or reply, no image or video attachment, and no support for latent styles.

  6. FireworksOfficialAI score22

    Normal Factory's CAD Arena joins the Specialized Intelligence Index

    AINormal Factory joins the Specialized Intelligence Index with CAD Arena, which tests whether AI agents can turn engineering drawings into accurate, editable CAD parts. The benchmark evaluates agents across five CAD platforms, extending the SII into engineering design.

    Image from @FireworksAI_HQ's post