Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 26

Sep 26Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    AIShanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

  3. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. LMSYS OrgAI score38

    SGLang adds multi-item scoring for faster decision model serving

    AISGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

Sep 24

Sep 24Thu
  1. ModelScopeAI score23

    NeoHorse-Jev-4B open model turns app states into structured decisions

    AIModelScope has released NeoHorse-Jev-4B, a compact open model that converts application states into structured decisions and probabilities. It scores 77.70 across six text decision benchmark groups, ranking first among four open-weight models with complete results in the comparison. Its prefill-only inference supports Choice, Noul, and Score primitives, accepts text or a single image with text, and is available under Apache 2.0 for deployment via vLLM, SGLang, Python, CLI, or HTTP.

  2. vLLMAI score42

    TileRT and vLLM hit 469 tok/s on GLM-5.3 with MI355X

    AIThe TileRT and AMD teams reached 469 tok/s single-user decode for GLM-5.3 on 8× MI355X using vLLM. The setup disaggregates work, with vLLM handling prefill and TileRT handling latency-critical decode through vLLM's V1 connector interface. SemiAnalysis's AgentX benchmark reports the configuration at 470 TPS on GLM 5.3 (FP8), over 40% faster than GB300 TRTLLM using FP4.

  3. Google ResearchAI score60

    Google Research details four agentic frameworks for coherent long-form video generation

    AIGoogle Research introduces four multi-agent frameworks for generating minutes-long videos with consistent characters and environments across shots. The frameworks include AI video co-director, CANVAS, A²RD, and VQQA, which are built as orchestration layers on Gemini and Veo and use SynthID watermarking. The post reports measured gains on benchmarks such as GenAD-Bench, HardContinuityBench, and LVBench-C, with the full architectures described in the linked papers.

    Why it matters: The post links four frameworks to specific failure modes in long video generation, such as semantic drift and cascading errors, making the design choices easier to compare.

  4. Goodfire ResearchAI score48

    Steering Along Manifolds Beats Linear Steering for Controlling Llama's Days-of-Week Behavior

    AIGoodfire Research shows that steering Llama-3.1 8B along the curved representation manifold of weekdays produces output probabilities that follow the model's natural cyclic behavior, shifting probability mass smoothly from Monday to Tuesday to Friday. Linear steering along a straight vector, by contrast, cuts across the behavior manifold and yields noisy off-target tokens, some not days of the week at all. The authors argue that representation geometry and behavior geometry are linked bidirectionally.

  5. LangChain BlogAI score50

    LangSmith Engine v2 adds red teaming and pre-validated agent fixes

    AILangChain released LangSmith Engine v2, an in-platform agent that scans production traces to detect agent issues and validates proposed fixes before human review. Engine v2 adds Red Teaming, currently in Private Beta for LangSmith Deployment users, which tests agents for weaknesses such as hallucinations and system-prompt violations before they reach production. Engine v2 is available in SaaS deployments for LangSmith Plus and Enterprise plans, with Self-Hosted support and BYOK for Engine coming later.

  6. Anthropic ResearchAI score60

    Anthropic study finds Claude agent trading limited by preference understanding

    AIAnthropic ran a controlled book-swapping market with 201 employees and Claude-powered agents, which reached 0.55 efficiency against a 0.89 optimum. Agents matched participants' own rankings on 61% of book pairs, and about 85% of the shortfall came from imprecise preference representation rather than the trading floor design. Stronger models produced more efficient markets than weaker ones, while instructions mattered less.

    Why it matters: The study separates agent misunderstanding of user preferences from negotiation failure, showing which failure mode limits outcomes in agent-run markets.

  7. LangChain BlogAI score44

    LangSmith Launches Trajectories for Readable, Chronological Agent Session Views

    AILangChain has launched Trajectories in LangSmith, a chronological, conversational view that aggregates human, AI, and tool messages across an agent and its subagents. Trajectories work with traces from LangChain, LangGraph, Deep Agents, OpenAI and Claude agent SDKs, and coding agents like Codex, Claude Code, and Cursor. The feature is available now on all plans in the US.

Sep 23

Sep 23Wed
  1. Philipp SchmidAI score62

    Gemini 3.8 Flash TTS guide shows how to create and reuse your own voice

    AIGemini 3.8 Flash TTS and Flash-Lite TTS are now available in the Gemini API and AI Studio, with a new feature to replicate your own voice or create one from a sentence. The guide shows recording two clips, one of 15-20 seconds of natural speech and one reading a required consent sentence, then creating a reusable voice ID. It also explains that input text is now spoken word for word, so delivery belongs in speech_metadata.style and short sounds inline.

  2. Dario AmodeiAI score76

    Claude Helps Discover a Possible New Gene Editing Enzyme System

    AIAnthropic announced that Claude, working mostly on its own, identified a previously unknown enzyme system in bacteriophage DNA that may represent a new gene editing mechanism. Claude read literature and genome data, proposed experiments, and Anthropic's team carried them out. The function and biotechnological utility of the system remain unclear.

    Why it matters: The post pairs a Claude-led discovery with the lab workflow used to verify it, showing how AI and humans split the research work in biology.

  3. Black Forest LabsAI score62

    Black Forest Labs releases FLUX 3 Action, a robot policy model

    AIBlack Forest Labs says its FLUX 3 Action, a single-step 7B checkpoint, outperforms every other open policy on RoboLab. It processes each second of robot motion 1.45× to 1.66× faster than Pi0.5, and uses a 2.13-second action horizon versus Pi0.5's 1 second. The company adds that its guidance-distilled checkpoint raises the state-of-the-art RoboLab success rate while running 2.85× to 3.15× faster than the previous leading open WAM.

  4. Black Forest LabsAI score67

    Black Forest Labs releases FLUX 3 Action, an open 7B world action model for robots

    AIBlack Forest Labs says FLUX 3 Action is an open-weights 7B world action model that ranks first on the RoboLab benchmark. The company says it outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. The model predicts video and actions together, and the company is releasing the weights, code, fine-tuning recipe, benchmarks, and examples. It also integrated the model into Hugging Face's LeRobot with NVIDIA, with edge deployment on NVIDIA Jetson.

    Why it matters: The release pairs benchmark results with the trade-off it claims to remove between world action model performance and VLA speed, which is useful context for robotics teams weighing open models.

  5. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.

  6. Mike KnoopAI score57

    Tufa Labs reaches 83.06% on ARC-AGI-2, 2% short of the grand prize

    AIMike Knoop says the top ARC Prize 2026 ARC-AGI-2 score of 83.06% by Tufa Labs is only 2% short of the 85% grand prize threshold. The challenge runs under strict Kaggle compute limits with no internet access, and the winning solution is set to be open sourced. The image shows the leaderboard with RabbitHole at 76.94%, nvbanana at 74.17%, Yi-Chia Chen at 55.14%, and Kha Vo at 37.50%.

  7. Philipp SchmidAI score62

    Gemini 3.8 Flash TTS adds voice replication and prompt-designed voices

    AIGoogle launched Gemini 3.8 Flash TTS and Flash-Lite TTS, letting users replicate a voice from 30 seconds of audio or design one from a text description. The post claims #1 on Hume's Voice Design Benchmark and top placement in Voice Arena across 6 languages. Voices can be directed line by line with style and inline tags, with a consent check on replication and SynthID on every clip. Availability is through the Gemini API and Google AI Studio.

  8. QwenAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

  9. ModelScopeAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

  10. ModelScopeAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

  11. KrASIA · Big TechAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

  12. Mike KnoopAI score25

    Formal verification gains ground, but human understanding remains an alignment gap

    AIMike Knoop argues that formal verification is becoming feasible and is important for security. He adds that it does not automatically build human understanding, which he calls an even bigger alignment problem. The post is framed as a reply to Boris Cherny's report that Claude Opus 5.5 helped formally verify the Claude Agent SDK in Lean, producing 16 bug-fix PRs.

Sep 22

Sep 22Tue
  1. TinkerAI score25

    Tinker fine-tunes Qwen3.6 for Jev-style probability prompts in 10 minutes

    AITinker says an open LLM can serve a Jev-like interface that takes discrete options and returns fast probabilities, since next-token prediction is already a probabilistic classifier. A post by @ekzhang1 reports that a $5, 10-minute supervised fine-tuning run on Tinker improved Qwen3.6-35B-A3B's handling of Jev-style prompts, with +8% on GPQA Diamond and +12% on MMLU-Pro.

  2. Redwood Research BlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.

  3. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.