Skip to content
Trending storyDeveloping

Artificial Analysis compares Kimi K3 and Muse Spark 1.3 on criterion pass rate and hallucinations

2 articles1 sourceLast article 9h ago

Overview

AI overview

Artificial Analysis reports that Kimi K3 (max) completes task criteria at a 93.0% Criterion Pass Rate but averages 2.09 material hallucinations per task, while Muse Spark 1.3 (max) reaches 96.0% with 1.68 material hallucinations per task.

The publisher states that completing criteria and avoiding hallucinations are different skills, so the two measures are reported separately.

In a separate Artificial Analysis ranking of models with a Hallucination-Gated All-Pass Rate above 0%, four models sit on the Pareto frontier of score versus cost per task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) costs about $9.50 per task and Muse Spark 1.3 (max) about $4.20, while the three Claude models cost roughly $18 to $22 per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%.

AIWritten by AI from the articles below · overview updated Oct 8, 9:09 PM ET

Check the sources:

Developments

2 developments

  1. Oct 8, 2:33 PM ET · 1 article
    Artificial Analysis posts Pareto frontier of score vs. cost per task across models with Hallucination-Gated All-Pass Rate above 0%
    Artificial Analysis: Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.
  2. Oct 8, 2:33 PM ET · 1 article
    Artificial Analysis compares Kimi K3 and Muse Spark 1.3 on criterion pass rate and hallucinations
    Artificial Analysis: Completing the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Artificial Analysis
    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

  2. Artificial Analysis
    Completing the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.

    Completing the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.

Heat trend

Not enough continuous observations to show a trend yet.