Artificial Analysis compares Kimi K3 and Muse Spark 1.3 on criterion pass rate and hallucinations
Overview
Artificial Analysis reports that Kimi K3 (max) completes task criteria at a 93.0% Criterion Pass Rate but averages 2.09 material hallucinations per task, while Muse Spark 1.3 (max) reaches 96.0% with 1.68 material hallucinations per task.
The publisher states that completing criteria and avoiding hallucinations are different skills, so the two measures are reported separately.
In a separate Artificial Analysis ranking of models with a Hallucination-Gated All-Pass Rate above 0%, four models sit on the Pareto frontier of score versus cost per task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) costs about $9.50 per task and Muse Spark 1.3 (max) about $4.20, while the three Claude models cost roughly $18 to $22 per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%.
AIWritten by AI from the articles below · overview updated Oct 8, 9:09 PM ET
Check the sources:
Developments
2 developments
- Oct 8, 2:33 PM ET · 1 articleArtificial Analysis posts Pareto frontier of score vs. cost per task across models with Hallucination-Gated All-Pass Rate above 0%Artificial Analysis: Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.
- Oct 8, 2:33 PM ET · 1 articleArtificial Analysis compares Kimi K3 and Muse Spark 1.3 on criterion pass rate and hallucinationsArtificial Analysis: Completing the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Artificial AnalysisAmong models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.
Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.
- Artificial AnalysisCompleting the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.
Completing the criteria and not hallucinating are different skills. Kimi K3 (max) achieves a Criterion Pass Rate of 93.0%, but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) achieves 96.0% while averaging 1.68 material hallucinations per task.
Heat trend
Not enough continuous observations to show a trend yet.