Artificial Analysis compares six hallucination checkers, which flag hallucinations at different rates
Overview
Artificial Analysis compared six hallucination checkers on the same deliverables from a fixed subset of 20 tasks and eight models.
It reports that GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The firm says these counts reflect checker behavior and do not establish accuracy or rule out self-preference.
Separately, Artificial Analysis reports that accounting for hallucinations reshuffles model rankings. Muse Spark 1.3 (max) drops from 26.7% to 8.9%, leaving Grok 4.7 (xhigh) first on the headline metric at 9.4%. Kimi K3 (max) falls from 16.7% to 5.3%, Claude Sonnet 5.5 (max with fallback) from 11.7% to 2.8%, and GPT-6.1 Sol (max) declines least, from 7.5% to 6.9%.
Written by AI from the articles below · updated Oct 8, 9:01 PM ET
Check the sources:
Developments
2 developments
- Oct 8, 2:33 PM ET · 1 articleHallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse SparkArtificial Analysis: Hallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse Spark
- Oct 8, 2:33 PM ET · 1 articleArtificial Analysis compares six hallucination checkers on 20 tasks and eight modelsArtificial Analysis: Artificial Analysis compares six hallucination checkers on 20 shared tasks
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Artificial AnalysisHallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse Spark
AIOnce hallucinations are accounted for, Muse Spark 1.3 (max) drops from 26.7% to 8.9%, leaving Grok 4.7 (xhigh) first on the headline metric at 9.4%. Kimi K3 (max) falls from 16.7% to 5.3%, and Claude Sonnet 5.5 (max with fallback) falls from 11.7% to 2.8%. GPT-6.1 Sol (max) declines least, from 7.5% to 6.9%.
- Artificial AnalysisArtificial Analysis compares six hallucination checkers on 20 shared tasks
AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.
Heat trend
Not enough continuous observations to show a trend yet.