Skip to content
Trending storyDeveloping

Artificial Analysis compares six hallucination checkers, which flag hallucinations at different rates

2 articles1 sourcesince Oct 8Last article Yesterday ·

Overview

AISummary of 2 articles

Artificial Analysis compared six hallucination checkers on the same deliverables from a fixed subset of 20 tasks and eight models.

It reports that GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The firm says these counts reflect checker behavior and do not establish accuracy or rule out self-preference.

Separately, Artificial Analysis reports that accounting for hallucinations reshuffles model rankings. Muse Spark 1.3 (max) drops from 26.7% to 8.9%, leaving Grok 4.7 (xhigh) first on the headline metric at 9.4%. Kimi K3 (max) falls from 16.7% to 5.3%, Claude Sonnet 5.5 (max with fallback) from 11.7% to 2.8%, and GPT-6.1 Sol (max) declines least, from 7.5% to 6.9%.

Written by AI from the articles below · updated Oct 8, 9:01 PM ET

Check the sources:

Developments

2 developments

  1. Oct 8, 2:33 PM ET · 1 article
    Hallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse Spark
    Artificial Analysis: Hallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse Spark
  2. Oct 8, 2:33 PM ET · 1 article
    Artificial Analysis compares six hallucination checkers on 20 tasks and eight models
    Artificial Analysis: Artificial Analysis compares six hallucination checkers on 20 shared tasks

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Artificial Analysis
    Hallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse Spark

    AIOnce hallucinations are accounted for, Muse Spark 1.3 (max) drops from 26.7% to 8.9%, leaving Grok 4.7 (xhigh) first on the headline metric at 9.4%. Kimi K3 (max) falls from 16.7% to 5.3%, and Claude Sonnet 5.5 (max with fallback) falls from 11.7% to 2.8%. GPT-6.1 Sol (max) declines least, from 7.5% to 6.9%.

  2. Artificial Analysis
    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

Heat trend

Not enough continuous observations to show a trend yet.