Artificial Analysis compares six hallucination checkers on 20 shared tasks
Original titleWe compared six hallucination checkers on the same deliverables from a fixed subset of 20 tasks and eight models. We ran our two-stage ch...
AISummary
Artificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models.
GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet.
The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.
Source: Artificial Analysis · x.com