Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%
Overview
Harvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score.
GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.
Written by AI from one article, by Artificial Analysis Articles
Check the sources:
Developments
2 developments
- Oct 8, 2:33 PM ET · 1 articleArtificial Analysis launches Harvey LAB-AA legal agent evaluation built with HarveyArtificial Analysis: Harvey LAB-AA: Artificial Analysis benchmark for legal AI agents
- Oct 8, 12:00 AM ET · 2 articlesHarvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%Artificial Analysis Articles: Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Sherwin WuHarvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings
AIArtificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination. The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).
- Artificial AnalysisHarvey LAB-AA: Artificial Analysis benchmark for legal AI agents
AIArtificial Analysis has released Harvey LAB-AA, an evaluation built on Harvey's LAB dataset and developed in collaboration with Harvey. Full results are published on the Artificial Analysis evaluations page, alongside Harvey's commentary on the benchmark and human expert preferences.
- Artificial Analysis ArticlesHarvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark
AIHarvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.
Heat trend
Not enough continuous observations to show a trend yet.