Skip to content
Trending storyDeveloping

Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%

3 articles3 sourcessince Oct 8Last article Yesterday ·

Overview

AISummary of one article

Harvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score.

GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Written by AI from one article, by Artificial Analysis Articles

Check the sources:

Developments

2 developments

  1. Oct 8, 2:33 PM ET · 1 article
    Artificial Analysis launches Harvey LAB-AA legal agent evaluation built with Harvey
    Artificial Analysis: Harvey LAB-AA: Artificial Analysis benchmark for legal AI agents
  2. Oct 8, 12:00 AM ET · 2 articles
    Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%
    Artificial Analysis Articles: Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Sherwin Wu
    Harvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings

    AIArtificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination. The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).

  2. Artificial Analysis
    Harvey LAB-AA: Artificial Analysis benchmark for legal AI agents

    AIArtificial Analysis has released Harvey LAB-AA, an evaluation built on Harvey's LAB dataset and developed in collaboration with Harvey. Full results are published on the Artificial Analysis evaluations page, alongside Harvey's commentary on the benchmark and human expert preferences.

  3. Artificial Analysis Articles
    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    AIHarvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Heat trend

Not enough continuous observations to show a trend yet.