Harvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings
AIArtificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination. The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).