Skip to content
View original post on X: Jerry Liu· 44/100AI score44/100

LlamaIndex Extract v2.5 hits 93–96% on dense table extraction benchmarks

AISummary

LlamaIndex released Extract v2.5, a set of document extraction agents that it says reach 93%–96%+ accuracy on long-list extraction, including records spanning pages.

The post claims the agents outperform frontier VLMs, which it says stop early, miss repeated records, and struggle to attribute values to sources, while LlamaIndex attributes every extracted value to its source. The agents are available through LlamaParse.

Post on XView on X
@jerryjliu0

one of the most challenging tasks for frontier models is being able to extract thousands of values from extremely dense tables in documents.

our new Extract v2.5 agents are able to get 93%-96%+ on long-list extraction, including cells that fall in between pages.

in contrast, astra gets ~30% accuracy over this benchmark.

in these cases, frontier VLMs can stop early or miss repeated records.

they also struggle to attribute each value back to the source. we are able to attribute every single extracted value back to the source.

check out the video below, the blog post, and llamaparse!

blog: https://www.llamaindex.ai/blog/introducing-extract-v2-5

llamaparse: https://cloud.llamaindex.ai/

Jerry Liu@jerryjliu0
Today we’re introducing Extract v2.5 - a series of frontier agents tuned for document extraction. The agents (cost-effective, agentic, agentic plus) are tuned for value accuracy and grounding. Our extraction agents outperform Opus 5.5 and GPT-6 Sol while being 30%-4x cheaper. We’ve made massive improvements on complex extraction over * long lists (86.1% -> 95.5% on our agentic tier) * records spanning pages (85.5% -> 96.5%) * scanned forms (90.9% -> 95.7%) We’ve also launched the following features: ✅ Advanced citations: we locate bounding boxes for all supporting values for an inferred field, even if there's not an exact match. ✅ Structural Reasoning: we tailor document extraction algorithms depending on the type, layout, and information Our extraction agents are SOTA in price-performance on ExtractBench, across a wide range of cost points. We are the best tool for document extraction across documents of any complexity. Blog: https://www.llamaindex.ai/blog/introducing-extract-v2-5 All of these are available on LlamaParse: https://cloud.llamaindex.ai/ . Come check it out!
View quoted post on X

Source: Jerry Liu · x.comPublished · added here