one of the most challenging tasks for frontier models is being able to extract thousands of values from extremely dense tables in documents.
our new Extract v2.5 agents are able to get 93%-96%+ on long-list extraction, including cells that fall in between pages.
in contrast, astra gets ~30% accuracy over this benchmark.
in these cases, frontier VLMs can stop early or miss repeated records.
they also struggle to attribute each value back to the source. we are able to attribute every single extracted value back to the source.
check out the video below, the blog post, and llamaparse!
blog: https://www.llamaindex.ai/blog/introducing-extract-v2-5
llamaparse: https://cloud.llamaindex.ai/

