Video wird geladen...
Video konnte nicht geladen werden
Today we’re introducing Extract v2.5 - a series of frontier agents tuned for document extraction. The agents (cost-effective, agentic, agentic plus) are tuned for value accuracy and grounding. Our extraction agents outperform Opus 5.5 and GPT-6 Sol while being 30%-4x cheaper. We’ve made massive improvements on complex extraction over... show more
50,581 Aufrufe • vor 6 Tagen •via X (Twitter)
30 Kommentare

Extract v2.5 was a huge effort from the team (entire engines got rebuilt!), proud of this launch 💪

Would love to collab on something with GLiNER!

@jerryjliu0 I'm building an open benchmark for tax document classification: W-2, 1099 or broker statement, and whose? It's the step before extraction. Would love to run Extract v2.5 on it with @llama_index. cc @disiok @murtazakhomusi @imaanxsultan @LoganMarkewich @tissemlk

Congrats 🎉

Awesome release!

@jerryjliu0 the final boss for document’s data

Always pushing the frontier!

@llama_index Your pricing model is horrible tho.. let me pay per use. Options for either $50/mo and under using or manual reload only is horrible.

@llama_index completely understand the concerns. we're massively simplifying our pricing in the next ~2-3 weeks to make it easier to do pay as you go, with the option to do larger commits with volume discounts + SLAs please stay tuned!

always on top lfggggg

Advanced citations feature sounds very useful

rip to my custom parsing scripts that took 3 weeks to build

Really strong release. Complex extraction is the part that's hard to move, and this moves it.

the 'cost-effective' claim only holds if accuracy stays above the threshold humans would accept. how did recall numbers move as you tuned for speed?

Does it improve latency as well?

Extract v2.5 — frontier doc agents tuned for grounding, not just cheaper tokens. Accuracy is the product. #RAG #Agents

yo those list accuracy jumps are huge

Were the Opus 5.5 and GPT-6 Sol comparisons run on ExtractBench too?

@andrewdsouza take a look: LlamaIndex’s Extract v2.5 helps teams with complex document extraction. I can help them reach the buyers who need it.

AI data extraction tool that is 30%-4x cheaper than the competition to check out later.

citations solve trust on the first read. the failure that compounds is the re-issued document: nothing re-extracts, and downstream keeps serving values from the old revision. record which revision each extracted value came from, or the staleness stays invisible.

cheaper and ahead of Opus 5.5 and Sol on extraction. the page-spanning records number is the one I'd look at first

Cheaper and more accurate at the same time is rare. Is the test set you used public?

@grok compare this against r-1 from @reductoai

grounded extraction with citations is essential for pipelines: flat text parsers hallucinate numbers and break downstream charts. in our pipeline targeting brazil, extracting macro reports into verified tabular schemas keeps our $0 visual renders 100% accurate. huge unlock

Tools amplify; they don't choose. Your line makes that obvious.

the bounding box citations are what get an extraction signed off, nobody trusts a value they cant trace back to the page

Reached out for a demo.

What is the score on OlmOCR benchmark

ExtractBench is tagged English only. Has anyone run v2.5 on German scans?
