Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

one of the most challenging tasks for frontier models is being able to extract thousands of values from extremely dense tables in documents. our new Extract v2.5 agents are able to get 93%-96%+ on long-list extraction, including cells that fall in between pages. in contrast, astra gets ~30% accuracy...

13,103 Aufrufe • vor 8 Tagen •via X (Twitter)

7 Kommentare

Profilbild von VisiveAI
VisiveAIvor 7 Tagen

Long-list extract at 93–96% with source attribution — dense tables still chew up frontier VLMs that bail early.#LlamaIndex #DocAI #Extract

Profilbild von Manu | 🥥
Manu | 🥥vor 7 Tagen

Cells that cross a page break are the real test for document extraction, so it's good to see that case measured. How does accuracy hold up when the table has merged headers?

Profilbild von Thomas
Thomasvor 7 Tagen

93 to 96 with a plus. the plus is doing overtime.

Profilbild von Aleksandar Janca
Aleksandar Jancavor 8 Tagen

attributing every value back to the source is the part I care about, a number I cant trace I cant put in front of a client

Profilbild von Jatin Garg
Jatin Gargvor 7 Tagen

93-96% on long-list extraction is strong. what do the remaining 4-7% failures look like - are they random cell misses or systematic issues with specific table structures?

Profilbild von Winston B.
Winston B.vor 7 Tagen

Early stopping is the nasty one on long tables, since 87 of 238 holdings still looks like a complete answer until someone counts rows. Per-value attribution is what makes that check cheap. Does it resolve to the cell or just the page?

Profilbild von Harley Lewis Foote
Harley Lewis Footevor 7 Tagen

$𝟬.𝟬𝟮 per page (per jerryjliu0) makes frontier-grade extraction economically viable, but accuracy benchmarks measure recall, not field-level trust. In agent commerce, a mis-extracted account number or KYC field doesn't average into a score.

Ähnliche Videos

Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A well-tuned extractor must parse multi-page filings without dropping rows, emit exact spatial citations for auditability, and handle messy scans. Also they must do all of this at a viable per-page cost so that you can scale this to millions of docs in production (you can’t be paying upwards of $1 in tokens per page!) Existing extraction benchmarks fall short: they are not large/diverse enough in document domain (finance, energy, gov, auto), elements (long records, scans, grounding), and schemas. So our applied research team built ExtractBench. We evaluated 14 systems: frontier VLMs, coding agents, and specialized extraction APIs, against 370 enterprise documents: 4,869 pages, 67 document types. Our biggest finding 🧪: Short documents mask critical system flaws. On files past 50 pages, commercial VLMs collapse below 35% recall due to silent list truncation. They hold high precision, but lose output attention and drop most of the table rows. ExtractBench evaluates value accuracy, long-record completeness, spatial grounding, and per-page cost with zero LLM judges. It is 100% deterministic and reproducible. In tandem with ExtractBench, we’re also introducing 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗣𝗹𝘂𝘀, a new Extract tier in LlamaParse that debuts at #1 on the leaderboard: 95.6% value accuracy, at less than a third the cost of the closest peer. Explore the findings, download the dataset, or run the harness: Blog: GitHub: HuggingFace: We will be actively evolving both our extraction benchmark as well as our extraction harness over time. If you check out either ExtractBench or LlamaParse, let us know your feedback!

Jerry Liu

75,710 Aufrufe • vor 2 Monaten