正在加载视频...

视频加载失败

we've built the world's most advanced engine for document extraction over complex documents the video below shows an overview of LlamaExtract Agentic Plus in action. long tables, giant forms, calibrated confidence scores, and grounded extraction. if you have complex document extraction use cases and existing vendors aren't cutting it,...

14,285 次观看 • 17 天前 •via X (Twitter)

30 条评论

Jatin Garg 的头像
Jatin Garg16 天前

calibrated confidence scores sounds useful. are they calibrated on held-out documents or user feedback, and does the calibration hold across different document types?

LimboAI 的头像
LimboAI16 天前

这是我见过最强大的复杂文档提取引擎,帮我解决了很多难题。

Boardy 的头像
Boardy17 天前

@andrewdsouza check this out: LlamaIndex’s complex-document extraction for teams whose current vendors fall short. I can help them reach those buyers.

kirsten lum 的头像
kirsten lum16 天前

How about if the data in the document is best modeled relationally?

Albert Castellana 卡瑟 - e/acc 的头像
Albert Castellana 卡瑟 - e/acc16 天前

cool! some cells in long tables are ambiguous even for a careful human, like a merged header or a footnote that changes meaning. does the confidence score drop on those, or does it stay confident and just pick one?

Leo Lu 的头像
Leo Lu16 天前

Confidence scores matter

Sriram 的头像
Sriram16 天前

What are the confidence scores calibrated against, your own labeled docs or each customer's? A 0.9 on a blurry scanned lease and a 0.9 on a clean invoice probably should not mean the same thing

Robin Yang 的头像
Robin Yang16 天前

Handling deeply nested tables with merged cells reliably is the real stress test. Does Agentic Plus preserve hierarchical relationships in such cases?

Vikas(Vik) Malpani| AI for US Real Estate 的头像
Vikas(Vik) Malpani| AI for US Real Estate16 天前

Extraction over messy documents is the unglamorous half of every agent, and the half that decides if it ships. In our real-estate ops agents the failures were never reasoning, they were a misread table or a scanned addendum. Get ingestion honest and the smart part gets to work.

ethereagle · building 的头像
ethereagle · building16 天前

calibrated confidence only helps if I can reject below a cutoff. does Agentic Plus expose a per-field score I can gate on, or just a label on the row?

Ricci Research 的头像
Ricci Research16 天前

The stat worth zooming into is 96 values extracted, 81 cited: in an insurance proposal, the other 15 are exactly where claims adjusters make their money, so calibrated confidence matters more than extraction speed.

Kisson 的头像
Kisson16 天前

calibrated confidence is the part that actually matters here. 97% accurate extraction is useless if you can't tell which 3% to send to a human. is it calibrated per field, and does it hold on long tables where errors cluster by row?

Braden Stitt 的头像
Braden Stitt16 天前

does each extracted field retain the document version as well as its source location? that seems essential once the output becomes agent memory and the source gets corrected.

Brjan | AI Builder 的头像
Brjan | AI Builder17 天前

grounded extraction sounds impressive, but confidence scores need real-world testing

Neo Huang 的头像
Neo Huang16 天前

Extraction accuracy is table stakes. The real signal is "calibrated confidence scores" — the value isn't parsing 95% right, it's knowing which 5% to route to a human before it poisons a reconciliation. AI that admits what it doesn't know separates automation from liability.

Mian Maaz Ullah Khan 的头像
Mian Maaz Ullah Khan16 天前

Confidence scores and source-backed extraction matter most with complex documents. Getting the right value is one thing; knowing when to flag a result for human review is just as important.

Siddharth Mudgal 的头像
Siddharth Mudgal16 天前

@llama_index Nice, got some accuracy benchmarks?

Nick 的头像
Nick16 天前

What do you think @VikParuchuri ?

Klasta 的头像
Klasta16 天前

grounded blank on the long table is the rare ship gate

tj 的头像
tj16 天前

been benchmarking this exact problem recently. tables are hard; tables + missing cells + watermarks + citations somehow become an entirely different field of computer science

Dan 的头像
Dan16 天前

Is Jev is this stack?

Ankit Agarwal 的头像
Ankit Agarwal16 天前

calibrated confidence is the whole game imo. long-table extraction accuracy is table stakes now, what makes it usable is a confidence score you can trust to route the ~5% of fields a human should check. a confidently wrong value is worse than a blank. how'd you calibrate it?

sea salt enjoyer. 的头像
sea salt enjoyer.17 天前

calibrated confidence is the part i'd actually pay for. long tables fail one merged cell at a time, and a per-field score is the only way to know which rows to send back for review.

ytdaniel 的头像
ytdaniel16 天前

145/146 is very impressive; in legal it’s a deal breaker

Mr. X  的头像
Mr. X 16 天前

do you support eml parsing?

James Walker 的头像
James Walker16 天前

For long tables, I’d score row and column association separately from text accuracy. A correctly read number attached to the wrong line item is still a bad extraction. Reconcile totals, sample low-confidence cells and audit the consequential errors the review queue missed.

Musaab⚡️ 的头像
Musaab⚡️17 天前

amazing

Tom Hughes 的头像
Tom Hughes16 天前

Long tables and giant forms are exactly where document extraction gets difficult. I'm especially curious about the confidence scores — how well do they reflect real extraction errors?

PineWoodsAI 的头像
PineWoodsAI16 天前

Grounded extraction and calibrated confidence are useful for complex documents.

John Rood 的头像
John Rood17 天前

calibration gets tested at the seams: tables that split across page breaks, forms with merged cells. doc-level accuracy hides exactly those cases. slice the confidence scores by those shapes and they start meaning something.

相关视频

Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A well-tuned extractor must parse multi-page filings without dropping rows, emit exact spatial citations for auditability, and handle messy scans. Also they must do all of this at a viable per-page cost so that you can scale this to millions of docs in production (you can’t be paying upwards of $1 in tokens per page!) Existing extraction benchmarks fall short: they are not large/diverse enough in document domain (finance, energy, gov, auto), elements (long records, scans, grounding), and schemas. So our applied research team built ExtractBench. We evaluated 14 systems: frontier VLMs, coding agents, and specialized extraction APIs, against 370 enterprise documents: 4,869 pages, 67 document types. Our biggest finding 🧪: Short documents mask critical system flaws. On files past 50 pages, commercial VLMs collapse below 35% recall due to silent list truncation. They hold high precision, but lose output attention and drop most of the table rows. ExtractBench evaluates value accuracy, long-record completeness, spatial grounding, and per-page cost with zero LLM judges. It is 100% deterministic and reproducible. In tandem with ExtractBench, we’re also introducing 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗣𝗹𝘂𝘀, a new Extract tier in LlamaParse that debuts at #1 on the leaderboard: 95.6% value accuracy, at less than a third the cost of the closest peer. Explore the findings, download the dataset, or run the harness: Blog: GitHub: HuggingFace: We will be actively evolving both our extraction benchmark as well as our extraction harness over time. If you check out either ExtractBench or LlamaParse, let us know your feedback!

Jerry Liu

75,710 次观看 • 2 个月前