Video wird geladen...
Video konnte nicht geladen werden
we've built the world's most advanced engine for document extraction over complex documents the video below shows an overview of LlamaExtract Agentic Plus in action. long tables, giant forms, calibrated confidence scores, and grounded extraction. if you have complex document extraction use cases and existing vendors aren't cutting it,... show more
14,285 Aufrufe • vor 16 Tagen •via X (Twitter)
30 Kommentare

calibrated confidence scores sounds useful. are they calibrated on held-out documents or user feedback, and does the calibration hold across different document types?

这是我见过最强大的复杂文档提取引擎,帮我解决了很多难题。

@andrewdsouza check this out: LlamaIndex’s complex-document extraction for teams whose current vendors fall short. I can help them reach those buyers.

How about if the data in the document is best modeled relationally?

cool! some cells in long tables are ambiguous even for a careful human, like a merged header or a footnote that changes meaning. does the confidence score drop on those, or does it stay confident and just pick one?

Confidence scores matter

What are the confidence scores calibrated against, your own labeled docs or each customer's? A 0.9 on a blurry scanned lease and a 0.9 on a clean invoice probably should not mean the same thing

Handling deeply nested tables with merged cells reliably is the real stress test. Does Agentic Plus preserve hierarchical relationships in such cases?

Extraction over messy documents is the unglamorous half of every agent, and the half that decides if it ships. In our real-estate ops agents the failures were never reasoning, they were a misread table or a scanned addendum. Get ingestion honest and the smart part gets to work.

calibrated confidence only helps if I can reject below a cutoff. does Agentic Plus expose a per-field score I can gate on, or just a label on the row?

The stat worth zooming into is 96 values extracted, 81 cited: in an insurance proposal, the other 15 are exactly where claims adjusters make their money, so calibrated confidence matters more than extraction speed.

calibrated confidence is the part that actually matters here. 97% accurate extraction is useless if you can't tell which 3% to send to a human. is it calibrated per field, and does it hold on long tables where errors cluster by row?

does each extracted field retain the document version as well as its source location? that seems essential once the output becomes agent memory and the source gets corrected.

grounded extraction sounds impressive, but confidence scores need real-world testing

Extraction accuracy is table stakes. The real signal is "calibrated confidence scores" — the value isn't parsing 95% right, it's knowing which 5% to route to a human before it poisons a reconciliation. AI that admits what it doesn't know separates automation from liability.

Confidence scores and source-backed extraction matter most with complex documents. Getting the right value is one thing; knowing when to flag a result for human review is just as important.

@llama_index Nice, got some accuracy benchmarks?

What do you think @VikParuchuri ?

grounded blank on the long table is the rare ship gate

been benchmarking this exact problem recently. tables are hard; tables + missing cells + watermarks + citations somehow become an entirely different field of computer science

Is Jev is this stack?

calibrated confidence is the whole game imo. long-table extraction accuracy is table stakes now, what makes it usable is a confidence score you can trust to route the ~5% of fields a human should check. a confidently wrong value is worse than a blank. how'd you calibrate it?

calibrated confidence is the part i'd actually pay for. long tables fail one merged cell at a time, and a per-field score is the only way to know which rows to send back for review.

145/146 is very impressive; in legal it’s a deal breaker

do you support eml parsing?

For long tables, I’d score row and column association separately from text accuracy. A correctly read number attached to the wrong line item is still a bad extraction. Reconcile totals, sample low-confidence cells and audit the consequential errors the review queue missed.

amazing

Long tables and giant forms are exactly where document extraction gets difficult. I'm especially curious about the confidence scores — how well do they reflect real extraction errors?

Grounded extraction and calibrated confidence are useful for complex documents.

calibration gets tested at the seams: tables that split across page breaks, forms with merged cells. doc-level accuracy hides exactly those cases. slice the confidence scores by those shapes and they start meaning something.

