Video yükleniyor...
Video Yüklenemedi
Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems — frontier VLMs, coding agents, extraction APIs — on 370 enterprise docs, 4,869 pages, 67 doc types. Zero LLM judges, fully deterministic. Biggest finding: past 50 pages, commercial VLMs... show more
32,699 görüntüleme • 2 ay önce •via X (Twitter)
16 Yorum

can you send me dm?

A rigorous benchmark exposing real weaknesses in document extraction

The value is in field-level acceptance, not one aggregate score. Break results out by document type, layout complexity, missing fields, confidence calibration, exception routing, reviewer correction time, cost, and whether the extracted record survives downstream use.

commercial VLMs keeping high precision while recall falls below 35% past 50 pages makes silent truncation easy to miss

Interesting results! It seems like Qwen3.6 35B is the best open-weight/local VLM for documents under 50 pages. Wondering if combining Qwen3.6 35B with Unlimited OCR could help improve performance on longer documents?

Love this benchmark. Real docs, real results.

deterministic grading is the actual news. it moves every argument from the grader to the labels, which is where it belongs, and it caps the benchmark at the quality of the labeling

𝟬% on 369 fields is hallucinated output that validates. Worse than crashes.

ExtractBench is the reality check for doc extraction 370 real enterprise docs prove most VLMs break past 50 pages

Benchmarking extraction across 14 systems on 370 enterprise docs is the right instrument. I have spent time on the parallel problem for web pages, where the quality gap is just as real. One dimension worth separating in results: extraction cost per page. VLMs are accurate but expensive per call. For structured web fields I built an alternative: derive a CSS selector once, cache it per schema, re-derive automatically when a site's layout changes. Repeated extraction cost becomes zero LLM calls.

370 docs, 4869 pages — that is a proper eval. Most benchmarks stop at 50 PDFs and call it a day. Curious: how much variance did you see across the 14 systems on table-heavy vs. narrative-heavy docs? 📊

A much needed benchmark exposing real enterprise extraction gaps

Congrats! Love the @vestaboard in there too!

The number I'd want next to recall: how often the system knew it had missed something. A dropped row in extraction doesn't look like a failure — the table still comes out complete-shaped. A benchmark that separates "missed and flagged" from "missed and silent" is the one that predicts production pain.

Silent row loss is the nasty failure mode here.

Finally a real benchmark that exposes the page limit

