Video yükleniyor...
Video Yüklenemedi
Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A well-tuned extractor must parse multi-page filings without dropping rows, emit exact... show more
75,697 görüntüleme • 1 ay önce •via X (Twitter)
25 Yorum

Extraction benchmarks are surprisingly hard to build, there are many more dimensions compared to pure parsing. Congrats to the team on shipping this!

Congrats! Very cool

Super useful!

Hill climbing on task-specific benchmarks helps create specialist agents that achieve higher cost-accuracy ratios than raw LLM APIs. If you're using a legacy OCR service or trying to use raw models, you're likely leaving money on the table and missing out on accuracy gains.

Super useful! Placing the over on this becoming an industry standard

Really good 👏

@llama_index @aman_unsiloed unsiloed score ?

The next frontier for AI is getting structured data out of messy enterprise docs at high accuracy and low cost.

save me so much time (and tokens)

thanks for the support!!

This is where enterprise AI gets real. The hard work is not answering a question once. It is pulling the right answer from messy documents every time.

Very cool!

Love your works.

One of the most overlooked aspects of digital privacy is the trade-off between efficiency, security, and cost under resource constraints.

this feels like a much more useful benchmark than another leaderboard on clean toy datasets

This is one I actually want to run. I like that completeness and grounding are treated separately here. I’m going to try the harness on a few messy long docs I’ve been working with. One thing I’d be curious to see next: a retrieval-first track where the extractor doesn’t get the full document. In practice, you can have perfectly grounded outputs and still miss critical context one step upstream.

True. Models follow instructions better every day. The moment one improvises off-spec, a human still signs off. I see it here in the States every day.

the hard part isn't even writing code, it's handling edge cases like messy docs that models just can't grasp yet

coding agents ace leetcode then choke on a scanned invoice. my scrapers felt that

ExtractBench fills a real gap — most benchmarks test coding or QA, not complex enterprise doc extraction where layout, tables, and multi-page context break models. Useful signal for anyone building RAG or agent pipelines over PDFs.

This is exactly what I’ve been hitting with my own apps – docs are a whole different beast than code. Curious if you saw big gaps between open vs closed models here?

the precision/recall split is the real finding, a system that's confidently wrong and silently drops rows is more dangerous in production than one that just fails loudly. we hit the same pattern building autof24 on scanned italian tax forms, the messy real doc is never the one in the demo.

Silent truncation matches what we see on long contracts. The model looks confident, quietly drops half the schedule rows, and nobody notices until a clause bites. Recall past 50 pages is the metric that actually matters in legal work. Good to see someone measuring it.

My denial agent reads payer letters where the reason code sits in a table split across two pages. It returns the page-one code at full confidence. A blank field is cheap; a confident wrong one costs a rework cycle. The number I want is how often it stays blank.

That sub-35% recall drop on 50+ page docs due to silent list truncation is so real. Nothing worse in production than a model confidently returning clean JSON while quietly dropping half the table rows in the middle.
