Loading video...

Video Failed to Load

Go Home

Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A well-tuned extractor must parse multi-page filings without dropping rows, emit exact...

75,697 views • 1 month ago •via X (Twitter)

25 Comments

Logan Markewich's profile picture
Logan Markewich1 month ago

Extraction benchmarks are surprisingly hard to build, there are many more dimensions compared to pure parsing. Congrats to the team on shipping this!

Mitchell Troyanovsky's profile picture
Mitchell Troyanovsky1 month ago

Congrats! Very cool

Jerry Chen's profile picture
Jerry Chen1 month ago

Super useful!

Murtaza Khomusi's profile picture
Murtaza Khomusi1 month ago

Hill climbing on task-specific benchmarks helps create specialist agents that achieve higher cost-accuracy ratios than raw LLM APIs. If you're using a legacy OCR service or trying to use raw models, you're likely leaving money on the table and missing out on accuracy gains.

Silen Naihin's profile picture
Silen Naihin1 month ago

Super useful! Placing the over on this becoming an industry standard

Abimael Martell's profile picture
Abimael Martell1 month ago

Really good 👏

Abhishek Eswaran's profile picture
Abhishek Eswaran1 month ago

@llama_index @aman_unsiloed unsiloed score ?

Simon Villanueva's profile picture
Simon Villanueva1 month ago

The next frontier for AI is getting structured data out of messy enterprise docs at high accuracy and low cost.

urik's profile picture
urik1 month ago

save me so much time (and tokens)

Jerry Liu's profile picture
Jerry Liu1 month ago

thanks for the support!!

Vesting ⚡'s profile picture
Vesting ⚡1 month ago

This is where enterprise AI gets real. The hard work is not answering a question once. It is pulling the right answer from messy documents every time.

Nimble's profile picture
Nimble1 month ago

Very cool!

Chris Han's profile picture
Chris Han1 month ago

Love your works.

กรรณิกา พรมโสภา's profile picture
กรรณิกา พรมโสภา1 month ago

One of the most overlooked aspects of digital privacy is the trade-off between efficiency, security, and cost under resource constraints.

Zenad's profile picture
Zenad1 month ago

this feels like a much more useful benchmark than another leaderboard on clean toy datasets

Atakan's profile picture
Atakan1 month ago

This is one I actually want to run. I like that completeness and grounding are treated separately here. I’m going to try the harness on a few messy long docs I’ve been working with. One thing I’d be curious to see next: a retrieval-first track where the extractor doesn’t get the full document. In practice, you can have perfectly grounded outputs and still miss critical context one step upstream.

M.Camisani-Calzolari's profile picture
M.Camisani-Calzolari1 month ago

True. Models follow instructions better every day. The moment one improvises off-spec, a human still signs off. I see it here in the States every day.

Adel Bucetta's profile picture
Adel Bucetta1 month ago

the hard part isn't even writing code, it's handling edge cases like messy docs that models just can't grasp yet

Marvin's profile picture
Marvin1 month ago

coding agents ace leetcode then choke on a scanned invoice. my scrapers felt that

Kyssta's profile picture
Kyssta1 month ago

ExtractBench fills a real gap — most benchmarks test coding or QA, not complex enterprise doc extraction where layout, tables, and multi-page context break models. Useful signal for anyone building RAG or agent pipelines over PDFs.

Henry Nguyen's profile picture
Henry Nguyen1 month ago

This is exactly what I’ve been hitting with my own apps – docs are a whole different beast than code. Curious if you saw big gaps between open vs closed models here?

saietta's profile picture
saietta1 month ago

the precision/recall split is the real finding, a system that's confidently wrong and silently drops rows is more dangerous in production than one that just fails loudly. we hit the same pattern building autof24 on scanned italian tax forms, the messy real doc is never the one in the demo.

Rafie Faruq's profile picture
Rafie Faruq1 month ago

Silent truncation matches what we see on long contracts. The model looks confident, quietly drops half the schedule rows, and nobody notices until a clause bites. Recall past 50 pages is the metric that actually matters in legal work. Good to see someone measuring it.

Nabbil Khan's profile picture
Nabbil Khan1 month ago

My denial agent reads payer letters where the reason code sits in a table split across two pages. It returns the page-one code at full confidence. A blank field is cheap; a confident wrong one costs a rework cycle. The number I want is how often it stays blank.

Habib's profile picture
Habib1 month ago

That sub-35% recall drop on 50+ page docs due to silent list truncation is so real. Nothing worse in production than a model confidently returning clean JSON while quietly dropping half the table rows in the middle.

Related Videos

We’re open sourcing the first document OCR benchmark for the agentic era, ParseBench. Document parsing is the foundation of every AI agent that works with real-world files. ParseBench is a benchmark that measures parsing quality specifically for agent knowledge work: ✅ It optimizes for semantic correctness (instead of exact similarity) ✅ It has the most comprehensive distribution of real-world enterprise documents It contains ~2,000 human-verified enterprise document pages with 167,000+ test rules across five dimensions that matter most: tables, charts, content faithfulness, semantic formatting, and visual grounding. We benchmarked 14 known document parsers on ParseBench, from frontier/OSS VLMs to specialized parsers to LlamaParse. Here are some of our findings: 💡 Increasing compute budget yields diminishing returns - Gemini/gpt-5-mini/haiku gain 3-5 points from minimal to high thinking, at 4x the cost. 💡 Charts are the most polarizing dimension for evaluation. Most specialized parsers score below 6%, while some VLM-based parsers do a bit better. 💡 VLMs are great at visual understanding but terrible at layout extraction. GPT-5-mini/haiku score below 10% on our visual grounding task, all specialized parsers do much better. 💡 No method crushes all 5 dimensions at once, but LlamaParse achieves the highest overall score at 84.9%, and is the leader in 4 out of the 5 dimensions. This is by far the deepest technical work that we’ve published as a company. I would encourage you to start with our blog and explore our links to Hugging Face to GitHub. All the details are in our full 35-page (!!) ArXiv whitepaper. 🌐: Blog: 📄 Paper: 💻 Code: 📊 Dataset: 🎥 YouTube:

Jerry Liu

108,093 views • 5 months ago

The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2️⃣ A “just-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the “out of the box” doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within LlamaIndex 🦙 to help any agent do two-pass document processing with higher accuracy and lower cost. 1️⃣ We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types 2️⃣ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a “zoom-in” pass. Come check it out! LiteParse: LlamaParse: All the relevant docs, including MCP, are here:

Jerry Liu

23,119 views • 1 month ago

Today, Box is announcing major new AI agent capabilities to let customers tap into the full value of their unstructured data. First, we’re announcing all new updates to the Box AI Studio to make it even easier to build AI agents that tap into your enterprise content for any job function, business process, or industry specific use case. We are also expanding our set of foundational agents that customers will be able to use to work with their enterprise content, including new features like search and research on unstructured data. Next, we’re announcing Box Extract to enable customers to use AI agents seamlessly for complex data extraction from any type of document or content. This makes it easier than ever to pull out data from contracts, invoices, research data, marketing assets, medical charts, and more. Finally, we’re introducing Box Automate, a new workflow automation solution within Box that lets you deploy AI agents across enterprise content-centric workflows. With Box Automate, you can design your business process in a simple drag and drop builder and then drop in AI agents at any step in the process. This ensures agents execute tasks at the right steps in a workflow every time. Best of all, our AI agents and workflow tools are designed to work across any system our customers work within, whether it’s leveraging pre-built integrations, Box APIs, or the new Box MCP Server. Ultimately, all of these capabilities come together to transform how companies can work with their enterprise content. Software has historically only been good at automating work that deals with structured data, which is why ERP, CRM, and HR systems have been mainstays of enterprise software for so long. The data in these systems fits neatly into a database, and the workflows are very ripe for automation. But it turns out most of the work in the world deals with unstructured data. It’s ideating through research documents, working with a client on contracts, reviewing details for a new product launch, looking at a patient’s healthcare record to make a diagnosis, working through due diligence documents for an M&A deal, and so on. For the first time ever, we can begin to bring all new insights and automation to this work with AI agents. At Box, we’re incredibly excited to be on this journey to help customers transform how they work with their most important data.

Aaron Levie

91,863 views • 1 year ago

3D-LLM: Injecting the 3D World into Large Language Models paper page: Large language models (LLMs) and Vision-Language Models (VLMs) have been proven to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physics, layout, and so on. In this work, we propose to inject the 3D world into large language models and introduce a whole new family of 3D-LLMs. Specifically, 3D-LLMs can take 3D point clouds and their features as input and perform a diverse set of 3D-related tasks, including captioning, dense captioning, 3D question answering, task decomposition, 3D grounding, 3D-assisted dialog, navigation, and so on. Using three types of prompting mechanisms that we design, we are able to collect over 300k 3D-language data covering these tasks. To efficiently train 3D-LLMs, we first utilize a 3D feature extractor that obtains 3D features from rendered multi- view images. Then, we use 2D VLMs as our backbones to train our 3D-LLMs. By introducing a 3D localization mechanism, 3D-LLMs can better capture 3D spatial information. Experiments on ScanQA show that our model outperforms state-of-the-art baselines by a large margin (e.g., the BLEU-1 score surpasses state-of-the-art score by 9%). Furthermore, experiments on our held-in datasets for 3D captioning, task composition, and 3D-assisted dialogue show that our model outperforms 2D VLMs. Qualitative examples also show that our model could perform more tasks beyond the scope of existing LLMs and VLMs.

AK

249,798 views • 3 years ago

🚀Introducing VisualWebBench: A Comprehensive Benchmark for Multimodal Web Page Understanding and Grounding. 🤔What's this all about? Why this benchmark? > Back in Nov 2023, when we released MMMU ( a comprehensive multimodal understanding benchmark, we received feedback that it included very few UI screenshots. Considering the growing importance of UI understanding, especially with the rise of powerful agents like Devin ( which is built on the strong vision capability of #GPT4, we recognized the need for a benchmark focused on UI screenshot understanding.📸👀 > Multimodal #LLMs have significantly boosted web agents' performance on benchmarks like Mind2Web and WebArena. For instance, the SeeAct agent ( showcases the power of integrating vision into web agents. However, these benchmarks primarily evaluate the end-to-end task execution ability of web agents rather than their understanding of web pages. 🌉 Bridging the Gap with VisualWebBench > To provide a comprehensive evaluation of multimodal LLMs' web page understanding capabilities, we introduce VisualWebBench. Our benchmark spans 139 websites 🌐 across 12 domains 🏷️ and 87 sub-domains 🔍, ensuring a diverse and representative dataset. It assesses MLLMs at three levels: website-level, element-level, and action-level 📊, and encompasses seven tasks designed to evaluate understanding, OCR, grounding, and reasoning abilities 🧠💡. 😮 Surprising Findings > 🎉 Open-source models are catching up: Even though closed-source MLLMs are still leading the leaderboard, we are happy to see open-source models like LLaVA 1.6 34B achieve comparable performance to Gemini Pro. > 🧠 Grounding ability, crucial for developing MLLM-based web applications, is a weakness for most MLLMs. > 🖼️ Importance of Image Resolution: The limited image resolution handling capabilities of most open-source MLLMs restrict their utility in web scenarios, where rich text and elements are prevalent. > 🧱 Relatively strong correlation with general understanding benchmarks like MMMU but weak correlation with web agent benchmarks like Mind2Web. Web agent benchmarks primarily evaluate the end-to-end task execution ability of web agents, which involves a series of actions to accomplish a goal. In contrast, VisualWebBench emphasizes evaluating the foundational skills of MLLMs such as understanding and grounding web page elements. 💡Fun Fact > Claude Sonnet is better than Opus on our benchmark :) 🎓 Conclusion > VisualWebBench serves as a valuable resource for the community, driving research and development in the field of multimodal web page understanding and grounding. As MLLMs continue to evolve and improve, we look forward to seeing new applications and breakthroughs. We believe that our benchmark will contribute to the development of more powerful MLLMs in the web domain, ultimately leading to a more intuitive and efficient user experience on the web. Kudos to the student leads Junpeng Liu Yifan Song and the team Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li! 👏 Check out more details in the Junpeng's thread👇

Xiang Yue

56,697 views • 2 years ago

Zack Polanski speaking at the Bakers, Food and Allied Workers Union, "The government are very good at recognising the problems, at recognising the crisis, the supply chain issues, the energy crisis in Iran, the energy crisis from Ukraine" "But very rarely do they seem to have solutions, things to actually do about it. And when they do have solutions, rarely are they solutions of the scale we need" "So if we look at the energy crisis, for instance, we've heard recently that energy bills in this country could go up 200 pound per year, on average for a household" "That's completely unacceptable" "And far too often I don't hear the solutions from the government that are just so obvious to ramp up our investment in renewable energy to make sure that we're insulating every single home in Britain that needs it" "So it is both warm in the winter and cool in the summer as well as creating hundreds of thousands of good green jobs that could be in public sectors that could be unionised so people are paid properly and treated with dignity and, and care and to remove the subsidies from the fossil fuel companies" "The same people who are destroying our planet should not be getting subsidised by the government at a time when we're in the climate crisis" "But as you know, well it's not just an energy crisis, it's a crisis for food too. Because what we've seen in Iran or implicated by Iran is a fertiliser crisis" "We know how devastating and damaging that already is for our supply chains and for the food that we produce" "And this badly needs intervention, it badly needs help. And what did we see this government do? Well, they cut tariffs on chocolates and biscuits" "Now don't get me wrong, there is room to do this and that will provide a small relief for some families" "That's not a long term plan for UK businesses and UK food production" "That's not a Long term plan to invest in resilience and in our food supply chains, in our energy" "It's not a long term plan that takes these issues seriously, not in the next few weeks or months, but goes we need to fundamentally rethink our systems change and how we provide food security as one of the most fundamental things in our society"

Farrukh

31,310 views • 4 months ago

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,508 views • 1 month ago

60 years ago, Ronald Reagan gave Americans a warning that has stood the test of time. This is from Reagan's "Time for Choosing" speech: "If we lose freedom here, there is no place to escape to. This is the last stand on Earth. And this idea that government is beholden to the people, that it has no other source of power except to sovereign people, is still the newest and most unique idea in all the long history of man's relation to man. This is the issue of this election. Whether we believe in our capacity for self-government or whether we abandon the American revolution and confess that a little intellectual elite in a far-distant capital can plan our lives for us better than we can plan them ourselves." "You and I are told increasingly that we have to choose between a left or right, but I would like to suggest that there is no such thing as a left or right. There is only an up or down--up to a man's age-old dream, the ultimate in individual freedom consistent with law and order--or down to the ant heap totalitarianism, and regardless of their sincerity, their humanitarian motives, those who would trade our freedom for security have embarked on this downward course." ... "Well, I for one resent it when a representative of the people refers to you and me--the free man and woman of this country--as 'the masses.' This is a term we haven't applied to ourselves in America. But beyond that, 'the full power of centralized government'--this was the very thing the Founding Fathers sought to minimize. They knew that governments don't control things. A government can't control the economy without controlling people. And they know when a government sets out to do that, it must use force and coercion to achieve its purpose. They also knew, those Founding Fathers, that outside of its legitimate functions, government does nothing as well or as economically as the private sector of the economy." ... "They say we are always 'against' things, never 'for' anything. Well, the trouble with our liberal friends is not that they are ignorant, but that they know so much that isn't so." ... "Mr. Democrat himself, Al Smith, the great American, came before the American people and charged that the leadership of his party was taking the part of Jefferson, Jackson, and Cleveland down the road under the banners of Marx, Lenin, and Stalin." ... "Our natural, inalienable rights are now considered to be a dispensation of government, and freedom has never been so fragile, so close to slipping from our grasp as it is at this moment." ... "You and I know and do not believe that life is so dear and peace so sweet as to be purchased at the price of chains and slavery. If nothing in life is worth dying for, when did this begin--just in the face of this enemy? Or should Moses have told the children of Israel to live in slavery under the pharaohs? Should Christ have refused the cross? Should the patriots at Concord Bridge have thrown down their guns and refused to fire the shot heard 'round the world? The martyrs of history were not fools, and our honored dead who gave their lives to stop the advance of the Nazis didn't die in vain. Where, then, is the road to peace? Well, it's a simple answer after all." "You and I have the courage to say to our enemies, 'There is a price we will not pay.' There is a point beyond which they must not advance." ... "You and I have a rendezvous with destiny. We will preserve for our children this, the last best hope of man on Earth, or we will sentence them to take the last step into a thousand years of darkness." Doesn't it feel like Reagan is speaking to us TODAY?

Kyle Becker

104,422 views • 2 years ago