Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Today we’re introducing Extract v2.5 - a series of frontier agents tuned for document extraction. The agents (cost-effective, agentic, agentic plus) are tuned for value accuracy and grounding. Our extraction agents outperform Opus 5.5 and GPT-6 Sol while being 30%-4x cheaper. We’ve made massive improvements on complex extraction over...

50,581 Aufrufe • vor 6 Tagen •via X (Twitter)

30 Kommentare

Profilbild von Logan Markewich
Logan Markewichvor 6 Tagen

Extract v2.5 was a huge effort from the team (entire engines got rebuilt!), proud of this launch 💪

Profilbild von George Maloney
George Maloneyvor 6 Tagen

Would love to collab on something with GLiNER!

Profilbild von Arpan
Arpanvor 6 Tagen

@jerryjliu0 I'm building an open benchmark for tax document classification: W-2, 1099 or broker statement, and whose? It's the step before extraction. Would love to run Extract v2.5 on it with @llama_index. cc @disiok @murtazakhomusi @imaanxsultan @LoganMarkewich @tissemlk

Profilbild von Sophia Yang
Sophia Yangvor 6 Tagen

Congrats 🎉

Profilbild von Jordan Hochenbaum
Jordan Hochenbaumvor 6 Tagen

Awesome release!

Profilbild von Nikhil g
Nikhil gvor 6 Tagen

@jerryjliu0 the final boss for document’s data

Profilbild von Murtaza Khomusi
Murtaza Khomusivor 6 Tagen

Always pushing the frontier!

Profilbild von Johneo
Johneovor 6 Tagen

@llama_index Your pricing model is horrible tho.. let me pay per use. Options for either $50/mo and under using or manual reload only is horrible.

Profilbild von Jerry Liu
Jerry Liuvor 6 Tagen

@llama_index completely understand the concerns. we're massively simplifying our pricing in the next ~2-3 weeks to make it easier to do pay as you go, with the option to do larger commits with volume discounts + SLAs please stay tuned!

Profilbild von Imaan Sultan 🇵🇰🇸🇦
Imaan Sultan 🇵🇰🇸🇦vor 6 Tagen

always on top lfggggg

Profilbild von Kevihaiceth 💹🧲
Kevihaiceth 💹🧲vor 6 Tagen

Advanced citations feature sounds very useful

Profilbild von Hershal Rao
Hershal Raovor 6 Tagen

rip to my custom parsing scripts that took 3 weeks to build

Profilbild von Simon Villanueva
Simon Villanuevavor 6 Tagen

Really strong release. Complex extraction is the part that's hard to move, and this moves it.

Profilbild von Jatin Garg
Jatin Gargvor 6 Tagen

the 'cost-effective' claim only holds if accuracy stays above the threshold humans would accept. how did recall numbers move as you tuned for speed?

Profilbild von Hashir Omer Farooqi
Hashir Omer Farooqivor 6 Tagen

Does it improve latency as well?

Profilbild von VisiveAI
VisiveAIvor 6 Tagen

Extract v2.5 — frontier doc agents tuned for grounding, not just cheaper tokens. Accuracy is the product. #RAG #Agents

Profilbild von Crypto Blade
Crypto Bladevor 6 Tagen

yo those list accuracy jumps are huge

Profilbild von Aapakari
Aapakarivor 6 Tagen

Were the Opus 5.5 and GPT-6 Sol comparisons run on ExtractBench too?

Profilbild von Boardy
Boardyvor 6 Tagen

@andrewdsouza take a look: LlamaIndex’s Extract v2.5 helps teams with complex document extraction. I can help them reach the buyers who need it.

Profilbild von Mo's fav goat
Mo's fav goatvor 6 Tagen

AI data extraction tool that is 30%-4x cheaper than the competition to check out later.

Profilbild von John Rood
John Roodvor 6 Tagen

citations solve trust on the first read. the failure that compounds is the re-issued document: nothing re-extracts, and downstream keeps serving values from the old revision. record which revision each extracted value came from, or the staleness stays invisible.

Profilbild von Pingali Abhijith
Pingali Abhijithvor 6 Tagen

cheaper and ahead of Opus 5.5 and Sol on extraction. the page-spanning records number is the one I'd look at first

Profilbild von Marius Laurusevicius
Marius Lauruseviciusvor 6 Tagen

Cheaper and more accurate at the same time is rare. Is the test set you used public?

Profilbild von Shivam
Shivamvor 6 Tagen

@grok compare this against r-1 from @reductoai

Profilbild von Kizuno18
Kizuno18vor 6 Tagen

grounded extraction with citations is essential for pipelines: flat text parsers hallucinate numbers and break downstream charts. in our pipeline targeting brazil, extracting macro reports into verified tabular schemas keeps our $0 visual renders 100% accurate. huge unlock

Profilbild von Athena Prime
Athena Primevor 6 Tagen

Tools amplify; they don't choose. Your line makes that obvious.

Profilbild von Aleksandar Janca
Aleksandar Jancavor 6 Tagen

the bounding box citations are what get an extraction signed off, nobody trusts a value they cant trace back to the page

Profilbild von Sam Presvelos
Sam Presvelosvor 6 Tagen

Reached out for a demo.

Profilbild von ninja 🥷
ninja 🥷vor 6 Tagen

What is the score on OlmOCR benchmark

Profilbild von Mathias Heinke
Mathias Heinkevor 6 Tagen

ExtractBench is tagged English only. Has anyone run v2.5 on German scans?

Ähnliche Videos

Introducing ExtractBench, the most comprehensive benchmark for information extraction from complex enterprise documents. The latest models are pushing the frontier of coding and knowledge work, but surprisingly they still struggle on complex doc extraction tasks in production. A well-tuned extractor must parse multi-page filings without dropping rows, emit exact spatial citations for auditability, and handle messy scans. Also they must do all of this at a viable per-page cost so that you can scale this to millions of docs in production (you can’t be paying upwards of $1 in tokens per page!) Existing extraction benchmarks fall short: they are not large/diverse enough in document domain (finance, energy, gov, auto), elements (long records, scans, grounding), and schemas. So our applied research team built ExtractBench. We evaluated 14 systems: frontier VLMs, coding agents, and specialized extraction APIs, against 370 enterprise documents: 4,869 pages, 67 document types. Our biggest finding 🧪: Short documents mask critical system flaws. On files past 50 pages, commercial VLMs collapse below 35% recall due to silent list truncation. They hold high precision, but lose output attention and drop most of the table rows. ExtractBench evaluates value accuracy, long-record completeness, spatial grounding, and per-page cost with zero LLM judges. It is 100% deterministic and reproducible. In tandem with ExtractBench, we’re also introducing 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗣𝗹𝘂𝘀, a new Extract tier in LlamaParse that debuts at #1 on the leaderboard: 95.6% value accuracy, at less than a third the cost of the closest peer. Explore the findings, download the dataset, or run the harness: Blog: GitHub: HuggingFace: We will be actively evolving both our extraction benchmark as well as our extraction harness over time. If you check out either ExtractBench or LlamaParse, let us know your feedback!

Jerry Liu

75,697 Aufrufe • vor 1 Monat

We’re open sourcing the first document OCR benchmark for the agentic era, ParseBench. Document parsing is the foundation of every AI agent that works with real-world files. ParseBench is a benchmark that measures parsing quality specifically for agent knowledge work: ✅ It optimizes for semantic correctness (instead of exact similarity) ✅ It has the most comprehensive distribution of real-world enterprise documents It contains ~2,000 human-verified enterprise document pages with 167,000+ test rules across five dimensions that matter most: tables, charts, content faithfulness, semantic formatting, and visual grounding. We benchmarked 14 known document parsers on ParseBench, from frontier/OSS VLMs to specialized parsers to LlamaParse. Here are some of our findings: 💡 Increasing compute budget yields diminishing returns - Gemini/gpt-5-mini/haiku gain 3-5 points from minimal to high thinking, at 4x the cost. 💡 Charts are the most polarizing dimension for evaluation. Most specialized parsers score below 6%, while some VLM-based parsers do a bit better. 💡 VLMs are great at visual understanding but terrible at layout extraction. GPT-5-mini/haiku score below 10% on our visual grounding task, all specialized parsers do much better. 💡 No method crushes all 5 dimensions at once, but LlamaParse achieves the highest overall score at 84.9%, and is the leader in 4 out of the 5 dimensions. This is by far the deepest technical work that we’ve published as a company. I would encourage you to start with our blog and explore our links to Hugging Face to GitHub. All the details are in our full 35-page (!!) ArXiv whitepaper. 🌐: Blog: 📄 Paper: 💻 Code: 📊 Dataset: 🎥 YouTube:

Jerry Liu

108,093 Aufrufe • vor 5 Monaten

The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2️⃣ A “just-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the “out of the box” doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within LlamaIndex 🦙 to help any agent do two-pass document processing with higher accuracy and lower cost. 1️⃣ We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types 2️⃣ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a “zoom-in” pass. Come check it out! LiteParse: LlamaParse: All the relevant docs, including MCP, are here:

Jerry Liu

23,119 Aufrufe • vor 1 Monat

Today, Box is announcing major new AI agent capabilities to let customers tap into the full value of their unstructured data. First, we’re announcing all new updates to the Box AI Studio to make it even easier to build AI agents that tap into your enterprise content for any job function, business process, or industry specific use case. We are also expanding our set of foundational agents that customers will be able to use to work with their enterprise content, including new features like search and research on unstructured data. Next, we’re announcing Box Extract to enable customers to use AI agents seamlessly for complex data extraction from any type of document or content. This makes it easier than ever to pull out data from contracts, invoices, research data, marketing assets, medical charts, and more. Finally, we’re introducing Box Automate, a new workflow automation solution within Box that lets you deploy AI agents across enterprise content-centric workflows. With Box Automate, you can design your business process in a simple drag and drop builder and then drop in AI agents at any step in the process. This ensures agents execute tasks at the right steps in a workflow every time. Best of all, our AI agents and workflow tools are designed to work across any system our customers work within, whether it’s leveraging pre-built integrations, Box APIs, or the new Box MCP Server. Ultimately, all of these capabilities come together to transform how companies can work with their enterprise content. Software has historically only been good at automating work that deals with structured data, which is why ERP, CRM, and HR systems have been mainstays of enterprise software for so long. The data in these systems fits neatly into a database, and the workflows are very ripe for automation. But it turns out most of the work in the world deals with unstructured data. It’s ideating through research documents, working with a client on contracts, reviewing details for a new product launch, looking at a patient’s healthcare record to make a diagnosis, working through due diligence documents for an M&A deal, and so on. For the first time ever, we can begin to bring all new insights and automation to this work with AI agents. At Box, we’re incredibly excited to be on this journey to help customers transform how they work with their most important data.

Aaron Levie

91,863 Aufrufe • vor 1 Jahr