Loading video...

Video Failed to Load

Go Home

One of the hardest problems with using AI agents to automate meaningful document work (contracts, KYC, diligence, claims, and more) is not actually building the agent, but building the UI/UX audit trail so the human can understand decisions linking back to the source documents - in PDFs, Powerpoint, Word,...

19,319 views • 6 months ago •via X (Twitter)

7 Comments

United Records's profile picture
United Records6 months ago

crazy

LightShift Studio's profile picture
LightShift Studio6 months ago

Man, still trying to figure out why you're flipping through PDFs like it's 1999?

Creart's profile picture
Creart6 months ago

In our slice the same problem: one product shot as source → listing variants, ad formats. Outputs need to trace back to the asset. We're building that.

Contextually | Context Studio's profile picture
Contextually | Context Studio6 months ago

Underrated insight. The "last mile" of trust isn't model accuracy it's showing your work. Same principle applies to personal Al: if your assistant makes a recommendation, you need to see why it thought that. Context without transparency is just a black box making decisions for you. Contextually | Cue coming soon get it first:

Leo Pessoa's profile picture
Leo Pessoa4 months ago

The audit trail problem is why object-level grounding matters, not just agent-level. When the domain object carries its sources and updates fields from them, each value traces to a source by construction — you know what grounded what. takes this approach.

Carmelo schepis's profile picture
Carmelo schepis6 months ago

interesting

Thomas Tao's profile picture
Thomas Tao6 months ago

Yeah the audit trail is the product. When spans point to exact PDF regions, trust goes way up.

Related Videos

The latest RAG trend for the current agent harnesses (Codex, Cowork) is to do two passes of document processing to solve a knowledge work task over a data room of documents: 1️⃣ A fast and light pass, oftentimes using a free/OSS doc parsing tool. This can be cheaply run across 10-100-1k’s of files, and enables the agent to then do retrieval (e.g. grep, semantic) to find relevant subsets of context. 2️⃣ A “just-in-time” VLM-based pass. Once the agent finds the relevant pages of context, it will screenshot the documents can call its own VLM (or write code) to dissect the pages. The issue with only using VLM-based OCR tools over massive ad-hoc customer file dumps is that it’s slow and expensive. Doing JIT VLM OCR allows the agent to filter through the data cheaply, but still preserve accuracy for the context that’s needed for the task. The agent harnesses do two-pass document processing by default using off the shelf-tools: pdf2text as the first pass, and using itself (Opus 5) as the second pass. See the below video where Cowork runs over a bunch of PDFs to answer a question about a benchmark graph in the Kimi k3 paper. The main issues here with the “out of the box” doc processing these agents offer are: * Opus 5 is not the best VLM for OCR. It is also way too expensive at scale and lacks grounding * The OSS tools like pypdf, pdf2text, may not be versatile enough as the first pass. * The agent will write a lot of throwaway code to rewrite things an OCR tool would’ve provided out of the box, like chart processing, bounding boxes, confidence scores, leading to increased cost and speed. We have all the tools within LlamaIndex 🦙 to help any agent do two-pass document processing with higher accuracy and lower cost. 1️⃣ We have liteparse for the first pass - a free/OSS parser written in Rust that’s faster/more accurate than other OSS parsers, and supports 50+ document types 2️⃣ We have LlamaParse for the second pass - an agentic document engine that uses VLMs+harnesses to achieve SOTA in accuracy and cost across various doc parsing and extraction tasks. It can be called from any agent harness as an MCP or skill. It takes in page numbers as input, so that the agent can choose to run LlamaParse over a subset of the doc instead of the full doc as a “zoom-in” pass. Come check it out! LiteParse: LlamaParse: All the relevant docs, including MCP, are here:

Jerry Liu

22,763 views • 1 month ago

The same kinds of productivity gains we've seen in coding with AI agents are heading to the rest of knowledge work. This is the jump when you go from having a chatbot to being able to actually have an agent go off and do work for minutes or even hours and come back with a complete work output that you then review. Here's an example of the new Box Agent filling out an RFP response from an existing knowledge base. This process would normally take hours to fill out, and requires the full attention of the user doing the work. Now, you provide the Box Agent with the RFP questions, and it will go off, make a plan, extract all the relevant questions, read through existing source material to come up with an answer, and then generate a new word document as the final output. All while you're doing something else. The key to this architecture is that the agent is able to use all of the same tools in the background that a user uses to get work done. The agent can search for documents, read entire files, run scripts and tools in the background, and even be able to write code on the fly to automate tasks it hasn't seen before. And best of all, the Box Agent will (soon) work from the Box MCP and CLI so you can invoke it in any agentic system as a step in a process. This kind of agent complexity would have been impossible even 6 months ago. Models consistently failed at tracking long running tasks or using the right tools at the right moment for the task. But this is all now possible because of models like GPT-5.4, Opus 4.6, and Gemini 3, and is only getting better by the month. Just as we moved from engineers writing code and using AI as an assistant to answer questions, in many areas of knowledge work -like legal, finance, consulting, sales, marketing, and more- when we have a problem we'll just kick off the AI agent to just go work on it for us in the background.

Aaron Levie

24,728 views • 5 months ago