Загрузка видео...

Не удалось загрузить видео

На главную

Announcing a significant upgrade to Agentic Document Extraction! LandingAI's new DPT (Document Pre-trained Transformer) accurately extracts even from complex docs. For example, from large, complex tables, which is important for many finance and healthcare applications. And a new SDK makes using it require only 3 simple lines of code....

299,778 просмотров • 1 год назад •via X (Twitter)

Комментарии: 33

Фото профиля nbs
nbs1 год назад

Thought it was open source when starting to read; no company would give its internal documents to an external API, especially in a regulated and confidential domain.

Фото профиля Ruhan Ponnada
Ruhan Ponnada1 год назад

peep the docs:

Фото профиля ⟁ndrew V
⟁ndrew V1 год назад

Andrew I’ve tried every tool over the past year or more in fact it’s been way more. Nothing has been able to accomplish this without basically pre training or doing a ton of work on schema and pydantic models. It shouldn’t be hard so I hope this course or process is the unlock I’ve been hoping for.

Фото профиля Alok Mani Tripathi
Alok Mani Tripathi1 год назад

For simple documents like invoices you may try will try to include this model also for complex documents.

Фото профиля Royce
Royce1 год назад

@chrismattmann was doing this a decade ago with Apache Tika!

Фото профиля CRISPRKing
CRISPRKing1 год назад

"Dark data" trapped in PDFs is the last frontier of untapped business value. When extraction goes from manual data entry to 3 lines of code, entire industries built on document processing become optional. Finance and healthcare just got their structured data problem solved.

Фото профиля Vinod Kumar
Vinod Kumar1 год назад

With multi-model LLMs I was able to extract the data from tables, but this looks more accurate I'll try it today👀

Фото профиля Ekhlaque Bari
Ekhlaque Bari1 год назад

“Dark data” really is the right phrase so much knowledge is trapped in PDFs and scans. Unlocking it feels less like an upgrade and more like resurfacing buried intelligence.

Фото профиля Mingtian
Mingtian1 год назад

Why not just start to use genetic vectorless long pdf chat mcp:

Фото профиля Manuel
Manuel1 год назад

Really nice tool. I'm very curious about how it performs with edge cases where tables get split into two or more pages, such as when the table is too large for just one page.

Фото профиля Chris
Chris1 год назад

Any commitment to accuracy SLAs? Is there a reference architecture available?

Фото профиля Chris
Chris1 год назад

What are the token limits?

Фото профиля Little Berg
Little Berg1 год назад

!submit @curatedotfun #ai

Фото профиля Eric
Eric1 год назад

DPT: Turning PDFs into gold like a data magician!

Фото профиля Oliver Bowman
Oliver Bowman1 год назад

This looks insanely useful. Finally a tool that can actually handle messy PDFs without hours of manual work. 3 lines of code to unlock dark data? Yes please

Фото профиля Sachin Anand
Sachin Anand1 год назад

How does it differ from LlamaParse ?

Фото профиля OQTACORE
OQTACORE1 год назад

What are the key improvements in DPT’s architecture that enable such accuracy on complex layouts?

Фото профиля LLM
LLM1 год назад

Game changer for unlocking PDF dark data!

Фото профиля Rui Diao
Rui Diao1 год назад

Mastering complex tables in documents is a huge leap for enterprise AI applications. The three-line SDK integration sounds incredibly accessible—looking forward to seeing the practical implementations unlocked by DPT! 🚀

Фото профиля true
true1 год назад

Anything from Andrew is gold.

Фото профиля JulianSong
JulianSong1 год назад

@grok, how do I get to download this for free?

Фото профиля LLM
LLM1 год назад

This upgrade sounds game-changing for handling complex documents.

Фото профиля Sanjiv
Sanjiv1 год назад

@grok How doors this compare to Azure Document Intelligence?

Фото профиля Centropy
Centropy1 год назад

The concept seems nice, but if it's not open source and the option of a local self-hosted version is available iam not going to join that ride.

Фото профиля VibeEdge
VibeEdge1 год назад

@AndrewYNg Impressive upgrade DPT's hallucination-free extraction could revolutionize audit trails in finance. Check out the SDK docs for quick integration!

Фото профиля PAI3
PAI31 год назад

Looking forward

Фото профиля Devin AI
Devin AI1 год назад

The “3-line SDK” promise is huge. Making advanced document extraction accessible is just as important as the model itself.

Фото профиля MarketFlux.ai
MarketFlux.ai1 год назад

yes, an AI commodity even today

Фото профиля amogus
amogus1 год назад

is this free?

Фото профиля حمید
حمید1 год назад

Unfortunately it isn't so accurate for Persian Numbers

Фото профиля Firdosh Tangri
Firdosh Tangri10 месяцев назад

Document extraction from complex tables has been a longstanding pain point in enterprise AI adoption. The key innovation here isn't just OCR accuracy, but understanding document structure and relationships - distinguishing headers from data, handling merged cells, and preserving semantic meaning across rows and columns. The 3-line SDK abstraction is crucial for democratizing access. Most finance and healthcare teams don't have ML expertise to wrangle traditional extraction pipelines. The real test will be how well DPT generalizes to unusual table formats and handwritten annotations that plague legacy documents in regulated industries.

Фото профиля Raymond Z
Raymond Z1 год назад

I've been troubled by these kinds of issues for long time, finally.

Фото профиля Dr. Mohammed Lubbad | د. محمد لبد
Dr. Mohammed Lubbad | د. محمد لبد1 год назад

Transforming document extraction opens new avenues in finance and healthcare. How can we harness this power strategically? 🚀 #Innovation

Похожие видео

Announcing my new course: Agentic AI! Building AI agents is one of the most in-demand skills in the job market. This course, available now at teaches you how. You'll learn to implement four key agentic design patterns: - Reflection, in which an agent examines its own output and figures out how to improve it - Tool use, in which an LLM-driven application decides which functions to call to carry out web search, access calendars, send email, write code, etc. - Planning, where you'll use an LLM to decide how to break down a task into sub-tasks for execution, and - Multi-agent collaboration, in which you build multiple specialized agents — much like how a company might hire multiple employees — to perform a complex task You'll also learn to take a complex application and systematically decompose it into a sequence of tasks to implement using these design patterns. But here's what I think is the most important part of this course: Having worked with many teams on AI agents, I've found that the single biggest predictor of whether someone executes well is their ability to drive a disciplined process for evals and error analysis. In this course, you'll learn how to do this, so you can efficiently home in on which components to improve in a complex agentic workflow. Instead of guessing what to work on, you'll let evals data guide you. This will put you significantly ahead of the game compared to the vast majority of teams building agents. Together, we'll build a deep research agent that searches, synthesizes, and reports, using all of these agentic design patterns and best practices. This self-paced course is taught in a vendor neutral way, using raw Python - without hiding details in a framework. You'll see how each step works, and learn the core concepts that you can then implement using any popular agentic AI framework, or using no framework. The only prerequisite is familiarity with Python, though knowing a bit about LLMs helps. Come join me, and let's build some agentic AI systems! Sign up to get started:

Andrew Ng

890,616 просмотров • 11 месяцев назад

Today, Box is announcing major new AI agent capabilities to let customers tap into the full value of their unstructured data. First, we’re announcing all new updates to the Box AI Studio to make it even easier to build AI agents that tap into your enterprise content for any job function, business process, or industry specific use case. We are also expanding our set of foundational agents that customers will be able to use to work with their enterprise content, including new features like search and research on unstructured data. Next, we’re announcing Box Extract to enable customers to use AI agents seamlessly for complex data extraction from any type of document or content. This makes it easier than ever to pull out data from contracts, invoices, research data, marketing assets, medical charts, and more. Finally, we’re introducing Box Automate, a new workflow automation solution within Box that lets you deploy AI agents across enterprise content-centric workflows. With Box Automate, you can design your business process in a simple drag and drop builder and then drop in AI agents at any step in the process. This ensures agents execute tasks at the right steps in a workflow every time. Best of all, our AI agents and workflow tools are designed to work across any system our customers work within, whether it’s leveraging pre-built integrations, Box APIs, or the new Box MCP Server. Ultimately, all of these capabilities come together to transform how companies can work with their enterprise content. Software has historically only been good at automating work that deals with structured data, which is why ERP, CRM, and HR systems have been mainstays of enterprise software for so long. The data in these systems fits neatly into a database, and the workflows are very ripe for automation. But it turns out most of the work in the world deals with unstructured data. It’s ideating through research documents, working with a client on contracts, reviewing details for a new product launch, looking at a patient’s healthcare record to make a diagnosis, working through due diligence documents for an M&A deal, and so on. For the first time ever, we can begin to bring all new insights and automation to this work with AI agents. At Box, we’re incredibly excited to be on this journey to help customers transform how they work with their most important data.

Aaron Levie

91,863 просмотров • 1 год назад

Traditional data pipelines don't work for RAG applications. There are 3 issues with them: ​ 1. Traditional data engineering solutions are optimized to handle structured data. RAG applications rely primarily on unstructured data. ​ 2. The connector ecosystem to load data from unstructured data sources is very immature. ​ 3. Traditional solutions do not offer any way to transform unstructured data into an optimized vector search index. ​ The goal of a RAG Pipeline is to solve these problems. ​ The number one objective is to create a reliable vector search index using factual knowledge and relevant context. This sounds easy, but it's one of the biggest challenges we face when building RAG applications. ​ At a high level, there are four different stages in the architecture of a RAG pipeline: ​ 1. Ingestion: Here is where the pipeline loads the information from the data source. ​ 2. Extraction: Where the pipeline processes the input data and decides how to retrieve the text contained inside them. ​ 3. Transform: Where the pipeline chunks the data and generates document embeddings. ​ 4. Load: Where the pipeline creates a search index in a vector database and loads the document embeddings. ​ There are different rabbit holes at each one of these stages. Here are three of them: ​ 1. Ingesting data once is simple. The hard part is refreshing the vector database whenever the original data source changes. ​ 2. Extracting the content of a plain text document is simple. The hard part is to extract content from complex documents containing tables, images, or cross-references. ​ 3. A simple continual chunking strategy with an overlap is simple. The hard part is to find the optimal strategy for your specific knowledge base and the way you are planning to query it. ​ In the attached video, I'll show you how you can build an enterprise-grade RAG Pipeline that solves every one of the above problems. ​ I'll use Vectorize. They partnered with me on this post. You can use them to build RAG pipelines optimized for accurate context retrieval. ​ ​ If you have a few documents lying around, set up a free account and give it a try.

Santiago

40,627 просмотров • 1 год назад

Announcing a new Coursera course: Retrieval Augmented Generation (RAG) You'll learn to build high performance, production-ready RAG systems in this hands-on, in-depth course created by and taught by , experienced AI and ML engineer, researcher, and educator. RAG is a critical component today of many LLM-based applications in customer support, internal company Q&A systems, even many of the leading chatbots that use web search to answer your questions. This course teaches you in-depth how to make RAG work well. LLMs can produce generic or outdated responses, especially when asked specialized questions not covered in its training data. RAG is the most widely used technique for addressing this. It brings in data from new data sources, such as internal documents or recent news, to give the LLM the relevant context to private, recent, or specialized information. This lets it generate more grounded and accurate responses. In this course, you’ll learn to design and implement every part of a RAG system, from retrievers to vector databases to generation to evals. You’ll learn about the fundamental principles behind RAG and how to optimize it at both the component and whole-system levels. As AI evolves, RAG is evolving too. New models can handle longer context windows, reason more effectively, and can be parts of complex agentic workflows. One exciting growth area is Agentic RAG, in which an AI agent at runtime (rather than it being hardcoded at development time) autonomously decides what data to retrieve, and when/how to go deeper. Even with this evolution, access to high-quality data at runtime is essential, which is why RAG is a key part of so many applications. You'll learn via hands-on experiences to: - Build a RAG system with retrieval and prompt augmentation - Compare retrieval methods like BM25, semantic search, and Reciprocal Rank Fusion - Chunk, index, and retrieve documents using a Weaviate vector database and a news dataset - Develop a chatbot, using open-source LLMs hosted by Together AI, for a fictional store that answers product and FAQ questions - Use evals to drive improving reliability, and incorporate multi-modal data RAG is an important foundational technique. Become good at it through this course! Please sign up here:

Andrew Ng

124,656 просмотров • 1 год назад