Загрузка видео...
Не удалось загрузить видео
Announcing a significant upgrade to Agentic Document Extraction! LandingAI's new DPT (Document Pre-trained Transformer) accurately extracts even from complex docs. For example, from large, complex tables, which is important for many finance and healthcare applications. And a new SDK makes using it require only 3 simple lines of code.... show more
299,778 просмотров • 1 год назад •via X (Twitter)
Комментарии: 33

Thought it was open source when starting to read; no company would give its internal documents to an external API, especially in a regulated and confidential domain.

peep the docs:

Andrew I’ve tried every tool over the past year or more in fact it’s been way more. Nothing has been able to accomplish this without basically pre training or doing a ton of work on schema and pydantic models. It shouldn’t be hard so I hope this course or process is the unlock I’ve been hoping for.

For simple documents like invoices you may try will try to include this model also for complex documents.

@chrismattmann was doing this a decade ago with Apache Tika!

"Dark data" trapped in PDFs is the last frontier of untapped business value. When extraction goes from manual data entry to 3 lines of code, entire industries built on document processing become optional. Finance and healthcare just got their structured data problem solved.

With multi-model LLMs I was able to extract the data from tables, but this looks more accurate I'll try it today👀

“Dark data” really is the right phrase so much knowledge is trapped in PDFs and scans. Unlocking it feels less like an upgrade and more like resurfacing buried intelligence.

Why not just start to use genetic vectorless long pdf chat mcp:

Really nice tool. I'm very curious about how it performs with edge cases where tables get split into two or more pages, such as when the table is too large for just one page.

Any commitment to accuracy SLAs? Is there a reference architecture available?

What are the token limits?

!submit @curatedotfun #ai

DPT: Turning PDFs into gold like a data magician!

This looks insanely useful. Finally a tool that can actually handle messy PDFs without hours of manual work. 3 lines of code to unlock dark data? Yes please

How does it differ from LlamaParse ?

What are the key improvements in DPT’s architecture that enable such accuracy on complex layouts?

Game changer for unlocking PDF dark data!

Mastering complex tables in documents is a huge leap for enterprise AI applications. The three-line SDK integration sounds incredibly accessible—looking forward to seeing the practical implementations unlocked by DPT! 🚀

Anything from Andrew is gold.

@grok, how do I get to download this for free?

This upgrade sounds game-changing for handling complex documents.

@grok How doors this compare to Azure Document Intelligence?

The concept seems nice, but if it's not open source and the option of a local self-hosted version is available iam not going to join that ride.

@AndrewYNg Impressive upgrade DPT's hallucination-free extraction could revolutionize audit trails in finance. Check out the SDK docs for quick integration!

Looking forward

The “3-line SDK” promise is huge. Making advanced document extraction accessible is just as important as the model itself.

yes, an AI commodity even today

is this free?

Unfortunately it isn't so accurate for Persian Numbers

Document extraction from complex tables has been a longstanding pain point in enterprise AI adoption. The key innovation here isn't just OCR accuracy, but understanding document structure and relationships - distinguishing headers from data, handling merged cells, and preserving semantic meaning across rows and columns. The 3-line SDK abstraction is crucial for democratizing access. Most finance and healthcare teams don't have ML expertise to wrangle traditional extraction pipelines. The real test will be how well DPT generalizes to unusual table formats and handwritten annotations that plague legacy documents in regulated industries.

I've been troubled by these kinds of issues for long time, finally.

Transforming document extraction opens new avenues in finance and healthcare. How can we harness this power strategically? 🚀 #Innovation
