正在加载视频...

视频加载失败

We got extremely good at extracting a scanned from layer by layer 📋 - from printed fields, overlaid text, handwriting, and reviewer annotations. For instance, on an annotated permit, our Extract agents can extract out both the original permit number but also any drawn/handwritten annotations on top. they're also...

11,257 次观看 • 3 天前 •via X (Twitter)

11 条评论

KShivendu 🌁 的头像
KShivendu 🌁3 天前

Love your demo videos! 🔥

Mike|AI SOP Lab 的头像
Mike|AI SOP Lab3 天前

Can grounding point to an image region when handwriting has no source text layer? Keeping that link through extraction and retrieval would make a wrong field much easier to audit.

Jatin Garg 的头像
Jatin Garg3 天前

the hard part with layered extraction is usually conflicting signals. how do you handle cases where handwriting corrections contradict the printed field value?

Cspoggers | AI notes 的头像
Cspoggers | AI notes3 天前

for the permit example, a useful robustness check would be rescanning the same form at different resolutions and angles. score the value and its assignment to original/reviewer fields separately, so a readable number in the wrong field doesn't count as success.

Vikas(Vik) Malpani| AI for US Real Estate 的头像
Vikas(Vik) Malpani| AI for US Real Estate3 天前

The annotated-permit case is exactly where generic OCR dies. In real-estate closings the binding change often lives in a handwritten margin note or a stamp, not the printed field. Nailing the annotation layer is what separates a demo from something a closing team trusts.

Robert 的头像
Robert2 天前

Separating the annotation layer is the hard part. The edge case is an annotation that is itself wrong, a reviewer overwriting a correct printed digit. Keeping both values and flagging the disagreement beats letting the later layer win silently.

AgentSearch 的头像
AgentSearch3 天前

Layer by layer is a good mental model for web pages too: the server HTML, then what JavaScript adds, then the overlays (cookie walls, modals, sticky nav). Most bad web extractions are the overlay layer winning. Picking which layer you actually want before parsing saves a lot of schema pain.

Ryan Nguyen 的头像
Ryan Nguyen3 天前

This looks awesome.

ZANDAR 2610 的头像
ZANDAR 26103 天前

Will extract v2.5 be open source for use to use in our respective regions for data residency?

iafineden 的头像
iafineden3 天前

Layers are the hard part on web pages too. My reading app pulls the article out of messy pages, and when it grabs a comment or a sidebar by mistake, nothing looks broken. The quiz just asks about the wrong text and the reader probably blames themselves.

Joel Citron 🟦 的头像
Joel Citron 🟦2 天前

My constraint on scanned forms is the handwritten layer — printed fields are table stakes, but the reviewer scribble is where our frontier calls quietly burn money and still miss the permit number. A 2x+ error cut at a fraction of the price is the kind of specialized agent I'd rather pay for than keep prompting a general model. How brittle is the grounding when the annotation ink crosses three fields at once?

相关视频

Bash is all you need! Which is why I'm introducing my holiday project: just-bash just-bash is a pretty complete implementation of bash in TypeScript designed to be used as a bash tool by AI agents. Because it turns out agents love exploring data via shell scripts, even beyond coding. It comes with grep, sed, awk and the 99th percentile features that an agent like Claude Code or Cursor would use. In fact, Claude Code can use it for secure bash execution. In the package - A bash-tool for AI SDK - A binary for use by yourself or your coding agents - An overlay filesystem to feed files to your agent securely - A Vercel Sandbox compatible API, so you can quickly upgrade to a real VM if you need to run binaries - An example AI agent that explores the just-bash code base using just-bash - I imported the Oils shell bash compatibility suite and just-bash passes a very good chunk What is interesting about this codebase: It was essentially entirely written by Opus 4.5. Coding agents love bash and they are good at reproducing it. They are also great at text-book recursive descent parsers and AST tweet-walk interpreters. That said, it is, like, a lot of code and I didn't read it all 😅. This is very much a hack, but it also seems to be _really_ useful. I haven't really found anything agents want to use that it doesn't support and it's fast and secure (caveats apply). It doesn't have write access to your computer and the filesystem is given a root that the agent cannot escape from. Find it at Related: Our recent blog post how we migrated our data analysis agent to bash tools and achieved incredible quality improvements The video shows the example agent investigating the just-bash code base

Malte Ubl

125,326 次观看 • 9 个月前

🚨 THIS IS ACTUALLY INSANE Your AI agent can have access to the web. But if it can't reliably read what’s actually on the page, that access is almost useless. We looked at Firecrawl as the web layer for AI agents and the numbers are hard to ignore. The setup is simple: Give it a URL, search query, or website. Firecrawl handles the ugly part — crawling, scraping, rendering, extracting, and turning web content into something an AI model can actually use. The headline numbers: → 173,000+ GitHub stars → Search, scrape and interact with the web at scale → Supports web pages, PDFs, DOCX and other content → Structured data extraction for AI workflows → MCP support for connecting it directly to AI agents The workflow looks like this: Search → Scrape → Crawl → Extract → Feed the agent Three things stand out: 1. Scraping becomes an infrastructure layer Instead of maintaining your own pile of HTTP clients, parsers, browser automation and retry logic, you can treat web access as an API. 2. Agents get more than raw HTML The goal isn't just downloading a webpage. It's turning messy web content into clean context that an LLM can reason over. 3. The same layer works across different agent workflows Research agents. RAG pipelines. AI search. Competitive intelligence. Web-data extraction. The interesting shift: AI agents don't just need better models. They need better access to the information those models are supposed to reason about. Firecrawl is building that layer. Save this repo.

Vikas gupta

20,066 次观看 • 1 个月前