Video yükleniyor...
Video Yüklenemedi
We got extremely good at extracting a scanned from layer by layer 📋 - from printed fields, overlaid text, handwriting, and reviewer annotations. For instance, on an annotated permit, our Extract agents can extract out both the original permit number but also any drawn/handwritten annotations on top. they're also... show more
11,257 görüntüleme • 3 gün önce •via X (Twitter)
11 Yorum

Love your demo videos! 🔥

Can grounding point to an image region when handwriting has no source text layer? Keeping that link through extraction and retrieval would make a wrong field much easier to audit.

the hard part with layered extraction is usually conflicting signals. how do you handle cases where handwriting corrections contradict the printed field value?

for the permit example, a useful robustness check would be rescanning the same form at different resolutions and angles. score the value and its assignment to original/reviewer fields separately, so a readable number in the wrong field doesn't count as success.

The annotated-permit case is exactly where generic OCR dies. In real-estate closings the binding change often lives in a handwritten margin note or a stamp, not the printed field. Nailing the annotation layer is what separates a demo from something a closing team trusts.

Separating the annotation layer is the hard part. The edge case is an annotation that is itself wrong, a reviewer overwriting a correct printed digit. Keeping both values and flagging the disagreement beats letting the later layer win silently.

Layer by layer is a good mental model for web pages too: the server HTML, then what JavaScript adds, then the overlays (cookie walls, modals, sticky nav). Most bad web extractions are the overlay layer winning. Picking which layer you actually want before parsing saves a lot of schema pain.

This looks awesome.

Will extract v2.5 be open source for use to use in our respective regions for data residency?

Layers are the hard part on web pages too. My reading app pulls the article out of messy pages, and when it grabs a comment or a sidebar by mistake, nothing looks broken. The quiz just asks about the wrong text and the reader probably blames themselves.

My constraint on scanned forms is the handwritten layer — printed fields are table stakes, but the reviewer scribble is where our frontier calls quietly burn money and still miss the permit number. A 2x+ error cut at a fraction of the price is the kind of specialized agent I'd rather pay for than keep prompting a general model. How brittle is the grounding when the annotation ink crosses three fields at once?


