Video wird geladen...
Video konnte nicht geladen werden
everyone is sleeping on this new JSON extraction model! - open-weights - matches Gemini Flash 3.5 - while being 3x faster pulling structured data out of documents usually takes a chain of OCR, an LLM, and extra code to repair broken JSON. lift, from Datalab, does all of this... show more
57,515 Aufrufe • vor 5 Tagen •via X (Twitter)
8 Kommentare

@datalabto how does it handle deeply nested structures? most extraction models start hallucinating keys past 4-5 levels.

@datalabto Returning null instead of guessing is the underrated part: a missing field can route to a human review queue, while a hallucinated value just flows downstream looking valid.

@datalabto 90.2% across 11,000 fields. how many invoices came back with every field right?

@datalabto How does it handle a table that runs across two pages?

@datalabto the ocr then llm then repair chain is where most extraction bugs come from, one model doing all of it would cut a lot

@datalabto Returning null for missing fields is a useful detail. I’d be curious to see examples with messy tables and values that carry over between pages.
@datalabto looks abandoned? 4m in AI time is eternity!

@datalabto Returning null when a field is missing instead of guessing is the line that sold me.
