Video wird geladen...
Video konnte nicht geladen werden
I've been training OCR models for 2 years. The progress we've made in handwriting is astonishing - from barely readable to almost perfect transcription.
108,889 Aufrufe • vor 3 Tagen •via X (Twitter)
43 Kommentare

I've been lucky enough to play a small part in the field improving, and I'm excited to solve more edge cases (although this page is close to perfect now, many are not!).

You're doing great job🙏🏾 But it would be more honest; if you benchmarked against Chinese counterparts like GLM-OCR, PaddleOCR etc as well, instead of staying in the western range of models. Their speed & accuracy is clearly beating the ones you are comparing yourself against.

The issue here is that we think olmocr is the most credible ocr benchmark, and chinese labs typically bypass it. We'll figure out a fix though.

Both glm & paddle extract signatures, if I remember correctly. Seattle based, Allen Institute funded olm-ocr is exactly whats limiting true global benchmarking. Chinese have their own standards, which is exactly why benchmarking with them is important.

I think @allenai_org would be surprised to learn that they're holding back global benchmarking

As a Russian native, I can read it, but I'm 100% sure any model can't

I’m so dumb ha I was watching the number in the upper right hand corner thinking “damn. It’s getting worse and worse every model…”

Can it transcribe a doctor’s handwriting?

Depends on the doctor :)

so cool, did you see the dude to solved the old Marmont -napoleon cipher with Astra this week?

Yes, very nice result!

some of the best work I have seen bro

Impressive work

Handwriting is the one I thought would stay broken for years. What moved it?

I have been disappointed to a crazy extent by Tesseract and only use VLMs !

Great job! I'll see how I can use this on my platform! Thanks for sharing.

Vik, congrats on your accomplishments and progress, just wondering if you had compared DataLab API to properly prompted Gemini Flash and how does it compare in accuracy?

does it handle messy cursive too or mostly print?

Cursive works!

I have worked on Many use case where this could be a great fit !

OlmOCR > OmniDocbench

Any plan for chandra ocr 3?

Yes

well done!

Now do it on the voynich manuscript please

Finally my profs can check my answer sheets

Hoes does Datalab do with watermarked documents? Have you ever benchmarked against any of the big LLM vision models like Astra? My experience is they're pretty good at reading handwriting.

Yes, frontier models are getting better - cost/latency/determinism/reliability all matter though

What about GLM?

I haven't tested it extensively on handwriting, but I think people have found it to be good

Awesome 👍

Fuck u and ur progress..just use chatgpt

Isso deve ser excelente para genealogia, história, pesquisa, tradução de documentos antigos etc.

Can we train it with our own dataset? Example: blurred OHT numbers

It is amazing. People like you are needed in Bihar govt offices which is purely shit.

I'd rather keep a few [unclear] markers in the transcription. A guessed surname is harder to catch once the whole letter looks readable.

Yes, that's a good point - we do word confidence, which is a nice proxy for this (some other APIs do the same) -

Great example. And it can even run on CPU (see with quantized models 100% offline and <1 second/page.

great work keep it up bro

so all i do is pick accurate mode to get this? (ie i don't need to configure any settings to be handwriting-specific?)

Has anyone brought you a family letter they’d given up trying to read?

This progress is really incredible, OCR for messy handwriting has always been a tough problem.

Out of curiosity, would you be able to share some difference between Chandra 2.0 and the Datalab Api accurate version of the model ? Essentially, what explains the difference in 1.0% error rate between the 2 models ? IS it just the orchestration pre and post processing of docs ?


