Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I've been training OCR models for 2 years. The progress we've made in handwriting is astonishing - from barely readable to almost perfect transcription.

108,889 Aufrufe • vor 3 Tagen •via X (Twitter)

43 Kommentare

Profilbild von Vik Paruchuri
Vik Paruchurivor 3 Tagen

I've been lucky enough to play a small part in the field improving, and I'm excited to solve more edge cases (although this page is close to perfect now, many are not!).

Profilbild von Nisargdatta
Nisargdattavor 3 Tagen

You're doing great job🙏🏾 But it would be more honest; if you benchmarked against Chinese counterparts like GLM-OCR, PaddleOCR etc as well, instead of staying in the western range of models. Their speed & accuracy is clearly beating the ones you are comparing yourself against.

Profilbild von Vik Paruchuri
Vik Paruchurivor 3 Tagen

The issue here is that we think olmocr is the most credible ocr benchmark, and chinese labs typically bypass it. We'll figure out a fix though.

Profilbild von Nisargdatta
Nisargdattavor 2 Tagen

Both glm & paddle extract signatures, if I remember correctly. Seattle based, Allen Institute funded olm-ocr is exactly whats limiting true global benchmarking. Chinese have their own standards, which is exactly why benchmarking with them is important.

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

I think @allenai_org would be surprised to learn that they're holding back global benchmarking

Profilbild von Nikita 🤙
Nikita 🤙vor 2 Tagen

As a Russian native, I can read it, but I'm 100% sure any model can't

Profilbild von Michael Kuznetsov
Michael Kuznetsovvor 2 Tagen

I’m so dumb ha I was watching the number in the upper right hand corner thinking “damn. It’s getting worse and worse every model…”

Profilbild von pratyush
pratyushvor 2 Tagen

Can it transcribe a doctor’s handwriting?

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

Depends on the doctor :)

Profilbild von Peter James
Peter Jamesvor 2 Tagen

so cool, did you see the dude to solved the old Marmont -napoleon cipher with Astra this week?

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

Yes, very nice result!

Profilbild von takeyourmeds
takeyourmedsvor 2 Tagen

some of the best work I have seen bro

Profilbild von Gaurav Vyas
Gaurav Vyasvor 2 Tagen

Impressive work

Profilbild von AI Quanting
AI Quantingvor 2 Tagen

Handwriting is the one I thought would stay broken for years. What moved it?

Profilbild von 𝛼
𝛼vor 2 Tagen

I have been disappointed to a crazy extent by Tesseract and only use VLMs !

Profilbild von Isaac | cvats
Isaac | cvatsvor 2 Tagen

Great job! I'll see how I can use this on my platform! Thanks for sharing.

Profilbild von Principal Mahler Appreciator 🇺🇦
Principal Mahler Appreciator 🇺🇦vor 2 Tagen

Vik, congrats on your accomplishments and progress, just wondering if you had compared DataLab API to properly prompted Gemini Flash and how does it compare in accuracy?

Profilbild von Harrison
Harrisonvor 2 Tagen

does it handle messy cursive too or mostly print?

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

Cursive works!

Profilbild von send
sendvor 2 Tagen

I have worked on Many use case where this could be a great fit !

Profilbild von ninja 🥷
ninja 🥷vor 2 Tagen

OlmOCR > OmniDocbench

Profilbild von noob गणितज्ञ
noob गणितज्ञvor 2 Tagen

Any plan for chandra ocr 3?

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

Yes

Profilbild von Zhilin Wang
Zhilin Wangvor 2 Tagen

well done!

Profilbild von Lexeme
Lexemevor 2 Tagen

Now do it on the voynich manuscript please

Profilbild von MAKEki
MAKEkivor 2 Tagen

Finally my profs can check my answer sheets

Profilbild von Charlie Martin
Charlie Martinvor 2 Tagen

Hoes does Datalab do with watermarked documents? Have you ever benchmarked against any of the big LLM vision models like Astra? My experience is they're pretty good at reading handwriting.

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

Yes, frontier models are getting better - cost/latency/determinism/reliability all matter though

Profilbild von Mahmut Kaynar
Mahmut Kaynarvor 2 Tagen

What about GLM?

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

I haven't tested it extensively on handwriting, but I think people have found it to be good

Profilbild von Raj Gupta
Raj Guptavor 2 Tagen

Awesome 👍

Profilbild von Daniel
Danielvor 2 Tagen

Fuck u and ur progress..just use chatgpt

Profilbild von Mateus Ferreira
Mateus Ferreiravor 2 Tagen

Isso deve ser excelente para genealogia, história, pesquisa, tradução de documentos antigos etc.

Profilbild von abdshomad
abdshomadvor 2 Tagen

Can we train it with our own dataset? Example: blurred OHT numbers

Profilbild von homee
homeevor 2 Tagen

It is amazing. People like you are needed in Bihar govt offices which is purely shit.

Profilbild von Hexabl0b
Hexabl0bvor 2 Tagen

I'd rather keep a few [unclear] markers in the transcription. A guessed surname is harder to catch once the whole letter looks readable.

Profilbild von Vik Paruchuri
Vik Paruchurivor 2 Tagen

Yes, that's a good point - we do word confidence, which is a nice proxy for this (some other APIs do the same) -

Profilbild von Omar
Omarvor 2 Tagen

Great example. And it can even run on CPU (see with quantized models 100% offline and <1 second/page.

Profilbild von ZenGPT
ZenGPTvor 2 Tagen

great work keep it up bro

Profilbild von Raymond Weitekamp
Raymond Weitekampvor 2 Tagen

so all i do is pick accurate mode to get this? (ie i don't need to configure any settings to be handwriting-specific?)

Profilbild von Ryan Gilpatric
Ryan Gilpatricvor 2 Tagen

Has anyone brought you a family letter they’d given up trying to read?

Profilbild von Crio Songo
Crio Songovor 2 Tagen

This progress is really incredible, OCR for messy handwriting has always been a tough problem.

Profilbild von FreePalestine!!!!
FreePalestine!!!!vor 2 Tagen

Out of curiosity, would you be able to share some difference between Chandra 2.0 and the Datalab Api accurate version of the model ? Essentially, what explains the difference in 1.0% error rate between the 2 models ? IS it just the orchestration pre and post processing of docs ?

Ähnliche Videos