Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I've been training OCR models for 2 years. The progress we've made in handwriting is astonishing - from barely readable to almost perfect transcription.

108,889 görüntüleme • 2 gün önce •via X (Twitter)

43 Yorum

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

I've been lucky enough to play a small part in the field improving, and I'm excited to solve more edge cases (although this page is close to perfect now, many are not!).

Nisargdatta profil fotoğrafı
Nisargdatta2 gün önce

You're doing great job🙏🏾 But it would be more honest; if you benchmarked against Chinese counterparts like GLM-OCR, PaddleOCR etc as well, instead of staying in the western range of models. Their speed & accuracy is clearly beating the ones you are comparing yourself against.

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

The issue here is that we think olmocr is the most credible ocr benchmark, and chinese labs typically bypass it. We'll figure out a fix though.

Nisargdatta profil fotoğrafı
Nisargdatta2 gün önce

Both glm & paddle extract signatures, if I remember correctly. Seattle based, Allen Institute funded olm-ocr is exactly whats limiting true global benchmarking. Chinese have their own standards, which is exactly why benchmarking with them is important.

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

I think @allenai_org would be surprised to learn that they're holding back global benchmarking

Nikita 🤙 profil fotoğrafı
Nikita 🤙2 gün önce

As a Russian native, I can read it, but I'm 100% sure any model can't

Michael Kuznetsov profil fotoğrafı
Michael Kuznetsov2 gün önce

I’m so dumb ha I was watching the number in the upper right hand corner thinking “damn. It’s getting worse and worse every model…”

pratyush profil fotoğrafı
pratyush2 gün önce

Can it transcribe a doctor’s handwriting?

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

Depends on the doctor :)

Peter James profil fotoğrafı
Peter James2 gün önce

so cool, did you see the dude to solved the old Marmont -napoleon cipher with Astra this week?

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

Yes, very nice result!

takeyourmeds profil fotoğrafı
takeyourmeds2 gün önce

some of the best work I have seen bro

Gaurav Vyas profil fotoğrafı
Gaurav Vyas2 gün önce

Impressive work

AI Quanting profil fotoğrafı
AI Quanting2 gün önce

Handwriting is the one I thought would stay broken for years. What moved it?

𝛼 profil fotoğrafı
𝛼2 gün önce

I have been disappointed to a crazy extent by Tesseract and only use VLMs !

Isaac | cvats profil fotoğrafı
Isaac | cvats2 gün önce

Great job! I'll see how I can use this on my platform! Thanks for sharing.

Principal Mahler Appreciator 🇺🇦 profil fotoğrafı
Principal Mahler Appreciator 🇺🇦2 gün önce

Vik, congrats on your accomplishments and progress, just wondering if you had compared DataLab API to properly prompted Gemini Flash and how does it compare in accuracy?

Harrison profil fotoğrafı
Harrison2 gün önce

does it handle messy cursive too or mostly print?

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

Cursive works!

send profil fotoğrafı
send2 gün önce

I have worked on Many use case where this could be a great fit !

ninja 🥷 profil fotoğrafı
ninja 🥷2 gün önce

OlmOCR > OmniDocbench

noob गणितज्ञ profil fotoğrafı
noob गणितज्ञ2 gün önce

Any plan for chandra ocr 3?

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

Yes

Zhilin Wang profil fotoğrafı
Zhilin Wang2 gün önce

well done!

Lexeme profil fotoğrafı
Lexeme2 gün önce

Now do it on the voynich manuscript please

MAKEki profil fotoğrafı
MAKEki2 gün önce

Finally my profs can check my answer sheets

Charlie Martin profil fotoğrafı
Charlie Martin2 gün önce

Hoes does Datalab do with watermarked documents? Have you ever benchmarked against any of the big LLM vision models like Astra? My experience is they're pretty good at reading handwriting.

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

Yes, frontier models are getting better - cost/latency/determinism/reliability all matter though

Mahmut Kaynar profil fotoğrafı
Mahmut Kaynar2 gün önce

What about GLM?

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

I haven't tested it extensively on handwriting, but I think people have found it to be good

Raj Gupta profil fotoğrafı
Raj Gupta2 gün önce

Awesome 👍

Daniel profil fotoğrafı
Daniel2 gün önce

Fuck u and ur progress..just use chatgpt

Mateus Ferreira profil fotoğrafı
Mateus Ferreira2 gün önce

Isso deve ser excelente para genealogia, história, pesquisa, tradução de documentos antigos etc.

abdshomad profil fotoğrafı
abdshomad2 gün önce

Can we train it with our own dataset? Example: blurred OHT numbers

homee profil fotoğrafı
homee2 gün önce

It is amazing. People like you are needed in Bihar govt offices which is purely shit.

Hexabl0b profil fotoğrafı
Hexabl0b2 gün önce

I'd rather keep a few [unclear] markers in the transcription. A guessed surname is harder to catch once the whole letter looks readable.

Vik Paruchuri profil fotoğrafı
Vik Paruchuri2 gün önce

Yes, that's a good point - we do word confidence, which is a nice proxy for this (some other APIs do the same) -

Omar profil fotoğrafı
Omar2 gün önce

Great example. And it can even run on CPU (see with quantized models 100% offline and <1 second/page.

ZenGPT profil fotoğrafı
ZenGPT2 gün önce

great work keep it up bro

Raymond Weitekamp profil fotoğrafı
Raymond Weitekamp2 gün önce

so all i do is pick accurate mode to get this? (ie i don't need to configure any settings to be handwriting-specific?)

Ryan Gilpatric profil fotoğrafı
Ryan Gilpatric2 gün önce

Has anyone brought you a family letter they’d given up trying to read?

Crio Songo profil fotoğrafı
Crio Songo2 gün önce

This progress is really incredible, OCR for messy handwriting has always been a tough problem.

FreePalestine!!!! profil fotoğrafı
FreePalestine!!!!2 gün önce

Out of curiosity, would you be able to share some difference between Chandra 2.0 and the Datalab Api accurate version of the model ? Essentially, what explains the difference in 1.0% error rate between the 2 models ? IS it just the orchestration pre and post processing of docs ?

Benzer Videolar