Загрузка видео...

Не удалось загрузить видео

На главную

I've been training OCR models for 2 years. The progress we've made in handwriting is astonishing - from barely readable to almost perfect transcription.

108,889 просмотров • 2 дней назад •via X (Twitter)

Комментарии: 43

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

I've been lucky enough to play a small part in the field improving, and I'm excited to solve more edge cases (although this page is close to perfect now, many are not!).

Фото профиля Nisargdatta
Nisargdatta2 дней назад

You're doing great job🙏🏾 But it would be more honest; if you benchmarked against Chinese counterparts like GLM-OCR, PaddleOCR etc as well, instead of staying in the western range of models. Their speed & accuracy is clearly beating the ones you are comparing yourself against.

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

The issue here is that we think olmocr is the most credible ocr benchmark, and chinese labs typically bypass it. We'll figure out a fix though.

Фото профиля Nisargdatta
Nisargdatta2 дней назад

Both glm & paddle extract signatures, if I remember correctly. Seattle based, Allen Institute funded olm-ocr is exactly whats limiting true global benchmarking. Chinese have their own standards, which is exactly why benchmarking with them is important.

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

I think @allenai_org would be surprised to learn that they're holding back global benchmarking

Фото профиля Nikita 🤙
Nikita 🤙2 дней назад

As a Russian native, I can read it, but I'm 100% sure any model can't

Фото профиля Michael Kuznetsov
Michael Kuznetsov2 дней назад

I’m so dumb ha I was watching the number in the upper right hand corner thinking “damn. It’s getting worse and worse every model…”

Фото профиля pratyush
pratyush2 дней назад

Can it transcribe a doctor’s handwriting?

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

Depends on the doctor :)

Фото профиля Peter James
Peter James2 дней назад

so cool, did you see the dude to solved the old Marmont -napoleon cipher with Astra this week?

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

Yes, very nice result!

Фото профиля takeyourmeds
takeyourmeds2 дней назад

some of the best work I have seen bro

Фото профиля Gaurav Vyas
Gaurav Vyas2 дней назад

Impressive work

Фото профиля AI Quanting
AI Quanting2 дней назад

Handwriting is the one I thought would stay broken for years. What moved it?

Фото профиля 𝛼
𝛼2 дней назад

I have been disappointed to a crazy extent by Tesseract and only use VLMs !

Фото профиля Isaac | cvats
Isaac | cvats2 дней назад

Great job! I'll see how I can use this on my platform! Thanks for sharing.

Фото профиля Principal Mahler Appreciator 🇺🇦
Principal Mahler Appreciator 🇺🇦2 дней назад

Vik, congrats on your accomplishments and progress, just wondering if you had compared DataLab API to properly prompted Gemini Flash and how does it compare in accuracy?

Фото профиля Harrison
Harrison2 дней назад

does it handle messy cursive too or mostly print?

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

Cursive works!

Фото профиля send
send2 дней назад

I have worked on Many use case where this could be a great fit !

Фото профиля ninja 🥷
ninja 🥷2 дней назад

OlmOCR > OmniDocbench

Фото профиля noob गणितज्ञ
noob गणितज्ञ2 дней назад

Any plan for chandra ocr 3?

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

Yes

Фото профиля Zhilin Wang
Zhilin Wang2 дней назад

well done!

Фото профиля Lexeme
Lexeme2 дней назад

Now do it on the voynich manuscript please

Фото профиля MAKEki
MAKEki2 дней назад

Finally my profs can check my answer sheets

Фото профиля Charlie Martin
Charlie Martin2 дней назад

Hoes does Datalab do with watermarked documents? Have you ever benchmarked against any of the big LLM vision models like Astra? My experience is they're pretty good at reading handwriting.

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

Yes, frontier models are getting better - cost/latency/determinism/reliability all matter though

Фото профиля Mahmut Kaynar
Mahmut Kaynar2 дней назад

What about GLM?

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

I haven't tested it extensively on handwriting, but I think people have found it to be good

Фото профиля Raj Gupta
Raj Gupta2 дней назад

Awesome 👍

Фото профиля Daniel
Daniel2 дней назад

Fuck u and ur progress..just use chatgpt

Фото профиля Mateus Ferreira
Mateus Ferreira2 дней назад

Isso deve ser excelente para genealogia, história, pesquisa, tradução de documentos antigos etc.

Фото профиля abdshomad
abdshomad2 дней назад

Can we train it with our own dataset? Example: blurred OHT numbers

Фото профиля homee
homee2 дней назад

It is amazing. People like you are needed in Bihar govt offices which is purely shit.

Фото профиля Hexabl0b
Hexabl0b2 дней назад

I'd rather keep a few [unclear] markers in the transcription. A guessed surname is harder to catch once the whole letter looks readable.

Фото профиля Vik Paruchuri
Vik Paruchuri2 дней назад

Yes, that's a good point - we do word confidence, which is a nice proxy for this (some other APIs do the same) -

Фото профиля Omar
Omar2 дней назад

Great example. And it can even run on CPU (see with quantized models 100% offline and <1 second/page.

Фото профиля ZenGPT
ZenGPT2 дней назад

great work keep it up bro

Фото профиля Raymond Weitekamp
Raymond Weitekamp2 дней назад

so all i do is pick accurate mode to get this? (ie i don't need to configure any settings to be handwriting-specific?)

Фото профиля Ryan Gilpatric
Ryan Gilpatric2 дней назад

Has anyone brought you a family letter they’d given up trying to read?

Фото профиля Crio Songo
Crio Songo2 дней назад

This progress is really incredible, OCR for messy handwriting has always been a tough problem.

Фото профиля FreePalestine!!!!
FreePalestine!!!!2 дней назад

Out of curiosity, would you be able to share some difference between Chandra 2.0 and the Datalab Api accurate version of the model ? Essentially, what explains the difference in 1.0% error rate between the 2 models ? IS it just the orchestration pre and post processing of docs ?

Похожие видео