Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I just created my own OCR app using Llama 3.2 vision! Upload an image, and it converts it into structured markdown using Llama 3.2 multimodal! Here's what I used: - Ollama for serving Llama 3.2 vision locally. - Streamlit for the UI. Everything in just 50 lines of code!...

131,139 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

Avi Chawla profil fotoğrafı
Avi Chawla1 yıl önce

Code:

Matias Perelli profil fotoğrafı
Matias Perelli1 yıl önce

I recorded a 31-minute video on how 7-8 figure eCom brands can add $143k using 3 Klaviyo "tweaks". Watch the exact method we used recently with a client. Click below.

Ivan Fioravanti ᯅ profil fotoğrafı
Ivan Fioravanti ᯅ1 yıl önce

Show me the code 🤣

Avi Chawla profil fotoğrafı
Avi Chawla1 yıl önce

Here you go:

hoka profil fotoğrafı
hoka1 yıl önce

Bro why

Avi Chawla profil fotoğrafı
Avi Chawla1 yıl önce

Akshay and I are co-founders of the same newsletter. So we work together.

remy profil fotoğrafı
remy1 yıl önce

Not as pretty but I actually did one last night to help someone. One shot ~ich with Cursor. :-)

Paul Lemaistre profil fotoğrafı
Paul Lemaistre1 yıl önce

Why not use stepfun-ai/GOT-OCR2_0? As far as I know it’s SOTA

Jayson Gent profil fotoğrafı
Jayson Gent1 yıl önce

@akshay_pachaar Awesome! Didn't even know a vision model was available. I need more vram..

Hemant C. Sharma profil fotoğrafı
Hemant C. Sharma1 yıl önce

Can anyone give me a solution that will run locally on a mobile app, without needing to connect. Thank you in advance

jatin profil fotoğrafı
jatin1 yıl önce

Product idea: Integrate with Concur and other expense filing platforms. Extract this data as a json and pass it on to their API to automatically file expenses.

Benzer Videolar

"Introducing Multimodal Llama 3.2": As promised two weeks ago, here's the short course on Meta's latest open model! This short course is created with Meta and taught by Amit Sangani, Director of AI Partner Engineering at Meta. Meta’s Llama family of models is leading the way in open models, allowing anyone to download, customize, fine-tune, or build new applications on top of them. Learn about the vision capabilities of the Llama 3.2, and use it for image classification, prompting, tokenization, tool-calling. You'll also learn about the open-source Llama stack, which gives building blocks for many different stages of the LLM application life cycle. In detail, you’ll: - Learn what are the features of Meta's four newest models, and when to use which Llama model. - Learn best practices for multimodal prompting, with applications to advanced image reasoning, illustrated by many examples: Understanding errors on a car dashboard, adding up the total of photographed restaurant receipts, grading written math homework. - Use different roles—system, user, assistant, ipython—in the Llama 3.1 and 3.2 models and the prompt format that identifies those roles. - Understand how Llama uses the tiktoken tokenizer, and how it has expanded to a 128k vocabulary size that improves encoding efficiency and multilingual support. - Learn how to prompt Llama to call built-in and custom tools (functions) with examples for web search and solving math equations. - Learn about Llama Stack, a standardized interface for common toolchain components like fine-tuning or synthetic data generation, useful for building agentic applications. By the end of this course, you’ll be equipped to build out new applications with the new Llama 3.2. Thank you to Ahmad Al-Dahle, Amit Sangani, and the whole AI at Meta team AI at Meta for all the hard work on Llama 3.2 — we’re excited to make these open models even more accessible to more developers with this new course! Please sign up here!

Andrew Ng

131,767 görüntüleme • 1 yıl önce