正在加载视频...

视频加载失败

I just created my own OCR app using Llama 3.2 vision! Upload an image, and it converts it into structured markdown using Llama 3.2 multimodal! Here's what I used: - Ollama for serving Llama 3.2 vision locally. - Streamlit for the UI. Everything in just 50 lines of code!...

131,139 次观看 • 1 年前 •via X (Twitter)

11 条评论

Avi Chawla 的头像
Avi Chawla1 年前

Code:

Matias Perelli 的头像
Matias Perelli1 年前

I recorded a 31-minute video on how 7-8 figure eCom brands can add $143k using 3 Klaviyo "tweaks". Watch the exact method we used recently with a client. Click below.

Ivan Fioravanti ᯅ 的头像
Ivan Fioravanti ᯅ1 年前

Show me the code 🤣

Avi Chawla 的头像
Avi Chawla1 年前

Here you go:

hoka 的头像
hoka1 年前

Bro why

Avi Chawla 的头像
Avi Chawla1 年前

Akshay and I are co-founders of the same newsletter. So we work together.

remy 的头像
remy1 年前

Not as pretty but I actually did one last night to help someone. One shot ~ich with Cursor. :-)

Paul Lemaistre 的头像
Paul Lemaistre1 年前

Why not use stepfun-ai/GOT-OCR2_0? As far as I know it’s SOTA

Jayson Gent 的头像
Jayson Gent1 年前

@akshay_pachaar Awesome! Didn't even know a vision model was available. I need more vram..

Hemant C. Sharma 的头像
Hemant C. Sharma1 年前

Can anyone give me a solution that will run locally on a mobile app, without needing to connect. Thank you in advance

jatin 的头像
jatin1 年前

Product idea: Integrate with Concur and other expense filing platforms. Extract this data as a json and pass it on to their API to automatically file expenses.

相关视频

"Introducing Multimodal Llama 3.2": As promised two weeks ago, here's the short course on Meta's latest open model! This short course is created with Meta and taught by Amit Sangani, Director of AI Partner Engineering at Meta. Meta’s Llama family of models is leading the way in open models, allowing anyone to download, customize, fine-tune, or build new applications on top of them. Learn about the vision capabilities of the Llama 3.2, and use it for image classification, prompting, tokenization, tool-calling. You'll also learn about the open-source Llama stack, which gives building blocks for many different stages of the LLM application life cycle. In detail, you’ll: - Learn what are the features of Meta's four newest models, and when to use which Llama model. - Learn best practices for multimodal prompting, with applications to advanced image reasoning, illustrated by many examples: Understanding errors on a car dashboard, adding up the total of photographed restaurant receipts, grading written math homework. - Use different roles—system, user, assistant, ipython—in the Llama 3.1 and 3.2 models and the prompt format that identifies those roles. - Understand how Llama uses the tiktoken tokenizer, and how it has expanded to a 128k vocabulary size that improves encoding efficiency and multilingual support. - Learn how to prompt Llama to call built-in and custom tools (functions) with examples for web search and solving math equations. - Learn about Llama Stack, a standardized interface for common toolchain components like fine-tuning or synthetic data generation, useful for building agentic applications. By the end of this course, you’ll be equipped to build out new applications with the new Llama 3.2. Thank you to Ahmad Al-Dahle, Amit Sangani, and the whole AI at Meta team AI at Meta for all the hard work on Llama 3.2 — we’re excited to make these open models even more accessible to more developers with this new course! Please sign up here!

Andrew Ng

131,846 次观看 • 2 年前