Andi Marafioti's banner
Andi Marafioti's profile picture

Andi Marafioti

@andimarafioti7,586 subscribers

leading multimodal research @huggingface (prev @unity)

Shorts

Can a VLM see without a vision encoder? We trained one for $100, inspired by Gemma 4 12B. Latency on an M3 Pro MacBook: 112 ms -> 1.1 ms for the image path 30% lower end-to-end image+LLM The architecture is just: patchify the image -> linear projection with pos embeddings -> LLM Writeup:

Can a VLM see without a vision encoder? We trained one for $100, inspired by Gemma 4 12B. Latency on an M3 Pro MacBook: 112 ms -> 1.1 ms for the image path 30% lower end-to-end image+LLM The architecture is just: patchify the image -> linear projection with pos embeddings -> LLM Writeup:

60,106 次观看

ok who put Hugging Face in Rick and Morty

ok who put Hugging Face in Rick and Morty

47,931 次观看

Real-time SmolVLM in a web-browser with transformers.js. All running locally with no installs. Just open the website.

Real-time SmolVLM in a web-browser with transformers.js. All running locally with no installs. Just open the website.

86,299 次观看

Videos

没有更多内容可加载