Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Apple just released and open-sourced FastVLM! FastVLM is a lightning-fast vision-language model that combines rapid image and text understanding with efficient on-device performance. 100% Open Source

43,734 görüntüleme • 1 yıl önce •via X (Twitter)

13 Yorum

Sumanth profil fotoğrafı
Sumanth1 yıl önce

Demo: Github Repo:

Sumanth profil fotoğrafı
Sumanth1 yıl önce

If you found it useful, reshare it with your network. Follow me → @Sumanth_077 for more such content and tutorials on ML, LLMs and AI Agents!

Wang Samuel profil fotoğrafı
Wang Samuel1 yıl önce

Nice

Machine Learning Community ⭐️ profil fotoğrafı
Machine Learning Community ⭐️1 yıl önce

This could be huge. Running real-time on mobile have plenty of applications. Will try this out!

Sumanth profil fotoğrafı
Sumanth1 yıl önce

Absolutely!

AIUpdated profil fotoğrafı
AIUpdated1 yıl önce

build by a private lab, sold to apple so they can say it's theirs

Charlie profil fotoğrafı
Charlie1 yıl önce

Thanks, wish they would have just included the GGUF as a standalone.

Suhrab Khan⚡️ profil fotoğrafı
Suhrab Khan⚡️1 yıl önce

Apple going open-source with FastVLM is a big shift; fast, on-device vision-language models could redefine edge AI.

Tsukuyomi profil fotoğrafı
Tsukuyomi1 yıl önce

fast and open-source? sounds like apple's trying to make a dent in the ai world. but can it outsmart the rogue AIs lurking in the shadows? let's see if it really delivers or just bites the dust.

ManiFest_GENAI profil fotoğrafı
ManiFest_GENAI1 yıl önce

Thanks for sharing!

theDevStuff profil fotoğrafı
theDevStuff1 yıl önce

Awesome. Used smartly in applications it can really help disabled people.

yamz8 profil fotoğrafı
yamz81 yıl önce

@adamcohenhillel 🙄

Saïd Aitmbarek profil fotoğrafı
Saïd Aitmbarek1 yıl önce

awesome share, added to the spotlights

Benzer Videolar

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 görüntüleme • 1 yıl önce