Loading video...

Video Failed to Load

Go Home

zero-shot image segmentation with Florence-2 + SAM-2 combo just added new mode to my Hugging Face; now you can run open vocabulary detection with Florence-2 + box to mask with SAM2 link:

66,018 views • 2 years ago •via X (Twitter)

12 Comments

SkalskiP's profile picture
SkalskiP2 years ago

last year I released a YouTube tutorial showing how to do a similar thing but with GroundingDINO + SAM back then it was taking 2-4 seconds per image; now it takes 0.5 second

Alex ZAP's profile picture
Alex ZAP1 year ago

🚀 Revolutionize your QA testing with ZAPTEST AI! ZAPTEST automates testing, slashes costs & boosts ROI up to 10X. No coding needed—just faster, smarter automation. 📅 Book your live demo now:

Dan Benyamin (Æ)'s profile picture
Dan Benyamin (Æ)2 years ago

@huggingface Here's a challenge! Find the labels and arrows in this image 😅

SkalskiP's profile picture
SkalskiP2 years ago

@huggingface It’s the same challenge. You just need to fine-tune Florence to answer questions like this.

Itay Bachman's profile picture
Itay Bachman2 years ago

@huggingface besides being faster, is it more accurate? this could allow for use cases like generic crop where i tell what i want to crop and it just crops it which could be useful

SkalskiP's profile picture
SkalskiP2 years ago

@huggingface hm... depends on what your query would be. But in general, yes.

🇮🇳 Chockalingam's profile picture
🇮🇳 Chockalingam2 years ago

@huggingface is there a tutorial on how you did this?

SkalskiP's profile picture
SkalskiP2 years ago

@huggingface Nope. Would you like to see one?

˗ˏˋ⚡️ˎˊ-'s profile picture
˗ˏˋ⚡️ˎˊ-2 years ago

@huggingface You send the bbox center or SAM takes bbox input? I think points are a better UX

SkalskiP's profile picture
SkalskiP2 years ago

@huggingface It can take both boxes or points

Rishabh Jalan's profile picture
Rishabh Jalan2 years ago

@huggingface Betterr than owlv2 +sam and dino+sam?

SkalskiP's profile picture
SkalskiP2 years ago

@huggingface I like it more. But I know @mervenoyann likes OWL a lot!

Related Videos

🚀 The Segment Anything Model (SAM) has been upgraded to SAM2, featuring an efficient image encoder for segmenting images and videos. But does SAM2 outperform SAM1 in medical image and video segmentation? We're thrilled to present our paper "Segment Anything in Medical Images and Videos: Benchmark and Deployment"! We comprehensively benchmark SAM2 across 11 medical image modalities and videos. 📄 Paper: 💻 Code: **Highlights:** 1. SAM2 doesn’t always outperform SAM1 in 2D medical images, but excels in video segmentation, making it more accurate and efficient for 3D images, such as CT and MR scans. 2. MedSAM still outperforms SAM2 on most 2D modalities, but SAM2 surpasses MedSAM for 3D image segmentation in a slice-by-slice approach. 3. Segmentation performance varies with model size; sometimes the smallest model outperforms larger ones. 4. Fine-tuning SAM2 significantly boosts its performance for medical image segmentation. While SAM2 may struggle with challenging objects that have unclear boundaries or low contrast, it excels in generating good initial segmentation masks for common medical images and videos. However, the official interface doesn’t support medical data formats and has limitations on video length. To address this, we've developed a 3D Slicer Plugin and Gradio API for efficient 3D medical image and video segmentation. We invite you to try them out and provide feedback! 🔧 Deployment: - 3D Slicer Plugin: - Gradio API: (Note: Due to GPU limitations, the online API is available for only 12 hours and may be slow. We highly recommend deploying the Gradio API with your own computing resources: A big shoutout to Jun Ma (JunMa) who recently joined our UHN AI hub (UHN AI Hub) as Machine Learning Lead, and kudos to all co-authors: Sumin Kim, Feifei Li, Mohammed Baharoon (Mohammed Baharoon), Reza Asakereh, and Hongwei Lyu! This is true teamwork! Looking forward to collaborating with the community to advance 3D medical image and video segmentation foundation models! University Health Network U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology Temerty Centre for AI in Medicine (T-CAIREM) Vector Institute #MedTech #AIinHealthcare #DeepLearning #MedicalImaging #SAM2 #MedSAM #AIResearch

Bo Wang

178,556 views • 2 years ago