Loading video...
Video Failed to Load
zero-shot image segmentation with Florence-2 + SAM-2 combo just added new mode to my Hugging Face; now you can run open vocabulary detection with Florence-2 + box to mask with SAM2 link:
66,018 views • 2 years ago •via X (Twitter)
12 Comments

last year I released a YouTube tutorial showing how to do a similar thing but with GroundingDINO + SAM back then it was taking 2-4 seconds per image; now it takes 0.5 second

🚀 Revolutionize your QA testing with ZAPTEST AI! ZAPTEST automates testing, slashes costs & boosts ROI up to 10X. No coding needed—just faster, smarter automation. 📅 Book your live demo now:

@huggingface Here's a challenge! Find the labels and arrows in this image 😅

@huggingface It’s the same challenge. You just need to fine-tune Florence to answer questions like this.

@huggingface besides being faster, is it more accurate? this could allow for use cases like generic crop where i tell what i want to crop and it just crops it which could be useful

@huggingface hm... depends on what your query would be. But in general, yes.

@huggingface is there a tutorial on how you did this?

@huggingface Nope. Would you like to see one?

@huggingface You send the bbox center or SAM takes bbox input? I think points are a better UX

@huggingface It can take both boxes or points

@huggingface Betterr than owlv2 +sam and dino+sam?

@huggingface I like it more. But I know @mervenoyann likes OWL a lot!

