Loading video...

Video Failed to Load

Go Home

This week, grounding DINO 1.5 was released It is a new model that uses text prompts to detect objects from videos and images in real-time Examples & demo to try below:

56,027 views • 2 years ago •via X (Twitter)

10 Comments

Allen T.'s profile picture
Allen T.2 years ago

1) Video object detection

Allen T.'s profile picture
Allen T.2 years ago

2) Image object detection

Allen T.'s profile picture
Allen T.2 years ago

3) Paper: Website: Playground:

Allen T.'s profile picture
Allen T.2 years ago

4) Grounding DINO 1.5

LEYVERSE's profile picture
LEYVERSE2 years ago

Someone, please combine this with Segment Anything and replace the boxes with masks of the objects. It would be incredible!

Allen T.'s profile picture
Allen T.2 years ago

I agree! That would be super helpful for editing

Dustin Hollywood's profile picture
Dustin Hollywood2 years ago

This would also rule out copyright violations because the law doesn’t distinguish between human and bot/code “seeing” and learning. What a great loop hole that was prolly not intended but awesome haha 😆

Allen T.'s profile picture
Allen T.2 years ago

I never even considered this angle, Dustin! 👀😯I wonder if that is going to be an argument that a lot of these companies that have been asking for permission to train on user phone and car data will use. It is allowing their AI to see the world without scraping

Happy's profile picture
Happy2 years ago

wow

Allen T.'s profile picture
Allen T.2 years ago

Agree!

Related Videos