Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

MetaAI's SAM 2 struggles when things move fast or when there are crowded, fast-moving objects! Introducing SAMURAI: An adaptation of the Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory. 100% Open Source

324,599 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

Sumanth profil fotoğrafı
Sumanth1 yıl önce

Github Repo:

Sumanth profil fotoğrafı
Sumanth1 yıl önce

If you find this useful, RT to share it with your friends. Don't forget to follow me @Sumanth_077 for more such content and tutorials on Python, AI & ML!

Page to Pixel Publishing profil fotoğrafı
Page to Pixel Publishing2 yıl önce

The Art of Flight is a homage to 80s/90s arcade action shmups with a fresh twist on the genre. Pilot multiple ships at the same time to take on oncoming waves of enemies in this fast paced space shooter. Wishlist on Steam today!

Rethynk AI profil fotoğrafı
Rethynk AI1 yıl önce

That’s a brilliant upgrade! SAMURAI sounds like a game-changer for dynamic environments where traditional SAM models fall short. Motion-aware memory could make zero-shot visual tracking far more robust, especially in real-world applications like sports analysis or autonomous vehicles.

Sumanth profil fotoğrafı
Sumanth1 yıl önce

Absolutely!

anarki🌟 profil fotoğrafı
anarki🌟1 yıl önce

instant @arXivBangers haha let’s gooooo!!!!!

Pérry Odé 🇪🇺🇩🇪🇺🇦🇮🇱 profil fotoğrafı
Pérry Odé 🇪🇺🇩🇪🇺🇦🇮🇱1 yıl önce

If I were a producer of self-shooting and AI-driven drones, SAMURAI would be preferable in this case. 👍

Machine Learning Community ⭐️ profil fotoğrafı
Machine Learning Community ⭐️1 yıl önce

Impressive!

Sumanth profil fotoğrafı
Sumanth1 yıl önce

Indeed!

Appy Pie profil fotoğrafı
Appy Pie1 yıl önce

MetaAI's SAM 2 meets its match with fast motion and crowded scenes, but SAMURAI steps in! Motion-aware memory and zero-shot visual tracking make it a game-changer. Plus, it's 100% open source!

tarama profil fotoğrafı
tarama1 yıl önce

@BenjaminDEKR fyi

Benzer Videolar

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,421 görüntüleme • 8 ay önce