Loading video...

Video Failed to Load

Go Home

Meta’s Segment Anything Model (SAM) 3.1 is now available on Meta Model API, giving developers a fast and lightweight model for detection, segmentation and tracking in a single call on inference tuned for SAM 3.1's architecture. Use a short phrase to find objects in images and video. One API...

558,660 views • 4 days ago •via X (Twitter)

14 Comments

Kate | Neuromancer's profile picture
Kate | Neuromancer4 days ago

one api call for detection, segmentation and tracking is pretty nice. being able to tell it what to find in an image or video and keep the same object tracked feels like a very useful upgrade for devs building visual tools.

Revhunt's profile picture
Revhunt4 days ago

ok, but what to do with it?

Shourjo's profile picture
Shourjo4 days ago

Now OpenAI should make ZUCC 3.1

Unfair Stack's profile picture
Unfair Stack4 days ago

Detection, segmentation, and tracking in a single API call makes vision workflows so much cleaner. 👏

djx's profile picture
djx4 days ago

SAM 3.1 上 Meta Model API 了短短语就能在图/视频里找对象,一次推理搞定检测分割跟踪钉一下:这是分割模型的 API 可达不是又换了个聊天大模型名

EKOS _ AGI 🦊 🇮🇷's profile picture
EKOS _ AGI 🦊 🇮🇷4 days ago

@AIatMeta Great perception layer for agents. The next question is what happens to those detections after they become part of an agent’s knowledge — can we trace the evidence behind the decisions they influence? @finkd

The Servo Report's profile picture
The Servo Report4 days ago

The robotics read on this: perception is the one layer where the compute is genuinely cheap. In Asimov 1's published BOM, every board in the robot — motion control, power, head, the Pi hat — comes to $818 out of roughly $16,000. The actuators and structure are $14,000 of it. So a fast, lightweight detection and tracking model isn't competing for budget with anything. It's the only part of the stack where "just use a better model" is actually the cheap answer. Everywhere below the neck, the answer is still metal.

Sanskar Pandey's profile picture
Sanskar Pandey4 days ago

nice

MASA's profile picture
MASA4 days ago

How clean does tracking stay through heavy occlusions?

Jigs's profile picture
Jigs4 days ago

We need Dario 3.1

Deep's profile picture
Deep4 days ago

one call for detect, segment and track kills a lot of glue code. keep your text prompts to 2 or 3 words, longer ones get sloppy.

Akash's profile picture
Akash4 days ago

What are the use cases

Aapakari's profile picture
Aapakari4 days ago

One call kills detector-tracker stitching

Mike's profile picture
Mike4 days ago

This is cool! Can't wait to use this with Jev. So much cool shit to build.

Related Videos

🚀 The Segment Anything Model (SAM) has been upgraded to SAM2, featuring an efficient image encoder for segmenting images and videos. But does SAM2 outperform SAM1 in medical image and video segmentation? We're thrilled to present our paper "Segment Anything in Medical Images and Videos: Benchmark and Deployment"! We comprehensively benchmark SAM2 across 11 medical image modalities and videos. 📄 Paper: 💻 Code: **Highlights:** 1. SAM2 doesn’t always outperform SAM1 in 2D medical images, but excels in video segmentation, making it more accurate and efficient for 3D images, such as CT and MR scans. 2. MedSAM still outperforms SAM2 on most 2D modalities, but SAM2 surpasses MedSAM for 3D image segmentation in a slice-by-slice approach. 3. Segmentation performance varies with model size; sometimes the smallest model outperforms larger ones. 4. Fine-tuning SAM2 significantly boosts its performance for medical image segmentation. While SAM2 may struggle with challenging objects that have unclear boundaries or low contrast, it excels in generating good initial segmentation masks for common medical images and videos. However, the official interface doesn’t support medical data formats and has limitations on video length. To address this, we've developed a 3D Slicer Plugin and Gradio API for efficient 3D medical image and video segmentation. We invite you to try them out and provide feedback! 🔧 Deployment: - 3D Slicer Plugin: - Gradio API: (Note: Due to GPU limitations, the online API is available for only 12 hours and may be slow. We highly recommend deploying the Gradio API with your own computing resources: A big shoutout to Jun Ma (JunMa) who recently joined our UHN AI hub (UHN AI Hub) as Machine Learning Lead, and kudos to all co-authors: Sumin Kim, Feifei Li, Mohammed Baharoon (Mohammed Baharoon), Reza Asakereh, and Hongwei Lyu! This is true teamwork! Looking forward to collaborating with the community to advance 3D medical image and video segmentation foundation models! University Health Network U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology Temerty Centre for AI in Medicine (T-CAIREM) Vector Institute #MedTech #AIinHealthcare #DeepLearning #MedicalImaging #SAM2 #MedSAM #AIResearch

Bo Wang

178,666 views • 2 years ago