Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Meta’s Segment Anything Model (SAM) 3.1 is now available on Meta Model API, giving developers a fast and lightweight model for detection, segmentation and tracking in a single call on inference tuned for SAM 3.1's architecture. Use a short phrase to find objects in images and video. One API...

558,660 Aufrufe • vor 4 Tagen •via X (Twitter)

14 Kommentare

Profilbild von Kate | Neuromancer
Kate | Neuromancervor 4 Tagen

one api call for detection, segmentation and tracking is pretty nice. being able to tell it what to find in an image or video and keep the same object tracked feels like a very useful upgrade for devs building visual tools.

Profilbild von Revhunt
Revhuntvor 4 Tagen

ok, but what to do with it?

Profilbild von Shourjo
Shourjovor 4 Tagen

Now OpenAI should make ZUCC 3.1

Profilbild von Unfair Stack
Unfair Stackvor 4 Tagen

Detection, segmentation, and tracking in a single API call makes vision workflows so much cleaner. 👏

Profilbild von djx
djxvor 4 Tagen

SAM 3.1 上 Meta Model API 了短短语就能在图/视频里找对象,一次推理搞定检测分割跟踪钉一下:这是分割模型的 API 可达不是又换了个聊天大模型名

Profilbild von EKOS _ AGI 🦊 🇮🇷
EKOS _ AGI 🦊 🇮🇷vor 4 Tagen

@AIatMeta Great perception layer for agents. The next question is what happens to those detections after they become part of an agent’s knowledge — can we trace the evidence behind the decisions they influence? @finkd

Profilbild von The Servo Report
The Servo Reportvor 4 Tagen

The robotics read on this: perception is the one layer where the compute is genuinely cheap. In Asimov 1's published BOM, every board in the robot — motion control, power, head, the Pi hat — comes to $818 out of roughly $16,000. The actuators and structure are $14,000 of it. So a fast, lightweight detection and tracking model isn't competing for budget with anything. It's the only part of the stack where "just use a better model" is actually the cheap answer. Everywhere below the neck, the answer is still metal.

Profilbild von Sanskar Pandey
Sanskar Pandeyvor 4 Tagen

nice

Profilbild von MASA
MASAvor 4 Tagen

How clean does tracking stay through heavy occlusions?

Profilbild von Jigs
Jigsvor 4 Tagen

We need Dario 3.1

Profilbild von Deep
Deepvor 4 Tagen

one call for detect, segment and track kills a lot of glue code. keep your text prompts to 2 or 3 words, longer ones get sloppy.

Profilbild von Akash
Akashvor 4 Tagen

What are the use cases

Profilbild von Aapakari
Aapakarivor 4 Tagen

One call kills detector-tracker stitching

Profilbild von Mike
Mikevor 4 Tagen

This is cool! Can't wait to use this with Jev. So much cool shit to build.

Ähnliche Videos

🚀 The Segment Anything Model (SAM) has been upgraded to SAM2, featuring an efficient image encoder for segmenting images and videos. But does SAM2 outperform SAM1 in medical image and video segmentation? We're thrilled to present our paper "Segment Anything in Medical Images and Videos: Benchmark and Deployment"! We comprehensively benchmark SAM2 across 11 medical image modalities and videos. 📄 Paper: 💻 Code: **Highlights:** 1. SAM2 doesn’t always outperform SAM1 in 2D medical images, but excels in video segmentation, making it more accurate and efficient for 3D images, such as CT and MR scans. 2. MedSAM still outperforms SAM2 on most 2D modalities, but SAM2 surpasses MedSAM for 3D image segmentation in a slice-by-slice approach. 3. Segmentation performance varies with model size; sometimes the smallest model outperforms larger ones. 4. Fine-tuning SAM2 significantly boosts its performance for medical image segmentation. While SAM2 may struggle with challenging objects that have unclear boundaries or low contrast, it excels in generating good initial segmentation masks for common medical images and videos. However, the official interface doesn’t support medical data formats and has limitations on video length. To address this, we've developed a 3D Slicer Plugin and Gradio API for efficient 3D medical image and video segmentation. We invite you to try them out and provide feedback! 🔧 Deployment: - 3D Slicer Plugin: - Gradio API: (Note: Due to GPU limitations, the online API is available for only 12 hours and may be slow. We highly recommend deploying the Gradio API with your own computing resources: A big shoutout to Jun Ma (JunMa) who recently joined our UHN AI hub (UHN AI Hub) as Machine Learning Lead, and kudos to all co-authors: Sumin Kim, Feifei Li, Mohammed Baharoon (Mohammed Baharoon), Reza Asakereh, and Hongwei Lyu! This is true teamwork! Looking forward to collaborating with the community to advance 3D medical image and video segmentation foundation models! University Health Network U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology Temerty Centre for AI in Medicine (T-CAIREM) Vector Institute #MedTech #AIinHealthcare #DeepLearning #MedicalImaging #SAM2 #MedSAM #AIResearch

Bo Wang

178,666 Aufrufe • vor 2 Jahren