Loading video...
Video Failed to Load
Meta’s Segment Anything Model (SAM) 3.1 is now available on Meta Model API, giving developers a fast and lightweight model for detection, segmentation and tracking in a single call on inference tuned for SAM 3.1's architecture. Use a short phrase to find objects in images and video. One API... show more
558,660 views • 4 days ago •via X (Twitter)
14 Comments

one api call for detection, segmentation and tracking is pretty nice. being able to tell it what to find in an image or video and keep the same object tracked feels like a very useful upgrade for devs building visual tools.

ok, but what to do with it?

Now OpenAI should make ZUCC 3.1

Detection, segmentation, and tracking in a single API call makes vision workflows so much cleaner. 👏

SAM 3.1 上 Meta Model API 了短短语就能在图/视频里找对象,一次推理搞定检测分割跟踪钉一下:这是分割模型的 API 可达不是又换了个聊天大模型名

@AIatMeta Great perception layer for agents. The next question is what happens to those detections after they become part of an agent’s knowledge — can we trace the evidence behind the decisions they influence? @finkd

The robotics read on this: perception is the one layer where the compute is genuinely cheap. In Asimov 1's published BOM, every board in the robot — motion control, power, head, the Pi hat — comes to $818 out of roughly $16,000. The actuators and structure are $14,000 of it. So a fast, lightweight detection and tracking model isn't competing for budget with anything. It's the only part of the stack where "just use a better model" is actually the cheap answer. Everywhere below the neck, the answer is still metal.

nice

How clean does tracking stay through heavy occlusions?

We need Dario 3.1

one call for detect, segment and track kills a lot of glue code. keep your text prompts to 2 or 3 words, longer ones get sloppy.

What are the use cases

One call kills detector-tracker stitching

This is cool! Can't wait to use this with Jev. So much cool shit to build.
