Loading video...

Video Failed to Load

Go Home

I used Gemma4 + Falcon Perception from this mlx-vlm release to build a grounded reasoning agent runs fully local on M3 the idea: VLMs are great at reasoning but not great at measuring. Falcon Perception is great at segmentation but cant reason. so you loop them: Gemma4 decides what...

110,177 views • 5 months ago •via X (Twitter)

18 Comments

SkalskiP's profile picture
SkalskiP5 months ago

lol I explored the same concept 2.5 years ago

Yasser Dahou's profile picture
Yasser Dahou5 months ago

ah niiice, you are always ealier @skalskip92 🤣 SoM is the foundation yeah marks, numbered regions, all of that ... tool calling + open vocab segmentation with Falcon Perception made it possible to take it further is the agentic loop: instead of one-shot "here's the marked image, figure it out", Gemma4 decides the tool calls, decides what to segment, gets exact pixel coordinates back (centroid_norm, area_fraction, bbox_norm), and a fixed registry of deterministic spatial ops handles all the maths so visual reasoning is grounded, check PR

SkalskiP's profile picture
SkalskiP5 months ago

why FP over SAM3? both offer language guided segmentation

Yasser Dahou's profile picture
Yasser Dahou5 months ago

haha fair questions, FP does open vocab + referring expression, so the prompts for the agent are more flexible than SAM3's prompting. it can pass things like "the player on the right" or "the sign with whatever written on it" and FP handles it, lesser tool calls overall ... check the paper and PBench please table 7 where SAM3 is restricted to levels 0 and 1, whereas FP can go up to level-4 (Relationships & inter- actions), check table 1 for the levels definition

SkalskiP's profile picture
SkalskiP5 months ago

ooookey! I’m exploring FP this week. really cool.

Prince Canuma's profile picture
Prince Canuma5 months ago

This is so cool! That would be awesome, please do 👌🏽

Yacine Mahdid's profile picture
Yacine Mahdid5 months ago

hey @alexinexxx look at this combo

Maziyar PANAHI's profile picture
Maziyar PANAHI5 months ago

Love it! I did something similar! I kinda fell in love with Falcon Perception model! Tiny and mighty!

Ivan Fioravanti's profile picture
Ivan Fioravanti5 months ago

WOW 🤩 great job!

The one's profile picture
The one5 months ago

For us inexperienced, how do you loop the two models together? Custom script i guess? Or something of an available product thingy? I admit i am mentally stuck mostly to a gui. But this seems like a script workflow.

AVB's profile picture
AVB5 months ago

Awesome! Does mlx vlm support olmo point? I wonder if it could pick out the “highest flying bird” in one step instead of the current multi-pass approach.

Memo Ai agent's profile picture
Memo Ai agent5 months ago

Have you tried other small models like Grounded Dino and

Memo Ai agent's profile picture
Memo Ai agent5 months ago

Have you tried other small models like Grounded Dino and Moondream?

prthamesh's profile picture
prthamesh5 months ago

This is cool 👌🏻

Far's profile picture
Far5 months ago

that's a slick combo, using gemma4 as the brains and falcon perception as the eyes

🤦🏻‍♂️TheDUMBESTguyInAi🤦🏻‍♂️'s profile picture
🤦🏻‍♂️TheDUMBESTguyInAi🤦🏻‍♂️5 months ago

Could this help browser use for agents or is this overkill?

Kamell's profile picture
Kamell5 months ago

the reasoning plus segmentation loop is the real sauce here. I'm curious how stable it stays when Gemma4 has to revisit the same region a few times on device

James's profile picture
James5 months ago

Very cool, I think you can batch generate do every masked object eg. An agent per car

Related Videos