Maziyar PANAHI's banner
Maziyar PANAHI's profile picture

Maziyar PANAHI

@MaziyarPanahi19,098 subscribers

the on-device open AI guy | building @OpenMed_AI | making open models @arcee_ai | AI infra & research @CNRS

Shorts

Gemma 4 looks at a parking lot. Decides what to ask. Calls SAM 3.1. "Segment all vehicles." 64 found. "Now just the white ones." 23 found. One model reasoning and orchestrating. One model executing. Both running locally on a MacBook. MLX. No cloud. No API.

Gemma 4 looks at a parking lot. Decides what to ask. Calls SAM 3.1. "Segment all vehicles." 64 found. "Now just the white ones." 23 found. One model reasoning and orchestrating. One model executing. Both running locally on a MacBook. MLX. No cloud. No API.

594,576 views

Gemma 4 watches raw video. Understands the scene. Then prompts SAM 3 to segment and RF-DETR to track. One AI directing two others. Fighter jets. Crowds. Aerial defense footage. All three models running locally on a MacBook. No cloud. What scene should I point this at next?

Gemma 4 watches raw video. Understands the scene. Then prompts SAM 3 to segment and RF-DETR to track. One AI directing two others. Fighter jets. Crowds. Aerial defense footage. All three models running locally on a MacBook. No cloud. What scene should I point this at next?

370,665 views

🚨 BREAKING: Apple just acquired Meta's SAM team! On-device segmentation now ships natively in macOS 26.x This is what it looks like when you run SAM3 on MLX locally. Every car. Every fire. Real-time. M2 laptop. No cloud.

🚨 BREAKING: Apple just acquired Meta's SAM team! On-device segmentation now ships natively in macOS 26.x This is what it looks like when you run SAM3 on MLX locally. Every car. Every fire. Real-time. M2 laptop. No cloud.

329,986 views

Gemma 4 just dropped. I had it captioning video in real-time within an hour. Running locally on a MacBook. No cloud. No API. Real-time scene understanding. Oh and SAM3 is segmenting every object in the same frame. Same laptop.

Gemma 4 just dropped. I had it captioning video in real-time within an hour. Running locally on a MacBook. No cloud. No API. Real-time scene understanding. Oh and SAM3 is segmenting every object in the same frame. Same laptop.

196,846 views

Got GLM-5.2 running on my Mac Studio via llama.cpp, the reasoning behind all my medical agentic workflows. It orchestrates a swarm of tiny on-device OpenMed experts: oncology, meds, labs. No cloud, no rate limits, nobody can take it away. AI must be owned, not rented.

Got GLM-5.2 running on my Mac Studio via llama.cpp, the reasoning behind all my medical agentic workflows. It orchestrates a swarm of tiny on-device OpenMed experts: oncology, meds, labs. No cloud, no rate limits, nobody can take it away. AI must be owned, not rented.

87,461 views

Wow! This is amazing! Segmented every car locally in real time with Meta's SAM3 converted to MLX. Just on-device (M2 laptop) vision getting absurdly good. Local AI is moving faster than most people realize! What other models should we test? what kind of videos?

Wow! This is amazing! Segmented every car locally in real time with Meta's SAM3 converted to MLX. Just on-device (M2 laptop) vision getting absurdly good. Local AI is moving faster than most people realize! What other models should we test? what kind of videos?

177,270 views

Gemma 4 analyzes the video. Generates key questions. Calls Falcon Perception. "Find all the people." 156 found. "Detect only white cars." 8 found. A 26B model is running agentic multi-QA vision orchestration. The models are running locally on a MacBook with MLX. No API.

Gemma 4 analyzes the video. Generates key questions. Calls Falcon Perception. "Find all the people." 156 found. "Detect only white cars." 8 found. A 26B model is running agentic multi-QA vision orchestration. The models are running locally on a MacBook with MLX. No API.

158,988 views

I finally got an open model to do structural biology by itself 🔥 GLM-5.2 drives the Mol* viewer, judges its own render through Qwen3-VL, and refines until the drug pops in its pocket. Then I spun it in 3D. All open, on Hugging Face. What should it build next?

I finally got an open model to do structural biology by itself 🔥 GLM-5.2 drives the Mol* viewer, judges its own render through Qwen3-VL, and refines until the drug pops in its pocket. Then I spun it in 3D. All open, on Hugging Face. What should it build next?

61,797 views

I finally got GLM-5.2 to work an entire 3-year patient chart that only Bonsai 27B was allowed to read 🔥 292 encounters live inside Bonsai on my Mac Studio. llama.cpp, Metal, ternary, 7.2GB, Apache-2.0. The chart never leaves the machine. GLM-5.2 can only ask questions. It asked three. Bonsai answered each in ~2s with 19,398 tokens still cached. Then it caught the thing buried 17 months back: metformin + iodinated contrast at eGFR 39. Nephrology warned about it in 2025. The ED booked the CT anyway. A 27B-class model used to need a datacentre. PrismML say the 1-bit build is 3.9GB and fits an iPhone 17 Pro Max. The orchestrator never touched the data. That's the whole point. What should it read next?

I finally got GLM-5.2 to work an entire 3-year patient chart that only Bonsai 27B was allowed to read 🔥 292 encounters live inside Bonsai on my Mac Studio. llama.cpp, Metal, ternary, 7.2GB, Apache-2.0. The chart never leaves the machine. GLM-5.2 can only ask questions. It asked three. Bonsai answered each in ~2s with 19,398 tokens still cached. Then it caught the thing buried 17 months back: metformin + iodinated contrast at eGFR 39. Nephrology warned about it in 2025. The ED booked the CT anyway. A 27B-class model used to need a datacentre. PrismML say the 1-bit build is 3.9GB and fits an iPhone 17 Pro Max. The orchestrator never touched the data. That's the whole point. What should it read next?

53,274 views

Same text. Two privacy filters. OpenAI's model catches 8 categories. OpenMed catches 55+: medical record numbers, blood type, API keys, financial codes, demographics. Trained on Nemotron data by Nvidia. All on-device. All open-source. Coming soon! What's missing?

Same text. Two privacy filters. OpenAI's model catches 8 categories. OpenMed catches 55+: medical record numbers, blood type, API keys, financial codes, demographics. Trained on Nemotron data by Nvidia. All on-device. All open-source. Coming soon! What's missing?

122,267 views

I played Thinking Machines' new Inkling a 2-minute doctor's visit, on my Mac Studio 🔥 The patient came in about her knee. Inkling listened to the whole thing and flagged the heart failure instead. It heard the real audio, not a transcript. 975B params, the 1-bit GGUF (Unsloth AI), running on llama.cpp (Hugging Face). Nothing left the machine. The knee was the easy part. Buried in the small talk: out of puff on the hill, sleeping on four pillows, ankles like balloons, and she'd quietly stopped her water tablet in spring. Five signals, one diagnosis, none of it why she booked. Her last line: "Should I have said something? It's only the knee I came about." A 975B model just did the hardest thing in medicine: it listened to what the patient didn't think was worth mentioning. What should it hear next?

I played Thinking Machines' new Inkling a 2-minute doctor's visit, on my Mac Studio 🔥 The patient came in about her knee. Inkling listened to the whole thing and flagged the heart failure instead. It heard the real audio, not a transcript. 975B params, the 1-bit GGUF (Unsloth AI), running on llama.cpp (Hugging Face). Nothing left the machine. The knee was the easy part. Buried in the small talk: out of puff on the hill, sleeping on four pillows, ankles like balloons, and she'd quietly stopped her water tablet in spring. Five signals, one diagnosis, none of it why she booked. Her last line: "Should I have said something? It's only the knee I came about." A 975B model just did the hardest thing in medicine: it listened to what the patient didn't think was worth mentioning. What should it hear next?

50,749 views

i'm speechless! i got MiniMax-H3 running fully offline on a mac studio! 15 seconds, three shots, 32kHz stereo denoised jointly with the picture. took 39 minutes! Weights MiniMax Design (H3), Mac engine David Dalcu 🤗 I have ZERO knowledge of making any movies! Sound on and enjoy!

i'm speechless! i got MiniMax-H3 running fully offline on a mac studio! 15 seconds, three shots, 32kHz stereo denoised jointly with the picture. took 39 minutes! Weights MiniMax Design (H3), Mac engine David Dalcu 🤗 I have ZERO knowledge of making any movies! Sound on and enjoy!

30,101 views

it's crazy what a 1.5B model can do these days! "VibeThinker-1.5B is a 1.5-billion parameter dense language model. With a total training cost of only $7,800 USD, it achieves reasoning performance comparable to larger models like GPT OSS-20B Medium." runs perfectly on device!

it's crazy what a 1.5B model can do these days! "VibeThinker-1.5B is a 1.5-billion parameter dense language model. With a total training cost of only $7,800 USD, it achieves reasoning performance comparable to larger models like GPT OSS-20B Medium." runs perfectly on device!

202,387 views

I finally managed to use Unlimited-OCR and Gemma 4 together in OpenMed. 🔥 A real patient chart: read, de-identified, and mapped to FHIR on one laptop. No cloud, no API key, nothing leaves the machine. All via llama.cpp, all free on Hugging Face. 🤗 What do we do next?

I finally managed to use Unlimited-OCR and Gemma 4 together in OpenMed. 🔥 A real patient chart: read, de-identified, and mapped to FHIR on one laptop. No cloud, no API key, nothing leaves the machine. All via llama.cpp, all free on Hugging Face. 🤗 What do we do next?

42,263 views

I showed you SAM 3 all week. This is a 0.6B model that outperforms it. Falcon Perception. Type "detect the plane" and it segments every plane in the frame. Pixel-accurate masks from natural language. Fighter jets. Fire. Crowds. All on a MacBook via MLX. No cloud.

I showed you SAM 3 all week. This is a 0.6B model that outperforms it. Falcon Perception. Type "detect the plane" and it segments every plane in the frame. Pixel-accurate masks from natural language. Fighter jets. Fire. Crowds. All on a MacBook via MLX. No cloud.

63,029 views

From parked cars to an Airbus A321 at cruising altitude. Same laptop. MLX & Torch. SAM3 segments every vehicle on the ground. Yes, just cars. RF-DETR spots the Austrian Airlines jet overhead. Real-time detection. Two open-source models running locally. No cloud. No API.

From parked cars to an Airbus A321 at cruising altitude. Same laptop. MLX & Torch. SAM3 segments every vehicle on the ground. Yes, just cars. RF-DETR spots the Austrian Airlines jet overhead. Real-time detection. Two open-source models running locally. No cloud. No API.

47,210 views

Skills are so much fun. I wrote 70+ for OpenMed and gave one to Kimi K3 🔥 It just read a patient chart it was never allowed to see. 25 identifiers masked on my Mac, 0 left the machine, and it still got every call right. Everything open on Hugging Face 🤗

Skills are so much fun. I wrote 70+ for OpenMed and gave one to Kimi K3 🔥 It just read a patient chart it was never allowed to see. 25 identifiers masked on my Mac, 0 left the machine, and it still got every call right. Everything open on Hugging Face 🤗

14,092 views

Opus 4.8 just did the most important thing in clinical AI: it said no. Asked to reconcile 3 guideline bodies on aspirin, OpenMed Agent searched, found the guidelines weren't in its sources, and labeled 8 gaps instead of inventing them. Refusing to fabricate is the feature.

Opus 4.8 just did the most important thing in clinical AI: it said no. Asked to reconcile 3 guideline bodies on aspirin, OpenMed Agent searched, found the guidelines weren't in its sources, and labeled 8 gaps instead of inventing them. Refusing to fabricate is the feature.

27,991 views

1 week, 4 open-source medical AI shipments: → 35 PII models for Portuguese → openmed==1.1.0 (Brazilian + EU coverage) → OpenMedKit on iPhone: GLiNER + MLX → OpenAI's privacy-filter ported to MLX (24-33x faster) All Apache 2.0. All on-device. 15 seconds recap:

1 week, 4 open-source medical AI shipments: → 35 PII models for Portuguese → openmed==1.1.0 (Brazilian + EU coverage) → OpenMedKit on iPhone: GLiNER + MLX → OpenAI's privacy-filter ported to MLX (24-33x faster) All Apache 2.0. All on-device. 15 seconds recap:

35,507 views

SAM 3D Body on a gymnast. One RGB frame in. Full 3D body mesh out. 18,439 vertices. 36,874 faces. Rotating 360° around a real human. Locally on a 3-year old MacBook. What subject should I mesh next?

SAM 3D Body on a gymnast. One RGB frame in. Full 3D body mesh out. 18,439 vertices. 36,874 faces. Rotating 360° around a real human. Locally on a 3-year old MacBook. What subject should I mesh next?

30,418 views

Videos

No more content to load