M4 Mac Mini AI Cluster Uses EXO Labs with... Thunderbolt 5 interconnect (80Gbps) to run LLMs distributed across 4 M4 Pro Mac Minis. The cluster is small (iPhone for reference). It’s running Nemotron 70B at 8 tok/sec and scales to Llama 405B (benchmarks soon).show more

Alex Cheema
3,516,775 views • 1 year ago
M4 Mac Mini AI cluster running Llama 3.3 70B... Preliminary results show Gigabit ethernet switch (4 tok/sec) is slower than Thunderbolt-5 with EXO Labs Benchmarks including Thunderbolt-5 coming soonTM.show more

Alex Cheema
73,090 views • 1 year ago
Running DeepSeek-V3 on M4 Mac Mini AI Cluster 671B... MoE model distributed across 8 M4 Pro 64GB Mac Minis. Apple Silicon with unified memory is a great fit for MoE.show more

EXO Labs
719,293 views • 1 year ago
Preparing to run the new Llama 3.3 70B on... my mac cluster. Downloading model shards onto 3 x M4 Pro Mac Mini and 1 x M3 Max MacBook Pro. AI cluster is connected by Gigabit ethernet switch with EXO Labsshow more

Alex Cheema
181,917 views • 1 year ago
Running DeepSeek R1 on my desk Uses EXO Labs... with Thunderbolt 5 interconnect (80Gbps) to run the full (671B, 8-bit) DeepSeek R1 distributed across 2 M3 Ultra 512GB Mac Studios (1TB total Unified Memory). Runs at 11 tok/sec. Theoretical max is ~20 tok/sec.show more

Alex Cheema
992,353 views • 1 year ago
AGI at home Running DeepSeek R1 across my 7... M4 Pro Mac Minis and 1 M4 Max MacBook Pro. Total unified memory = 496GB. Uses EXO Labs distributed inference with 4-bit quantization. Next goal is fp8 (requires >700GB)show more

Alex Cheema
1,934,971 views • 1 year ago
How much faster is the new MacBook Pro for... AI inference? M4 Max is 27% faster with 72 tok/sec compared to 56 tok/sec of the M3 Max with MLX running Gemma 2 9B (4bit). The 27% speedup is the same with Llama-3.2-1b, Llama-3.2-3b and others. Next up: EXO Labs M4 cluster.show more

Alex Cheema
529,690 views • 1 year ago
Tested the new MacBook Pro M4 Pro vs. the... Mac mini M1 in LM Studio, running Llama 3.2 3B 4-bit MLX. Results: M1 (8-Core) ------> 25 tok/sec M4 Pro (14-Core) -> 104 tok/sec 🤯 – 4x faster!show more

01000010
111,446 views • 1 year ago
Day 1: Benchmarks We ran 1000+ LLM benchmarks on... real consumer devices. Data includes single-device and multi-device clusters with Tokens-Per-Second (TPS) and Time-To-First-Token (TTFT). Setups tested: 3x M4 Mac Mini cluster, iPhone 15 + S24, RTX4090 & more.show more

EXO Labs
254,968 views • 1 year ago
"It's on my todo to build my own stack... of apple minis and run Llama on my own little cluster on my desk." - Garry Tan Quick AI homelab tutorial: - Install EXO Labs on each mini (open source, steps in README) - Make sure all the minis are on the same WiFi network or for faster/more reliable results, connected over Thunderbolt/Ethernet. Devices will auto-discover each other and auto-shard the LLMs. - Now you can chat to your cluster with your preferred LLM GUI (the tiny corp tinychat ships with exo)show more

Alex Cheema - e/acc
243,248 views • 1 year ago
Distributed training on M4 Mac Mini cluster We implemented... Google DeepMind DiLoCo on Apple Silicon to train large models with 100-1000x less bandwidth compared to DDP baseline. AI is entering a new era where a distributed network of consumer devices can train large models.show more

EXO Labs
347,697 views • 1 year ago
EXO 1.0 at Apple NeurIPS booth. 4 x 512GB... M3 Ultra Mac Studios clustered all-to-all with RDMA over Thunderbolt 5. Runs DeepSeek v3.2 (8-bit) at 25 tok/sec with tensor parallelism.show more

EXO Labs
186,329 views • 7 months ago
Success! 🍎🤖 I can now control Unitree robots on... a Mac Mini M4 running macOS with a simple tweak! The unitree_sdk2_python ( thread implementation isn’t essential—getting the DDS subscriber and publisher to work is enough. I replaced Unitree’s Linux thread logic with a simple Python while loop for sensor reading and robot control. Here’s a video of me controlling a Go2 robot on macOS (M4 Mac Mini) using a Python script for low-level control and sensor reading.show more

Tairan He
57,603 views • 1 year ago
MARCUS CHEN STACKED 30 MAC MINIS INTO AN AI... SERVER FARM. ONE $599 MAC MINI REPLACES YOUR $200/MONTH CLAUDE CODE BILL WITH $3 IN ELECTRICITY two months ago a developer posted his claude code bill on reddit. $170 in 10 days. someone replied "i bought a mac mini m4. haven't paid anthropic since." apple stores ran out of mac minis the same week the m4 chip has 120 gb/s memory bandwidth and unified memory architecture. cpu and gpu share one pool so the model loads once and both read from it. a $599 mac mini runs ai faster than a $1,500 windows pc with a discrete gpu since january 2026 ollama supports the anthropic messages api format. claude code connects directly to your local mac mini with one environment variable. same interface, zero api costs, $0 per request a heavy developer pays $459 a month across claude code max, chatgpt pro, gemini, cursor and copilot. that's $5,508 a year. the mac mini pays off in 3 months and runs on $3 in electricity after that uber rolled out claude code to 5,000 engineers and burned through their $3.4 billion 2026 ai budget in 4 months. the people who own the hardware in 2026 are going to look very far ahead in 2028 bookmark this and read the article belowshow more

starmex
357,349 views • 1 month ago
THIS SHELF OF MAC MINIS REPLACES $4,080 A YEAR... IN AI SUBSCRIPTIONS 00:02 the camera pans across a shelf of stacked Mac minis and the trick is obvious: that silent little farm runs the models you rent every month most people pay 7 companies for AI and use 3 of the tools. they forget the rest on the credit card and call it a stack the Mac mini M4 ends that. one shared memory pool means a $599 box runs 7B and 8B models faster than Windows machines that cost twice as much ollama pull, one command. open webui in one docker line. point Claude Code at localhost and it just works it draws 10 to 30 watts, sits silent next to a router, and runs 24/7 for $3 a month in power it pays back a $20 ChatGPT Plus sub in 3 months, then saves you $4,000 a year while the frontier still rents you compute every month you wait is another $340 gone for compute that fits on a shelfshow more

Fokki
12,933 views • 27 days ago
A 27-year-old guy from China bought 100 Mac minis... and turned his apartment into a server hub Instead of furniture, the space is filled with neat racks where a hundred Mac minis hum like a living organism, processing gigabytes of data in a continuous stream Many small AI startups in Asia cannot afford to train models on NVIDIA H100 servers: it is insanely expensive He provides his Mac minis for running and fine-tuning small, highly specialized models. He takes on the dirty work of data labeling and quality assurance by using his network of agents that check each other's work > His initial investment was $59,900 > He was able to recoup his costs in 3 months Right now, his average income is $25,000–$30,000 per month He is not a genius: he simply found a niche with high demand and created an offershow more

Bober_smart
553,140 views • 12 days ago
Exploring diff visuals with Gen2. I'm amazed at how... it takes a single block image to build neon streets, structures, and entire cityscapes. Anybody else toying with AI tonight⁉️ AI Art Workflow: 1. Crafted starting image in #Midjourney 2. Generated a 4-sec video with #Gen2 (no text prompt) 3. Extended first 4 secs to 18 secs 4. Repeated steps 2-3 twice, each time using the final frame from the previous 18-sec video to start the next 18-secs run 5. Merged all three 18-sec videos in #FinalCut Pro 6. Adjusted speed for fluidity 7. Final version published here is a 54-sec journey distilled into a 16-sec video with 🎶 #AIart #AIArtCommunityshow more

Dave Villalva
47,946 views • 2 years ago
BREAKING: Iran fired ballistic missiles at Tel Aviv overnight... using cluster munitions. Process what that means. A conventional warhead hits one point. A cluster warhead releases dozens of submunitions, each weighing 2 to 5 kilograms of high-explosive fragmentation, scattering across a wide area after re-entry. One missile becomes dozens of weapons. The fire rate has collapsed from 90 missiles per day on Day 1 to approximately 10 on Day 24. Iran has 140 launchers remaining. The mathematics of 10 missiles per day sounds manageable until each missile contains a warhead that multiplies into dozens of independent kill zones across a city. Overnight impacts confirmed in the Tel Aviv and Ramat Gan area. Sirens across central Israel. Kiryat Shmona hit again in the north. Damage to buildings. Injuries reported. Magen David Adom responding. The IDF confirms cluster submunitions were deployed. Footage shows dispersal patterns consistent with Khorramshahr-4 warheads, each carrying 1,500 kilograms at Mach 8 to 16 with MIRV-capable cluster dispensers. Israel’s multi-layered defence is designed to intercept the parent missile before dispersal. Iron Dome handles short and medium range. David’s Sling covers the intermediate gap. Arrow-2 and Arrow-3 engage ballistic threats at altitude. The strategy works when the intercept happens above the dispersal altitude. When it does not, when a missile penetrates the outer layers and reaches its terminal phase, the submunitions scatter and no system can intercept dozens of 2-kilogram bomblets falling across a residential area simultaneously. Iran has lost 70 percent of its launchers. Its fire rate is 89 percent lower than Day 1. Its navy is destroyed. Its air force is gone. And it is still putting cluster munitions over Tel Aviv. That is not desperation. That is adaptation. Iran cannot match the volume of Day 1. So it changed the payload. Fewer missiles, each one carrying more weapons inside it. The fire rate declined. The lethality per round increased. The shift from conventional to cluster warheads is Iran compensating for launcher attrition with warhead complexity. The arithmetic of the war just changed from “how many missiles” to “how many submunitions per missile.” This is what makes the electricity ledger so dangerous. Iran has publicly absorbed strikes on hospitals, schools, and emergency centres without reciprocating against equivalent Israeli civilian infrastructure. It is conserving its 140 remaining launchers for the one target category it has reserved: electricity. If the 5-day power-plant pause collapses Saturday and Trump executes his threat, those 140 launchers will fire cluster-armed Khorramshahr-4s at Gulf and Israeli power stations, desalination plants, and grid nodes. A single cluster warhead detonating over a transformer yard does not damage it. It saturates it with dozens of bomblets that destroy equipment across the entire facility footprint. Cluster munitions are not designed for military targets. They are designed for area denial. And a power grid is the ultimate area target. Saturday is four days away. The launchers are armed with cluster warheads. The ledger is open. The targets are named. The only thing standing between the cluster munitions and the grid is a 5-day pause announced on a social media platform that Iran says does not correspond to any agreement it has made.show more

Shanaka Anslem Perera ⚡
72,348 views • 4 months ago
This guy built a mini AI farm out of... 4 Nvidia boxes It does not look like a data center. It looks like a stack of small machines sitting next to a laptop. But each box is a DGX Spark with Grace Blackwell inside, 128GB unified memory, and enough room to run models normal gaming GPUs cannot even open. Using the launch price from the article, 4 of them is almost $12,000 of local AI compute on one desk. That sounds expensive until you compare it to cloud GPUs. A serious AI builder can burn $1,500 to $3,000 a month renting A100s and H100s for client work, fine-tunes, agents and 70B models. He basically moved that bill from the cloud into hardware he owns. 4 Nvidia boxes. 512GB unified memory. No hourly meter running in the background. No rented GPUs eating the margin every time an agent runs too long. The funny part is most people still think local AI means a slow laptop running a toy model. Meanwhile guys like this are stacking compute at home. Save this, local AI is turning into the new mining farm.show more

Gipp 🦅
590,100 views • 1 month ago
What you’re seeing is likely a Shahab-3 variant, an... IRGC medium-range ballistic missile with a cluster payload, striking central Tel Aviv. Range: ~1,300 km. CEP: ~300 m. Submunitions: up to 1,400. Designed not for precision, but saturation. The cluster head doesn’t need a direct hit, just altitude. At 2–5 km above target, it blooms, raining steel across a 500–800 m footprint. Facades peel open. Cars ignite. Infrastructure cracks. The goal isn’t collapse, it’s disruption: trauma centers overwhelmed, emergency bandwidth flooded, command-and-control slowed. Israel’s Arrow interceptors are built for unitary warheads, not saturation blooms. The Iron Dome isn’t rated for this class. 1 missile is a strike. 2 is a message. Anything more, and you’re watching Tel Aviv’s deterrence assumptions unravel in real time.show more

Thomas Keith
319,855 views • 1 year ago
This Chinese developer launched Llama 70B locally on a... MacBook on a plane and for a full 11 hours without internet ran client projects. He was sitting by the window on a transatlantic flight with a MacBook Pro M4 with 64 GB of memory. WiFi on board cost $25 for the flight. He declined. No cloud API, no connection to Anthropic or OpenAI servers, no internet at all. Just a local Llama 3.3 70B on bf16 and his own orchestrator script. The model runs through llama.cpp. Generation speed, 71 tokens per second. Context around 60,000 tokens. Memory usage, 48.6 GiB out of 64. Battery at takeoff, 3 hours 21 minutes. And he gave the orchestrator this system prompt before takeoff: "You are an offline orchestrator running on a single MacBook. There is no network. The only resources you have are local files in /Users/dev/work, the Llama 70B inference server at localhost:8080, and a battery budget of 3 hours 21 minutes. Process the queue at /Users/dev/work/queue.jsonl (one client task per line). For each task: draft → run local evals → save artefact to /Users/dev/work/done/. Save context checkpoints every 12 tasks so you can resume after a battery swap. Stop only on empty queue or when battery drops below 5%." So the system knows exactly what resources it is running on. It knows it has no connection to the outside world for the next 11 hours. It knows it has finite memory and a finite battery. It knows the human will not intervene until the plane lands. The system runs in 1 loop. Takes a task from the queue, runs it through inference, saves the artifact, writes a checkpoint. Task after task, just like that. And only when the battery drops below 5% does the orchestrator automatically pause, waits for the laptop to switch to the backup power bank, and continues from the last checkpoint. Here is what the system actually writes in his log during the flight: "saved context checkpoint 8 of 12 (pos_min = 488, pos_max = 50118, size = 62.813 MiB)" "restored context checkpoint (pos_min = 488, pos_max = 50118)" "prompt processing progress: n_tokens = 50 / 60 818" "task 37016 done | tps = 71 s tokens text → /Users/dev/work/done/proposal_westside.md" Outside the window, clouds, blue sky, and no WiFi. On the tray, 1 MacBook, an open terminal on 2 screens, and an inference server on localhost. From what I have observed, this is the cleanest offline AI workflow I have seen in the past year: 11 hours of flight, $0 for WiFi, and the entire client queue closed before landing.show more

Blaze
1,839,572 views • 2 months ago