Loading video...

Video Failed to Load

Go Home

THIS DEVELOPER JUST RAN A TRILLION PARAMETER MODEL ON 4 MAC STUDIOS - 10X FASTER AND 5X CHEAPER THAN CLOUD CODE 19:00 he says it out loud. "we just ran a trillion parameter model. 30 something tokens per second. wow." RDMA over Thunderbolt made the cluster 10x faster than...

80,087 views • 3 months ago •via X (Twitter)

36 Comments

topmass's profile picture
topmass3 months ago

my brother how is 30 TPS 10x faster than claude code lol bait

// JERisBRISK //'s profile picture
// JERisBRISK //3 months ago

Great video. Please give @digitalix the credit he deserves for it.

V0id's profile picture
V0id3 months ago

No, he's not a developer, he's an addict at MicroCenter Right? @digitalix

Tim Messerschmidt's profile picture
Tim Messerschmidt3 months ago

@digitalix is awesome. One of my favorite creators in the AI space.

Kneeanderthul's profile picture
Kneeanderthul3 months ago

You can believe this fake post or you can just watch @digitalix videos to see what he actually says BTW Alex has NEVER mentioned 10x faster than Cloud code 🤣

netrunner's profile picture
netrunner3 months ago

hardware $0? Where do I get some of that hardware?!

LabelGuy's profile picture
LabelGuy3 months ago

@digitalix you made it! (I mean, you literally made this video)

Jey's profile picture
Jey3 months ago

bro that RDMA over thunderbolt + tensor parallelism combo is actually wild

Daniel's profile picture
Daniel3 months ago

At least tag him! @digitalix

The Texan Patriot's profile picture
The Texan Patriot3 months ago

4 Mac Studios = ~$20,000 Yeah, I'll get right on that.

No One's profile picture
No One3 months ago

Why are you not considering hardware costs? Cost of hardware amortized over the useful life = Claude Max subscription $200/month.

Matt Sagewood's profile picture
Matt Sagewood3 months ago

@grok who is the author of this video?

Hussain Hashim | Building SundayBack's profile picture
Hussain Hashim | Building SundayBack3 months ago

@noisyb0y1 that's insane. kinda makes me rethink my whole setup. time to get more Thunderbolt cables, I guess 😂

Marcos's profile picture
Marcos3 months ago

@grok find the original video on YouTube

SLOT.WIN's profile picture
SLOT.WIN3 months ago

i'm doubling down on mac studios for my next casino ai pit this changes the house odds forever

rewind's profile picture
rewind3 months ago

mac studios became ai racks

Noisy's profile picture
Noisy3 months ago

mac studio is a gem in 2026

Rio Brioo's profile picture
Rio Brioo3 months ago

Wait so are you saying we could basically build our own AI supercomputers in our garages now because Macs are that good?

Noisy's profile picture
Noisy3 months ago

yeah bro and we can create even more than this

🤯's profile picture
🤯3 months ago

@cyrilXBT What’s the cost of those 4xMac Studio?

James Bower's profile picture
James Bower3 months ago

I really enjoy his videos.

starmex's profile picture
starmex3 months ago

thanks man, this needs to be pushed so people understand and stop paying big money for subscriptions

Noisy's profile picture
Noisy3 months ago

yeah bro watch all 22 minutes first and then go about your day

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB3 months ago

四台 Mac Studio 这下真成小机房了

Za’fran 🇵🇸's profile picture
Za’fran 🇵🇸3 months ago

now mention the price of such a setup 😅

Shivam Kushwaha's profile picture
Shivam Kushwaha3 months ago

Yesterday: Rent intelligence. Today: Own intelligence. 🔥

アキラ's profile picture
アキラ3 months ago

Impressive achievement! Innovation at its finest.

HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️'s profile picture
HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️3 months ago

❤️

Crazy tennis's profile picture
Crazy tennis3 months ago

curious if thunderbolt latency becomes the bottleneck once you're batching real workloads — what's token-to-token variance actually look like?

dpakin's profile picture
dpakin3 months ago

They don't do the same lol

Primee32's profile picture
Primee323 months ago

66 watts total vs $1,900 a month in cloud costs. four mac studios just deleted an entire GPU rental budget

Bonsai 🌳's profile picture
Bonsai 🌳3 months ago

What if he sets up 8 Mac Minis?

Harry Tandy's profile picture
Harry Tandy3 months ago

apple silicon clusters are becoming absolutely insane now

leopardracer's profile picture
leopardracer3 months ago

imo worth watching

DC|use.fo's profile picture
DC|use.fo3 months ago

RDMA over Thunderbolt is basically building a high-speed highway between 4 houses to avoid the main road. Impressive, but is the bottleneck now the compute or the sheer audacity of the setup?

botguild's profile picture
botguild3 months ago

And you too can have this setup for $10k usd, wow that’s amazing

Related Videos

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

508,540 views • 5 months ago

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

343,270 views • 1 month ago