正在加载视频...

视频加载失败

THIS DEVELOPER JUST RAN A TRILLION PARAMETER MODEL ON 4 MAC STUDIOS - 10X FASTER AND 5X CHEAPER THAN CLOUD CODE 19:00 he says it out loud. "we just ran a trillion parameter model. 30 something tokens per second. wow." RDMA over Thunderbolt made the cluster 10x faster than...

80,087 次观看 • 3 个月前 •via X (Twitter)

36 条评论

topmass 的头像
topmass3 个月前

my brother how is 30 TPS 10x faster than claude code lol bait

// JERisBRISK // 的头像
// JERisBRISK //3 个月前

Great video. Please give @digitalix the credit he deserves for it.

V0id 的头像
V0id3 个月前

No, he's not a developer, he's an addict at MicroCenter Right? @digitalix

Tim Messerschmidt 的头像
Tim Messerschmidt3 个月前

@digitalix is awesome. One of my favorite creators in the AI space.

Kneeanderthul 的头像
Kneeanderthul3 个月前

You can believe this fake post or you can just watch @digitalix videos to see what he actually says BTW Alex has NEVER mentioned 10x faster than Cloud code 🤣

netrunner 的头像
netrunner3 个月前

hardware $0? Where do I get some of that hardware?!

LabelGuy 的头像
LabelGuy3 个月前

@digitalix you made it! (I mean, you literally made this video)

Jey 的头像
Jey3 个月前

bro that RDMA over thunderbolt + tensor parallelism combo is actually wild

Daniel 的头像
Daniel3 个月前

At least tag him! @digitalix

The Texan Patriot 的头像
The Texan Patriot3 个月前

4 Mac Studios = ~$20,000 Yeah, I'll get right on that.

No One 的头像
No One3 个月前

Why are you not considering hardware costs? Cost of hardware amortized over the useful life = Claude Max subscription $200/month.

Matt Sagewood 的头像
Matt Sagewood3 个月前

@grok who is the author of this video?

Hussain Hashim | Building SundayBack 的头像
Hussain Hashim | Building SundayBack3 个月前

@noisyb0y1 that's insane. kinda makes me rethink my whole setup. time to get more Thunderbolt cables, I guess 😂

Marcos 的头像
Marcos3 个月前

@grok find the original video on YouTube

SLOT.WIN 的头像
SLOT.WIN3 个月前

i'm doubling down on mac studios for my next casino ai pit this changes the house odds forever

rewind 的头像
rewind3 个月前

mac studios became ai racks

Noisy 的头像
Noisy3 个月前

mac studio is a gem in 2026

Rio Brioo 的头像
Rio Brioo3 个月前

Wait so are you saying we could basically build our own AI supercomputers in our garages now because Macs are that good?

Noisy 的头像
Noisy3 个月前

yeah bro and we can create even more than this

🤯 的头像
🤯3 个月前

@cyrilXBT What’s the cost of those 4xMac Studio?

James Bower 的头像
James Bower3 个月前

I really enjoy his videos.

starmex 的头像
starmex3 个月前

thanks man, this needs to be pushed so people understand and stop paying big money for subscriptions

Noisy 的头像
Noisy3 个月前

yeah bro watch all 22 minutes first and then go about your day

安叫兽|Bird🕊️ 🔶 BNB 的头像
安叫兽|Bird🕊️ 🔶 BNB3 个月前

四台 Mac Studio 这下真成小机房了

Za’fran 🇵🇸 的头像
Za’fran 🇵🇸3 个月前

now mention the price of such a setup 😅

Shivam Kushwaha 的头像
Shivam Kushwaha3 个月前

Yesterday: Rent intelligence. Today: Own intelligence. 🔥

アキラ 的头像
アキラ3 个月前

Impressive achievement! Innovation at its finest.

HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️ 的头像
HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️3 个月前

❤️

Crazy tennis 的头像
Crazy tennis3 个月前

curious if thunderbolt latency becomes the bottleneck once you're batching real workloads — what's token-to-token variance actually look like?

dpakin 的头像
dpakin3 个月前

They don't do the same lol

Primee32 的头像
Primee323 个月前

66 watts total vs $1,900 a month in cloud costs. four mac studios just deleted an entire GPU rental budget

Bonsai 🌳 的头像
Bonsai 🌳3 个月前

What if he sets up 8 Mac Minis?

Harry Tandy 的头像
Harry Tandy3 个月前

apple silicon clusters are becoming absolutely insane now

leopardracer 的头像
leopardracer3 个月前

imo worth watching

DC|use.fo 的头像
DC|use.fo3 个月前

RDMA over Thunderbolt is basically building a high-speed highway between 4 houses to avoid the main road. Impressive, but is the bottleneck now the compute or the sheer audacity of the setup?

botguild 的头像
botguild3 个月前

And you too can have this setup for $10k usd, wow that’s amazing

相关视频

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

508,540 次观看 • 5 个月前

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

343,270 次观看 • 1 个月前