Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

THIS DEVELOPER JUST RAN A TRILLION PARAMETER MODEL ON 4 MAC STUDIOS - 10X FASTER AND 5X CHEAPER THAN CLOUD CODE 19:00 he says it out loud. "we just ran a trillion parameter model. 30 something tokens per second. wow." RDMA over Thunderbolt made the cluster 10x faster than...

80,087 görüntüleme • 3 ay önce •via X (Twitter)

36 Yorum

topmass profil fotoğrafı
topmass3 ay önce

my brother how is 30 TPS 10x faster than claude code lol bait

// JERisBRISK // profil fotoğrafı
// JERisBRISK //3 ay önce

Great video. Please give @digitalix the credit he deserves for it.

V0id profil fotoğrafı
V0id3 ay önce

No, he's not a developer, he's an addict at MicroCenter Right? @digitalix

Tim Messerschmidt profil fotoğrafı
Tim Messerschmidt3 ay önce

@digitalix is awesome. One of my favorite creators in the AI space.

Kneeanderthul profil fotoğrafı
Kneeanderthul3 ay önce

You can believe this fake post or you can just watch @digitalix videos to see what he actually says BTW Alex has NEVER mentioned 10x faster than Cloud code 🤣

netrunner profil fotoğrafı
netrunner3 ay önce

hardware $0? Where do I get some of that hardware?!

LabelGuy profil fotoğrafı
LabelGuy3 ay önce

@digitalix you made it! (I mean, you literally made this video)

Jey profil fotoğrafı
Jey3 ay önce

bro that RDMA over thunderbolt + tensor parallelism combo is actually wild

Daniel profil fotoğrafı
Daniel3 ay önce

At least tag him! @digitalix

The Texan Patriot profil fotoğrafı
The Texan Patriot3 ay önce

4 Mac Studios = ~$20,000 Yeah, I'll get right on that.

No One profil fotoğrafı
No One3 ay önce

Why are you not considering hardware costs? Cost of hardware amortized over the useful life = Claude Max subscription $200/month.

Matt Sagewood profil fotoğrafı
Matt Sagewood3 ay önce

@grok who is the author of this video?

Hussain Hashim | Building SundayBack profil fotoğrafı
Hussain Hashim | Building SundayBack3 ay önce

@noisyb0y1 that's insane. kinda makes me rethink my whole setup. time to get more Thunderbolt cables, I guess 😂

Marcos profil fotoğrafı
Marcos3 ay önce

@grok find the original video on YouTube

SLOT.WIN profil fotoğrafı
SLOT.WIN3 ay önce

i'm doubling down on mac studios for my next casino ai pit this changes the house odds forever

rewind profil fotoğrafı
rewind3 ay önce

mac studios became ai racks

Noisy profil fotoğrafı
Noisy3 ay önce

mac studio is a gem in 2026

Rio Brioo profil fotoğrafı
Rio Brioo3 ay önce

Wait so are you saying we could basically build our own AI supercomputers in our garages now because Macs are that good?

Noisy profil fotoğrafı
Noisy3 ay önce

yeah bro and we can create even more than this

🤯 profil fotoğrafı
🤯3 ay önce

@cyrilXBT What’s the cost of those 4xMac Studio?

James Bower profil fotoğrafı
James Bower3 ay önce

I really enjoy his videos.

starmex profil fotoğrafı
starmex3 ay önce

thanks man, this needs to be pushed so people understand and stop paying big money for subscriptions

Noisy profil fotoğrafı
Noisy3 ay önce

yeah bro watch all 22 minutes first and then go about your day

安叫兽|Bird🕊️ 🔶 BNB profil fotoğrafı
安叫兽|Bird🕊️ 🔶 BNB3 ay önce

四台 Mac Studio 这下真成小机房了

Za’fran 🇵🇸 profil fotoğrafı
Za’fran 🇵🇸3 ay önce

now mention the price of such a setup 😅

Shivam Kushwaha profil fotoğrafı
Shivam Kushwaha3 ay önce

Yesterday: Rent intelligence. Today: Own intelligence. 🔥

アキラ profil fotoğrafı
アキラ3 ay önce

Impressive achievement! Innovation at its finest.

HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️ profil fotoğrafı
HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️3 ay önce

❤️

Crazy tennis profil fotoğrafı
Crazy tennis3 ay önce

curious if thunderbolt latency becomes the bottleneck once you're batching real workloads — what's token-to-token variance actually look like?

dpakin profil fotoğrafı
dpakin3 ay önce

They don't do the same lol

Primee32 profil fotoğrafı
Primee323 ay önce

66 watts total vs $1,900 a month in cloud costs. four mac studios just deleted an entire GPU rental budget

Bonsai 🌳 profil fotoğrafı
Bonsai 🌳3 ay önce

What if he sets up 8 Mac Minis?

Harry Tandy profil fotoğrafı
Harry Tandy3 ay önce

apple silicon clusters are becoming absolutely insane now

leopardracer profil fotoğrafı
leopardracer3 ay önce

imo worth watching

DC|use.fo profil fotoğrafı
DC|use.fo3 ay önce

RDMA over Thunderbolt is basically building a high-speed highway between 4 houses to avoid the main road. Impressive, but is the bottleneck now the compute or the sheer audacity of the setup?

botguild profil fotoğrafı
botguild3 ay önce

And you too can have this setup for $10k usd, wow that’s amazing

Benzer Videolar

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

508,540 görüntüleme • 5 ay önce

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

343,270 görüntüleme • 1 ay önce