Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

THIS DEVELOPER JUST RAN A TRILLION PARAMETER MODEL ON 4 MAC STUDIOS - 10X FASTER AND 5X CHEAPER THAN CLOUD CODE 19:00 he says it out loud. "we just ran a trillion parameter model. 30 something tokens per second. wow." RDMA over Thunderbolt made the cluster 10x faster than...

80,087 Aufrufe • vor 3 Monaten •via X (Twitter)

36 Kommentare

Profilbild von topmass
topmassvor 3 Monaten

my brother how is 30 TPS 10x faster than claude code lol bait

Profilbild von // JERisBRISK //
// JERisBRISK //vor 3 Monaten

Great video. Please give @digitalix the credit he deserves for it.

Profilbild von V0id
V0idvor 3 Monaten

No, he's not a developer, he's an addict at MicroCenter Right? @digitalix

Profilbild von Tim Messerschmidt
Tim Messerschmidtvor 3 Monaten

@digitalix is awesome. One of my favorite creators in the AI space.

Profilbild von Kneeanderthul
Kneeanderthulvor 3 Monaten

You can believe this fake post or you can just watch @digitalix videos to see what he actually says BTW Alex has NEVER mentioned 10x faster than Cloud code 🤣

Profilbild von netrunner
netrunnervor 3 Monaten

hardware $0? Where do I get some of that hardware?!

Profilbild von LabelGuy
LabelGuyvor 3 Monaten

@digitalix you made it! (I mean, you literally made this video)

Profilbild von Jey
Jeyvor 3 Monaten

bro that RDMA over thunderbolt + tensor parallelism combo is actually wild

Profilbild von Daniel
Danielvor 3 Monaten

At least tag him! @digitalix

Profilbild von The Texan Patriot
The Texan Patriotvor 3 Monaten

4 Mac Studios = ~$20,000 Yeah, I'll get right on that.

Profilbild von No One
No Onevor 3 Monaten

Why are you not considering hardware costs? Cost of hardware amortized over the useful life = Claude Max subscription $200/month.

Profilbild von Matt Sagewood
Matt Sagewoodvor 3 Monaten

@grok who is the author of this video?

Profilbild von Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBackvor 3 Monaten

@noisyb0y1 that's insane. kinda makes me rethink my whole setup. time to get more Thunderbolt cables, I guess 😂

Profilbild von Marcos
Marcosvor 3 Monaten

@grok find the original video on YouTube

Profilbild von SLOT.WIN
SLOT.WINvor 3 Monaten

i'm doubling down on mac studios for my next casino ai pit this changes the house odds forever

Profilbild von rewind
rewindvor 3 Monaten

mac studios became ai racks

Profilbild von Noisy
Noisyvor 3 Monaten

mac studio is a gem in 2026

Profilbild von Rio Brioo
Rio Brioovor 3 Monaten

Wait so are you saying we could basically build our own AI supercomputers in our garages now because Macs are that good?

Profilbild von Noisy
Noisyvor 3 Monaten

yeah bro and we can create even more than this

Profilbild von 🤯
🤯vor 3 Monaten

@cyrilXBT What’s the cost of those 4xMac Studio?

Profilbild von James Bower
James Bowervor 3 Monaten

I really enjoy his videos.

Profilbild von starmex
starmexvor 3 Monaten

thanks man, this needs to be pushed so people understand and stop paying big money for subscriptions

Profilbild von Noisy
Noisyvor 3 Monaten

yeah bro watch all 22 minutes first and then go about your day

Profilbild von 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNBvor 3 Monaten

四台 Mac Studio 这下真成小机房了

Profilbild von Za’fran 🇵🇸
Za’fran 🇵🇸vor 3 Monaten

now mention the price of such a setup 😅

Profilbild von Shivam Kushwaha
Shivam Kushwahavor 3 Monaten

Yesterday: Rent intelligence. Today: Own intelligence. 🔥

Profilbild von アキラ
アキラvor 3 Monaten

Impressive achievement! Innovation at its finest.

Profilbild von HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️
HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️vor 3 Monaten

❤️

Profilbild von Crazy tennis
Crazy tennisvor 3 Monaten

curious if thunderbolt latency becomes the bottleneck once you're batching real workloads — what's token-to-token variance actually look like?

Profilbild von dpakin
dpakinvor 3 Monaten

They don't do the same lol

Profilbild von Primee32
Primee32vor 3 Monaten

66 watts total vs $1,900 a month in cloud costs. four mac studios just deleted an entire GPU rental budget

Profilbild von Bonsai 🌳
Bonsai 🌳vor 3 Monaten

What if he sets up 8 Mac Minis?

Profilbild von Harry Tandy
Harry Tandyvor 3 Monaten

apple silicon clusters are becoming absolutely insane now

Profilbild von leopardracer
leopardracervor 3 Monaten

imo worth watching

Profilbild von DC|use.fo
DC|use.fovor 3 Monaten

RDMA over Thunderbolt is basically building a high-speed highway between 4 houses to avoid the main road. Impressive, but is the bottleneck now the compute or the sheer audacity of the setup?

Profilbild von botguild
botguildvor 3 Monaten

And you too can have this setup for $10k usd, wow that’s amazing

Ähnliche Videos

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

508,540 Aufrufe • vor 5 Monaten

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

343,270 Aufrufe • vor 1 Monat