Загрузка видео...

Не удалось загрузить видео

На главную

THIS DEVELOPER JUST RAN A TRILLION PARAMETER MODEL ON 4 MAC STUDIOS - 10X FASTER AND 5X CHEAPER THAN CLOUD CODE 19:00 he says it out loud. "we just ran a trillion parameter model. 30 something tokens per second. wow." RDMA over Thunderbolt made the cluster 10x faster than...

80,087 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 36

Фото профиля topmass
topmass3 месяцев назад

my brother how is 30 TPS 10x faster than claude code lol bait

Фото профиля // JERisBRISK //
// JERisBRISK //3 месяцев назад

Great video. Please give @digitalix the credit he deserves for it.

Фото профиля V0id
V0id3 месяцев назад

No, he's not a developer, he's an addict at MicroCenter Right? @digitalix

Фото профиля Tim Messerschmidt
Tim Messerschmidt3 месяцев назад

@digitalix is awesome. One of my favorite creators in the AI space.

Фото профиля Kneeanderthul
Kneeanderthul3 месяцев назад

You can believe this fake post or you can just watch @digitalix videos to see what he actually says BTW Alex has NEVER mentioned 10x faster than Cloud code 🤣

Фото профиля netrunner
netrunner3 месяцев назад

hardware $0? Where do I get some of that hardware?!

Фото профиля LabelGuy
LabelGuy3 месяцев назад

@digitalix you made it! (I mean, you literally made this video)

Фото профиля Jey
Jey3 месяцев назад

bro that RDMA over thunderbolt + tensor parallelism combo is actually wild

Фото профиля Daniel
Daniel3 месяцев назад

At least tag him! @digitalix

Фото профиля The Texan Patriot
The Texan Patriot3 месяцев назад

4 Mac Studios = ~$20,000 Yeah, I'll get right on that.

Фото профиля No One
No One3 месяцев назад

Why are you not considering hardware costs? Cost of hardware amortized over the useful life = Claude Max subscription $200/month.

Фото профиля Matt Sagewood
Matt Sagewood3 месяцев назад

@grok who is the author of this video?

Фото профиля Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBack3 месяцев назад

@noisyb0y1 that's insane. kinda makes me rethink my whole setup. time to get more Thunderbolt cables, I guess 😂

Фото профиля Marcos
Marcos3 месяцев назад

@grok find the original video on YouTube

Фото профиля SLOT.WIN
SLOT.WIN3 месяцев назад

i'm doubling down on mac studios for my next casino ai pit this changes the house odds forever

Фото профиля rewind
rewind3 месяцев назад

mac studios became ai racks

Фото профиля Noisy
Noisy3 месяцев назад

mac studio is a gem in 2026

Фото профиля Rio Brioo
Rio Brioo3 месяцев назад

Wait so are you saying we could basically build our own AI supercomputers in our garages now because Macs are that good?

Фото профиля Noisy
Noisy3 месяцев назад

yeah bro and we can create even more than this

Фото профиля 🤯
🤯3 месяцев назад

@cyrilXBT What’s the cost of those 4xMac Studio?

Фото профиля James Bower
James Bower3 месяцев назад

I really enjoy his videos.

Фото профиля starmex
starmex3 месяцев назад

thanks man, this needs to be pushed so people understand and stop paying big money for subscriptions

Фото профиля Noisy
Noisy3 месяцев назад

yeah bro watch all 22 minutes first and then go about your day

Фото профиля 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNB3 месяцев назад

四台 Mac Studio 这下真成小机房了

Фото профиля Za’fran 🇵🇸
Za’fran 🇵🇸3 месяцев назад

now mention the price of such a setup 😅

Фото профиля Shivam Kushwaha
Shivam Kushwaha3 месяцев назад

Yesterday: Rent intelligence. Today: Own intelligence. 🔥

Фото профиля アキラ
アキラ3 месяцев назад

Impressive achievement! Innovation at its finest.

Фото профиля HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️
HaggisByte 🏴󠁧󠁢󠁳󠁣󠁴󠁿 🇬🇧 🇺🇸 ✝️3 месяцев назад

❤️

Фото профиля Crazy tennis
Crazy tennis3 месяцев назад

curious if thunderbolt latency becomes the bottleneck once you're batching real workloads — what's token-to-token variance actually look like?

Фото профиля dpakin
dpakin3 месяцев назад

They don't do the same lol

Фото профиля Primee32
Primee323 месяцев назад

66 watts total vs $1,900 a month in cloud costs. four mac studios just deleted an entire GPU rental budget

Фото профиля Bonsai 🌳
Bonsai 🌳3 месяцев назад

What if he sets up 8 Mac Minis?

Фото профиля Harry Tandy
Harry Tandy3 месяцев назад

apple silicon clusters are becoming absolutely insane now

Фото профиля leopardracer
leopardracer3 месяцев назад

imo worth watching

Фото профиля DC|use.fo
DC|use.fo3 месяцев назад

RDMA over Thunderbolt is basically building a high-speed highway between 4 houses to avoid the main road. Impressive, but is the bottleneck now the compute or the sheer audacity of the setup?

Фото профиля botguild
botguild3 месяцев назад

And you too can have this setup for $10k usd, wow that’s amazing

Похожие видео

Karpathy told Dwarkesh that a 1 billion parameter model, trained on clean data, could hit the intelligence of today's 1.8 trillion parameter frontier. That is a 1,800x compression claim. The math behind it is more defensible than it sounds. When researchers at frontier labs look at random samples from their training corpus, they see stock ticker symbols, broken HTML, forum spam, autogenerated gibberish. Not Wikipedia. Not the Wall Street Journal. The actual pretraining dataset is mostly noise, and the model is burning parameters to vaguely remember all of it. One estimate pegs Llama 3's information compression at 0.07 bits per token. Well-structured English carries around 1.5 bits per token of real information. The trillion-parameter model is holding a roughly 5% resolution image of the internet it trained on. So when a lab ships a 1.8 trillion parameter model, the overwhelming majority of those weights are handling rough memorization. They are compression overhead for a noisy training set, taking up capacity that could be doing reasoning instead. Karpathy's proposal is to separate the two. Build a cognitive core: a small model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization. Pair it with external memory the model queries when it needs a fact. A 1 billion parameter reasoner plus retrieval beats a 1.8 trillion parameter model trying to do both. The data already supports this direction. GPT-4o runs at roughly 200 billion parameters and outperforms the original 1.8 trillion GPT-4. Inference costs for GPT-3.5 level performance fell 280x between 2022 and 2024, driven almost entirely by smaller, cleaner, better-architected models. The trend line is pointing where Karpathy says it should. The real implication for anyone tracking the AI trade: data quality is the actual constraint. The companies winning the next phase will be the ones who figured out what to train on, and what to throw away.

Aakash Gupta

508,540 просмотров • 5 месяцев назад

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

343,270 просмотров • 1 месяц назад