Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚨 I DON'T KNOW WHAT PEOPLE ARE WAITING FOR! This guy just deployed a full GPU compute farm in a data center Not renting cloud instances, not paying per-token, just racks of consumer GPUs on warehouse shelving doing AI inference 24/7 The setup is wild: open-frame rigs with rows...

16,305 görüntüleme • 3 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

A USED $700 RTX 3090 HAS MORE MEMORY AND FAR MORE SPEED THAN THE ~$1,000 A2 GOING INTO THIS SERVER. THE A2 STILL WINS - FOR ONE REASON THAT ISN'T PERFORMANCE. that clip is a dell poweredge r760 - a 2u rack server - getting an nvidia a2 dropped in through a riser. this is the enterprise route to local ai. the rung the desktop ladder skips entirely. the a2, verified: 16gb of memory, single-slot, low-profile, and just 40 to 60 watts it's an inference card, built to sit in a server that has no room or spare power for a real gpu so why not a 3090? a used rtx 3090 gives you 24gb and many times the compute for around $700. the a2 gives 16gb, much slower, for roughly a thousand. on raw local-llm value, the desktop card wins, and it isn't close. the a2's whole reason to exist is the thing you can't see on a spec sheet: it fits. a 2u server has no spare gpu power and no room for a triple-slot furnace. the a2 slips into one low-profile slot on 60 watts and lets an existing server do inference without a rebuild. the honest caveats: the r760 is not a desk machine. loaded, it idles at hundreds of watts, and its fans sound like a hair dryer that never turns off. it belongs in a rack or a closet 16gb is still 16gb. the same 7b-to-32b ceiling as the cheap cards. the enterprise badge doesn't buy you a bigger model what it does buy: ecc memory, redundant power, remote management, hot-swap everything. reliability, not speed so who this is for: someone who already runs a rack and wants to add private inference without touching the power or cooling budget. for anyone starting from a desk, the mac mini or the 3090 wins on every axis that matters. the real point: the "best" local ai box depends entirely on what you already own. on a desk, efficiency wins. in a rack, fit and reliability win. the a2 isn't a bad card - it's a card for a constraint the ladder never mentions. no fastest gpu here, no desk-friendly box, no bigger model than the cheap cards already run. save this before you buy enterprise gear for a desktop job.

Grimmer

24,287 görüntüleme • 2 ay önce

Demis Hassabis just explained why the real AI bottleneck has nothing to do with training runs. Most people picture the AI arms race as who can build the biggest model. GPT-4 or Gemini Ultra style training runs, a few hundred million in compute, fired once or twice a year. The constraint sits somewhere else. Every time a researcher has a new algorithmic idea, a new architecture, a new training technique, they can't just test it on a laptop. They have to run it at the scale where it would actually be deployed, because ideas that look promising at small scale fall apart completely when you put them into a real system. Every research hypothesis burns significant compute before a single line of production code gets written. At a lab like DeepMind, hundreds of researchers are running hundreds of ideas simultaneously. The demand for experimental compute is continuous. It never stops. Now layer the hardware reality on top. GPU lead times are currently 36 to 52 weeks for data center hardware. Global AI data centers are already drawing 29.6 gigawatts, equivalent to the peak power demand of the entire state of New York, and they still can't meet demand. Companies willing to pay any price can't just buy more compute. They wait in line. The speed of scientific discovery in AI is now gated by hardware availability. The next breakthrough is sitting in a researcher's head right now. Whether it gets validated fast enough to matter depends entirely on whether the compute is there when they need it. The AI race gets won by whoever can run the most experiments per month.

Aakash Gupta

32,150 görüntüleme • 5 ay önce

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

343,270 görüntüleme • 1 ay önce

this is what 12 gigs of VRAM built in 2026. a 9 billion parameter model running on a 5 year old RTX 3060 wrote a full space shooter from a single prompt. blank screen on first try. i came back with a bug list and the same model on the same card fixed every issue across 11 files without touching a single line myself. enemies still looked wrong so i pushed another iteration and now the game has pixel art octopi, particle effects, screen shake, projectile physics and a combo system. all running locally on a card that was designed to play fortnite. three iterations. zero cloud. zero API calls. every token generated on hardware sitting under my desk. the model reads its own code, finds what's broken, patches it, validates syntax and restarts the server. i just describe what's wrong and it handles the rest. people are paying monthly subscriptions to type into a browser tab and wait for a server farm to respond. meanwhile a GPU you can find used on ebay is running a full autonomous hermes agent framework with 31 tools, 128K context window and thinking mode generating at 29 tokens per second nonstop. the game still needs work. level upgrades don't trigger and boss fights need tuning. but the fact that i'm iterating on gameplay balance instead of debugging whether the code runs at all tells you where this is headed. every iteration the game gets better on the same hardware. same 12 gigs. same 9 billion parameters. same RTX 3060 from 5 years ago your GPU is not a gaming card anymore. it's a local AI lab that never sends your data anywhere.

Sudo su

170,848 görüntüleme • 6 ay önce