Загрузка видео...

Не удалось загрузить видео

На главную

THIS CHINESE BUILDER TURNED 8 RETIRED TESLA P40s INTO A 192GB PRIVATE AI SERVER THAT CAN WORK 5,760 GPU-HOURS EVERY MONTH. 00:06 he holds up one used Tesla P40 while seven more sit behind it, then installs the full stack inside a massive dual-Xeon server chassis. each card carries...

42,332 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

A 24-YEAR-OLD CHINESE DEVELOPER FROM HANGZHOU TURNED RTX 4090 / 3090-CLASS GPU RACKS INTO HIS OWN PRIVATE AI CLOUD. HIS $740/MONTH AI BILL DROPPED TO $31 IN ELECTRICITY he got tired of paying for chatgpt, claude, cursor, openai api credits and every “pro” tool that quietly turns into another monthly tax. long context runs, codebase scans, document parsing, agent loops. every workflow ended with a new invoice so he built a local llm rack instead. used server hardware, RTX 4090 / 3090-class GPU boxes, ollama for automation, lm studio for testing models, llama.cpp for heavier local runs. around $6,200 upfront, but after that the cost is mostly power and maintenance now his scripts hit localhost instead of a cloud api. code reviews, private docs, chinese contracts, sql cleanup, support replies and research tasks stay inside the room. no token panic, no rate-limit wall, no sensitive files leaving his own machines the funny part is that he did not replace claude completely. he just stopped using frontier models for dumb volume work. 65% of daily ai tasks do not need the smartest model alive. they need cheap tokens, privacy and a machine that can run all night cloud ai is still the brain. local ai is the engine room. once he separated those two, his monthly ai stack stopped looking like subscriptions and started looking like infrastructure by 2027, owning your own local ai rack will not look extreme. it will look like the moment people realized renting intelligence forever was the expensive option.

Gipp 🦅

21,320 просмотров • 2 месяцев назад

AMD might have disrupted Nvidia's entire cloud GPU rental business. In January at CES, AMD CEO Lisa Su demonstrated a $1,499 mini PC running the same class of AI model that currently costs companies $2,500 to $3,000 every month to rent from Nvidia-powered cloud servers. AMD's own branded version opened pre-orders this month at $3,999. Third party manufacturers have been selling the same chip since 2025 starting at $1,499. Here is exactly why this is dangerous for Nvidia. Nvidia's $75 billion quarterly revenue is built almost entirely on one business model, companies rent access to Nvidia GPUs through cloud providers like AWS and Lambda Labs to run AI. They pay monthly. Nvidia gets paid every time someone runs an AI model in the cloud. That recurring rental income is what turned Nvidia into a $5 trillion company. The AMD box eliminates that monthly fee permanently. One AI consultant switched from $2,800 per month in Nvidia cloud rental costs to $8 per month in electricity. The hardware paid for itself in 11 days. Over 8 months he generated $47,000 running the same AI workloads that previously left him paying Nvidia's ecosystem $2,800 every single month. Multiply that across thousands of enterprise customers and the revenue erosion becomes structural. Every business that buys this box stops paying cloud rental fees forever. Lawyers, doctors, banks, accountants, and financial advisors, businesses with sensitive data that cannot legally go to a cloud server represent billions in annual cloud GPU fees that Nvidia is now at risk of losing permanently. The threat is also closing in from the top. Google signed deals worth tens of billions with Anthropic and Meta to replace Nvidia with its own chips. Amazon built its own AI chips across AWS. Apple trained its AI on Google's chips, not Nvidia's. Custom silicon has grown from 21% of the AI chip market in 2025 to 28% in 2026. Nvidia's rental model only worked because serious AI compute had no alternative.

Bull Theory

26,765 просмотров • 2 месяцев назад

The creator of High Bandwidth Memory (HBM) put a number on the AI build that should stop every infra investor cold. A cluster of a million GPUs runs at roughly 10-20% utilization (Save this). Kim Jung-ho spent thirty years building what feeds the GPU, and his claim is that the GPU is barely working. Here is what is actually happening. Every time a model generates output, the data has to be read out of memory, computed, and written back. The read and the write swallow almost the entire cycle. While that data moves, the GPU does nothing. It sits there, fully powered, fully paid for, waiting. By Kim's estimate the memory is doing only about 30 percent of the work it needs to do. The processor idles the rest. So a million installed GPUs run at 10 to 20 percent. You are not compute constrained. You are memory constrained, and the expensive part is standing around. Adding more GPUs does not fix this. It gives you more processors starving for the same data. Here is the part that decides the next decade. Memory can grow. When a cell cannot shrink any further, you stack it into a high-rise, layer on layer. A GPU cannot be stacked. It runs too hot and needs a cooler bolted to its back, so the one move that rescues memory is closed to the processor. The thing that can keep stacking compounds. The thing that cannot plateaus. The marginal dollar in an AI build now buys more by fixing the memory path than by bolting on another idle GPU. Which is why the companies that control memory bandwidth and supply are not suppliers to the AI trade. They are the AI trade.

Fireside Alpha

38,370 просмотров • 1 месяц назад

NVIDIA quietly built two desktop boxes that delete a $25,000/year AI subscription bill You don't rewrite your stack, you don't rent another data center, you just plug both into the wall and switch one line of code One looks like a deck of cards, the other like a hardback novel, together they replace ChatGPT Plus, Claude Pro, Cursor Pro, the OpenAI API meter, and every cloud GPU you were renting for fine-tunes It's built on the same CUDA stack the data center runs, which means once you migrate one workflow the rest follow on the same code path The reason NVIDIA shipped this is simple The bigger you scale on cloud AI, the harder you get taxed, and a one-person operator paying $2,100/month is producing exactly $0 of asset value at the end of every month And their solution is to skip the rental meter entirely, push inference back onto your desk, and let you loop agents overnight without watching a number tick on someone else's invoice This is much cheaper, faster, and pays itself back in 6 weeks for anyone already running AI for work But there is still a question nobody has answered yet, what happens when the next frontier model drops and your local 70B falls 6 months behind mid-quarter Also, technically a stack of four of the big box runs a 1.6 trillion parameter model on a desk for under $12,000 Even a fraction of that compute is more than most people will ever need in a year Bookmark this, it's worth coming back to when you have time 👇

ZEUS⚡️

65,369 просмотров • 2 месяцев назад