
Antid
@antisadh • 2,038 subscribers
studying AI daily | helping you save & make money with it
Videos

NVIDIA CHARGES $30K FOR AN H100 WITH NVLINK. TENSTORRENT SELLS THE SAME 800GBIT/S INTERCONNECT IN A $2K CARD. HE PLUGGED 4 INTO PCIE x1 SLOTS AND RUNS OPUS-5 LOCALLY FOR $6/MONTH buy a tenstorrent card for $2k -> plug into any $50 pcie x1 slot -> connect cards via qsfp-dd 800g -> skip the nvidia pcie bandwidth wall -> pool 100gb+ of vram across cards for less than one h100 -> run opus-5 locally for $6/month. that loop is why every serious solo ai builder is quietly ordering tenstorrent cards and openai's infrastructure team went quiet for 3 days after keller's announcement. tenstorrent accelerator + qsfp-dd 800g interconnect + cheap pcie x1 slots + linux + no nvidia tax - that's the stack. watch and save it, then order a tenstorrent card this weekend.
Antid88,517 views • 26 days ago

ONE BUILDER WIRED 4 RASPBERRY PI MODULES INTO A €450 AI CLUSTER WITH 64GB OF RAM AND KILLED HIS $200/MONTH CHATGPT SUBSCRIPTION 00:37 he points at the terminal, "loaded, loaded, loaded - 4 nodes synced over gigabit, 64 gigabytes total, llama 3.1 partitioned across the cluster" he stacked 4 raspberry pi CM5 lite modules with 16 gigabytes each onto a sipeed nanocluster board, the whole rig runs off a single 65 watt power supply and a gigabit internal switch distributed llama partitions the model across the nodes with synchronized workloads handled over the internal fabric, no quantization needed, the cluster hits about 30 tokens per second on small models a heavy developer pays $200 a month for chatgpt pro and another $200 for claude code, this rig cost €450 in parts and breaks even on the stack in month 2 most people will keep paying anthropic and openai forever, a few will spend a weekend wiring 4 raspberry pi modules and never see a subscription invoice again the window is open, follow and bookmark before it closes
Antid64,291 views • 2 months ago

OPENAI PAYS NVIDIA $30K PER H100. INTEL SELLS THE SAME 32GB VRAM IN A $950 ARC PRO B70. HE STACKED 2 AND RUNS OPUS-5 LOCALLY FOR $3/MONTH buy 2 intel arc pro b70 for $1,900 -> plug them into any dual-x8 pcie motherboard -> boot ubuntu 26.04 -> gpus recognized instantly -> load a q4 moe model like qwen 35b-a3b -> run 64gb of local ai for $3/month in electricity that loop is why nvidia's 5090 monopoly on 32gb vram just quietly cracked and openai's inference costs are about to look ridiculous 2x intel arc pro b70 + asrock taichi lite + ryzen 5 9600 + 64gb ddr5 + ubuntu 26.04 + intel's ipex stack - that's the stack watch and save it, then order your first b70 this weekend
Antid28,739 views • 28 days ago

MICROSOFT JUST BROKE COMPATIBILITY WITH A HOMELAB GUY'S $80 AI RIG. HIS 90-SECOND FIX IS ALREADY IN A GITHUB REPO 4,700 PEOPLE FORKED old ai server -> windows 10 expiring -> microsoft blocks win11 upgrade -> tpm chip missing -> boot linux instead -> flash modded firmware -> keep running 15gb of local ai for $0/month that loop is why microsoft's tpm requirement just accidentally handed linux every serious homelab in 2026 linux + amd bc-250 + segfault firmware + moth enjoyer's docs + 6.8tb pcie ssd - that's the stack watch and save it, then move your ai server off windows this weekend
Antid28,509 views • 29 days ago

NVIDIA'S $30K H100 JUST GOT OUT-INFERENCED BY A $250 RYZEN. 18 TOKENS/SEC ON A 35B MODEL, $0.03 IN ELECTRICITY, $47B GONE FROM NVDA BY FRIDAY download qwen3 35b-a3b -> load in lm studio -> disable gpu entirely -> run on cpu + ram -> hit 18 tokens/sec on a 6-core ryzen -> pay nothing per token. that loop is why moe models just made every consumer cpu a viable ai host and nvidia's inference monopoly is quietly on borrowed time. qwen3 35b-a3b + lm studio + ryzen 5 9600x + 32gb ddr5 + zero gpu - that's the stack. watch and save it, then run a 35b model on your laptop this weekend.
Antid23,745 views • 29 days ago

A $40 BC250 BOARD WITH 16GB GDDR6 GETS A GITHUB FIRMWARE FLASH, THE DEFAULT 8/8 CPU/GPU SPLIT REWRITES TO 0.5/15.5 AND OPENS 15.5GB OF UNIFIED MEMORY TO OLLAMA, DEAD CRYPTO HARDWARE JUST DOUBLED ITS AI CEILING 01:08 the operator points at the chart, "you only give 512 megs to the GPU part and that's reserved, that means all the rest of the RAM is available to either the CPU or the GPU" a BC250 ships from the factory with 16GB GDDR6 split 8/8 between Oberon CPU cores and the RDNA 2 iGPU, the iGPU only ever uses 4-5GB during inference, the other 3-4GB sits stranded SEC Bolt's firmware mod on the moth-enjoyer GitHub flashes a 0.5/15.5 split, the GPU reserves 512MB and 15.5GB falls into a unified pool that ollama treats as VRAM, qwen 3.6 14B at 4 bit fits with 2GB context headroom 15.5GB on a $40 board with a $40 CH347 flasher beats the $249 Jetson Orin Nano's 8GB by nearly 2x at 1/6th the price, the firmware is open source, the dump backup process takes 4 minutes per board your map's tier zero floor was the OptiPlex at $35-50 with iGPU only, the modded BC250 sits at the same $40 mark with 15.5GB of GDDR6 and an actual RDNA 2 GPU, the buyer who flashes once unlocks a 13B class local AI host for the price of a dinner the window is open, follow and bookmark before it closes
Antid25,619 views • 1 month ago

A 24 BAY DELL POWEREDGE T550 TOWER SERVER FITS IN A HOME OFFICE WITHOUT A RACK, 384TB OF RAW SAS STORAGE HOSTS EVERY OPEN SOURCE LLM EVER RELEASED LOCALLY, THE HOME AI LIBRARY TIER UNDER YOUR MAP 00:25 the reseller spins the chassis around, "it's a 24 bay Dell T550 tower server man, look it's got a boss card for dual NVMe boot SSDs" a refurbished T550 ships at $900-1,100 with single Xeon Silver, 64GB DDR4 ECC and 24 hot swap SAS bays in a tower form factor that fits next to a desk without a server rack 24x 16TB SAS drives at $80 each on eBay equals 384TB raw storage for $1,920, the same array on enterprise Mac Studio Pro tier would cost $35,000 in thunderbolt JBOD enclosures Llama 3.3 70B is 42GB, Qwen3-235B is 110GB, DeepSeek-V3 is 100GB, Mistral Large is 80GB, the entire ollama public model library fits inside 4TB, you can mirror every open weights release for the next decade your map covers compute boxes that swap one model at a time, the T550 is the model warehouse tier that feeds them, a homelab operator wires the tower over 10GbE to a mac mini and the mac never waits 30 seconds to pull a swap the window is open, follow and bookmark before it closes
Antid18,526 views • 1 month ago
No more content to load