Wide Expert Parallelism increases the total memory bandwidth available... per MoE deployment. This means the model distributes the MoE expert weights across multiple GPUs, so each GPU only needs to load a tiny fraction of the weights. This translates to higher throughput per GPU, increasing perf per dollar and perf per watt.show more

SemiAnalysis
30,675 просмотров • 2 месяцев назад
MTP speedup Qwen by 2.5x in Atomic Chat Dense... vs MoE models on 2x RTX 5090 Qwen3.6 27B: 51 → 117 tps +137% Qwen3.6 35B-A3B: 218 → 267 tps +25% MTP drafts several tokens ahead and verifies them in one pass. The speedup depends on memory moved per pass. Dense 27B reads all 27B params per token, MoE 35B-A3B only reads 3B active. Dense had way more to save by batching. The baseline tps also differ (218 vs 51) for the same reason from the other side. Token generation is memory-bandwidth bound, and MoE moves ~8x less memory per token, so its baseline is already 4x ahead. ~80% draft acceptance. Zero accuracy loss. ~1 GB extra VRAM. Open-source code and local AI app – in the comments 👇show more

atomic.chat
171,139 просмотров • 3 месяцев назад
167 tok/s on a single RTX 4090. FreeToken just... pushed Qwen3.6-35B-A3B NVFP4 into ridiculous territory. → RTX 4090 24GB → 35B total parameters → ~3B active per token → NVFP4 → 167 tok/s decode → 23.55GB peak VRAM → No speculative decoding → No MTP → No draft model That last part is what makes this interesting. The speed isn’t coming from guessing future tokens. FreeToken is exploiting the MoE architecture directly, keeping the active experts on the GPU while using bandwidth-aware execution for the rest. A 35B model with only ~3B active parameters running at 167 tok/s on a consumer 4090 is exactly why MoE models are becoming so compelling for local AI.show more

FHILY👑
33,058 просмотров • 13 дней назад
GLM 5.2 INPUT FELL FROM $1.40 TO SEVEN CENTS... PER MILLION IN NINETY DAYS • what it costs now > $0.07 per million input at the cheapest of 20 providers, a 95% drop in three months. > direct still lists $1.40 in and $4.40 out, with cached input at $0.26. > Same model, same weights, twenty-fold spread depending on the door you walk through. • the free way in > New accounts on get 20M tokens: > 744B MoE, 1M context, MIT weights you can also just download and self-host. Nobody announced this. It happened one provider at a time -> and the model itself never changed. Check which provider you are actually routed through before you top up anywhere ↓show more

slash1s
30,117 просмотров • 27 дней назад
Native time-tracking in Notion This setup lets you track... multiple sessions per task, and shows the total time you've worked across all of them.show more

Thomas Frank
12,480 просмотров • 1 год назад
This is what happens when 48GB of VRAM isn't... enough. Someone put 4× NVIDIA RTX A6000 inside an HP Z8 Fury G5. That gives you 192GB of VRAM across four GPUs. But here's the interesting part: You don't magically get one 192GB GPU. The memory is split across four cards, and the software has to manage how the model is distributed between them. That's why these machines are interesting for local AI. The problem isn't always compute. Sometimes you just need somewhere to put the model. Would you build this today, or go for a newer GPU with more bandwidth?show more

KyzoroX
80,805 просмотров • 3 дней назад
AN AWS ENGINEER QUIETLY BUILT A 2 PETABYTE HOME... SERVER FOR $9/MONTH THAT KILLS A $3,400/MONTH CLOUD STORAGE BILL the lenovo thinkstation pgx ships nvidia's gb10 grace blackwell superchip and 128gb of unified memory in a box the size of a mac mini at 1.2kg it runs an 80b qwen3 coder model at 25 to 40 tokens per second and a 196b step-3.5-flash moe model at 20 tokens per second locally the gb10 packs 6,144 cuda cores, 192 fifth-generation tensor cores and rates at 1 petaflop of fp4 with sparsity from a single 240 watt usb-c power supply fine tuning qwen 2.5 7b with lora took 18 minutes and 41gb of unified memory while the gpu pulled 65 watts and peaked at 77 degrees the box pulls a docker container from nvidia's registry and serves a frontier model on your local network with tool calling and zero data leaving your desk bookmark this and read the article belowshow more

starmex
193,226 просмотров • 2 месяцев назад
this is the worst local AI will ever be.... tomorrow it gets faster. next month the models get smarter. next year your GPU runs what a data center runs today. Qwen3.5-35B-A3B on a single 3090. told it to visualize its own expert routing. 256 experts, 8 active per token, rendered in 3D on the same GPU running inference. no API key. no subscription. no permission needed. closed AI isn't losing ground. it's losing the argument.show more

Sudo su
106,822 просмотров • 6 месяцев назад
Newborns are so tiny. Most full term infants weigh... only 6-8 pounds at birth. But cherish (and document) these special moments while they last, because you won’t have a little baby for long. In fact, children’s growth and development in the early years is positively exponential. Over the course of a single year, infants transform from tiny and helpless to increasingly verbal, mobile toddlers, typically tripling their birth weights. So quickly is your baby growing, in fact, that between their second and sixth months many increase their height by up to a quarter inch PER WEEK (1 inch per month). You’ll only have a tiny baby for the blink of an eye. So make each moment count. This happy little guy (1 week) was shared to TT by tai.vieira.21.show more

Dan Wuori
203,718 просмотров • 2 лет назад
Run Gemma 4 26B MoE on 8GB VRAM with... 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware. Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card. The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?" Today, I’m delivering exactly that. I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!. If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed. The performance metrics are astonishing: - 20 tokens/sec flat decode throughput. - Stable, flat decode speed even with massive prompts. - I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame. # What about prefill? Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable. And this is running completely without Multi Token Prediction (MTP) active. How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4. The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse. # The Test Setup: CPU: Intel Core i7 RAM: 16GB System RAM GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) # The Secret Sauce (The -cmoe Flag) To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp. This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache. It prevents VRAM spillage and holds the throughput rock solid. # The flags: -m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking. Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the repliesshow more

Alok
292,770 просмотров • 3 месяцев назад
"Pakistan is the only country that has good relations... with US, Russia & China. Muslim world is giving great importance to PAK" So as per Arfa & her expert, Beggar nation Pakistan is an emerging power!! Imagine saying this without laughing 😂show more

BALA
301,541 просмотров • 5 месяцев назад
5 days ago it took 2 GPUs to build... this. today it takes 1. same prompt. same particle simulation. completely different model. Qwen-Coder-Next (80B) on 2x 3090s. 46 tok/s. 564 lines. 2 iterations to get it working. 48GB VRAM across two cards just to hold it. Qwen3.5-35B-A3B on a single 3090. 112 tok/s. 461 lines. first try. cleaner code, fewer lines, better structured. 19.7GB on disk with 4GB VRAM to spare. half the parameters. one GPU instead of two. 2.4x faster. and the output actually improved. this is what happens when architecture catches up to ambition. Gated Delta Networks(Mamba2 variant) hybrid with sparse MoE. 3B active params out of 35B per token. efficiency at the architecture level, not just quantization. the curve isn't flattening. it's steepening.show more

Sudo su
34,624 просмотров • 6 месяцев назад
Gemma 4 26B A4B MoE - 500+ t/s decode... - Single RTX 4090 (24 GB VRAM) - Llama.cpp concurrency 24 - q8 kv cache How many API users can you simultaneously host on a single RTX 4090 (24 GB VRAM) before it crashes? Yesterday, I proved you can host 14 active users using unquantized memory. Today, I used 8 bit KV Cache Quantization to hack the VRAM footprint. I successfully scaled to 24 concurrent users without a single dropped connection. A 71% server capacity boost for free. By adding the -ctk q8_0 -ctv q8_0 flags to llama.cpp, you compress the KV cache context memory from 16 bit to 8 bit. This unlocks massive concurrency limits on Gemma 4 26B (MoE) on a single 24GB consumer GPU. Here is the exact telemetry from pushing 8 bit quantization to its absolute physical edge: # TEST 1: The 24 User Concurrency Max Server Config: 24 slots (np 24) | 4,096 context per slot | 98,304 Total Context Client Load: 24 simultaneous requests (2,000 token prompt per user) Unquantized KV cache for this load requires 28GB+ VRAM (Instant OOM). Quantized to Q8, it allocated safely at 23.35 GB. The C++ engine crunched the entire batch in 28.5 seconds. Decode Speed: 21 t/s (Per User) | 500 t/s (Agg) # TEST 2: The 48 User Queue Overload What happens to a compressed cache during a traffic spike? Server Config: 24 slots (np 24) | 4,096 context per slot | 98,304 Total Context Client Load: 48 simultaneous requests (2k token prompt per user) Zero queue drops. The scheduler flushed and hot swapped the 8 bit memory flawlessly on the fly, completing all 48 users in 66.0 seconds (a perfect 2.3x queue scaling multiplier). Decode Speed: 18 t/s (Per User) | 430 t/s (Agg) # TEST 3: The 8 User RAG Slam Server Config: 8 slots (np 8) | 60,000 context per slot | 480,000 Total Context Client Load: 8 simultaneous requests (30k token prompt per user) It allocated 23.83 GB VRAM and chewed through ~240,000 prefill tokens in 46 seconds under massive memory pressure. Prefill Speed: 6,200 t/s (Agg) Decode Speed: 22 t/s (Per User) | 175 t/s (Agg) # The Engineering Alpha (The Quantization Tradeoff): You gain a massive 71% increase in server capacity, but what do you lose? Compute latency. Because the cache is stored in 8 bit, the GPU's cores have to dequantize the memory back to 16 bit on the fly during every single prefill step. In my unquantized tests yesterday, single slot prefill was hitting ~1,500+ t/s. Today, under the heavy 48-user Q8 load, prefill dropped as low as ~750 t/s. You trade a few seconds of initial prefill latency to essentially double your API hosting capacity. For production high volume SaaS, this is the ultimate unit economics cheat code. Here is the exact command to run a 24 user Q8 continuous batching server on your own single 4090, single 3090 or any 24gb vram rig: ./build/bin/llama-server -m gemma-4-26B-A4B-it.gguf -c 98304 -np 24 -b 2048 -ub 2048 -ngl 99 -fa on -ctk q8_0 -ctv q8_0 --port 8080 (Note: -c 98304 allocates exactly 4,096 tokens of context per user across 24 slots). Hugging Face links to the Unsloth Gemma 4 26B QAT quants along with performance graphs available in the replies. Would you trade 3 seconds of Time To First Token latency to double your active user capacity?show more

Alok
17,465 просмотров • 1 месяц назад
I can’t believe I got to design this dream... project. It’s a musicbox carrousel lamp all in one. The Prince Charming Regal Carrousel popcorn bucket is now available only #magickingdom select popcorn cart. (Limit 2 per person, per transaction. available while supplies last)show more

Joey Chou
61,678 просмотров • 1 год назад
THE CLOUD BILL WAS $14,000 A MONTH. EVERY MONTH.... JUST FOR GPU ACCESS. No ownership. No hardware. Just renting someone else's chips and watching the invoice grow. So this startup did the math and bought 1,000 Mac Mini M4s instead. $599 each. One-time. Total: ~$599,000 upfront. Sounds insane. Until you do the math the other way. $14K/mo in cloud GPUs is $168K a year. In 3.5 years you've burned through $599K and own nothing. The meter just keeps running. These guys own the hardware. Same output. Less power draw. And after year 4, every month of compute is basically free. The M4 was never built for data centers. But its performance per watt is so good that someone looked at a rack of a thousand of them and thought "why not." That's when it clicks. The AI infrastructure race isn't about who has the biggest GPU anymore. It's about who figured out they were overpaying for one. Would you make this switch?show more

Framez
18,004 просмотров • 1 месяц назад
The bulk/cut thing is literally so simple: Spend as... much time as you need getting to 10-12% body-fat one single time Baseline established! After that, simply oscillate between 10-12% body-fat and 15-17% body-fat Never go above the high end Only go below the low end if doing a show/just looking to show off When cutting, aim to lose .5-1% of total bodyweight per week (500-1000 calorie deficit per day) When bulking, aim to gain 1-2 pounds per month or .25-.5 pounds per week on average (100-200 calorie surplus per day) SIMPLE!show more

Dean Turner
144,011 просмотров • 14 дней назад
$IREN "we haven't disclosed the specific amount of GPUs"... 1. 🤮 reminds me of $NBIS 2. Setting a terrible precedent here for future deals 3. Making it purposely difficult, to not let analysts properly value your 2027 revenue 4. Increasing the polarized view on IREN by the market However: "approximately 60MW of air-cooled Blackwells" 1. You typically don't talk about gross capacity in a deployment like this 2. If it would be gross capacity, the GPU hour rate at IT level would be crazy high (at PUE 1.2, $680m / 50 = 13.6m/MW) 3. At 60MW IT load, and ~14kW draw at DGX server level, we can get to ~4,286 DGX systems with 8 GPUs per. 4. Based on this we can conclude that 60MW of IT load can run approximately 34k DGX B300. 5. 34k DGX B300 at $680m/yr, would represent a GPU hour price of $2.28 Now this is the problem with not disclosing your GPU quantity. You purposely make your business model look bad, because by approach, you get to a GPU hour price that would imply a payback period of 4 years, where only the last year of the contract is 100% margin. But of course, we can also take "the glass is half full" approach. IREN has ordered 50K B300s from Dell. They have 2 purchase orders for this, 1 between Dell Canada and IE CA Leasing Ltd for 4 phases, and 1 between Dell USA and IE US Hardware 1 Inc (amended from IE US Hardware 4 Inc on April 27, 2026). The order for Canada is divided in 4 phases, and are going to Mackenzie for 80MW of gross capacity, which happens to be 4 buildings of 20MW. The order for Childress is divided in 2 phases, and are going to DC35 and DC36, (as depicted in the earnings presentation) and those are 50MW gross. The purchase price of the order for Childress was $1.2B, and for Canada it was $2.3B If we go with 50,000 B300s for a total of $3.5B then $1.2 would represent 34.285% of the 50,000 GPUs, or 17,140 B300s rounded down. For this calculation I will consider that $IREN will deploy 17,140 GPUs in 50MW gross capacity in DC35 and DC36 of block 3 in Childress.. That would imply at 1.2 PUE, IREN can run 17,140 B300s in 41.67MW IT load. Now by that ratio, they can run 24,680 GPUs in 60MW IT load — a massive difference with 34k units through the Nvidia DGX reference calculation. If common sense is applied, you can still get to 2 completely different outcomes, that show a difference of more than 9k GPUs. The GPU hour rate at 24.68k GPUs would be $3.145 per B300, as MASSIVE difference from the earlier calculated $2.28. Sure, the DGX system may be a factor here. And I'm sure that the reality is somewhere in the middle. But I personally hate this as an investor, to be unable to calculate profitability on unit economic basis. After all, contracts are signed on a $/GPU hour basis. Why hide this from your investors? Not being able to calculate payback periods, unable to calculate ROIC. And most importantly, we cannot properly assess the $NVDA deal on a contract basis. I really hope the payback period of this contract is not 4 years. I want the glass to be half full, but by starting to censor the purchases, IREN is taking a step in the wrong direction. Not a fan of this.show more

Frans Bakker
148,167 просмотров • 4 месяцев назад