⚡️ We just added support for Nvidia G4 (RTX... Pro 6000 Blackwell Server Edition) GPUs in Colaboratory! 🔥 With a peak rating of 960 BF16 TFLOPs (~50% more than the A100-80G) and 96 GB of VRAM (20% more than the A100-80G), it’s the most efficient high-performance GPU we’ve ever released. ⏱️Time to make GPUs go brrrr 🏎️show more

Colaboratory
62,470 Aufrufe • vor 6 Monaten
245 tokens/sec on DeepSeek-V4-Flash and it’s only using 2... RTX 6000 Blackwell GPUs. tinycorp dropped a 2-GPU tinybox edition in honor of DeepSeek fully ready to expand to 4 GPUs later. The GPU middle class just got a serious upgrade.show more

Md Ismail Šojal 🕷️
16,148 Aufrufe • vor 1 Monat
This is what happens when 48GB of VRAM isn't... enough. Someone put 4× NVIDIA RTX A6000 inside an HP Z8 Fury G5. That gives you 192GB of VRAM across four GPUs. But here's the interesting part: You don't magically get one 192GB GPU. The memory is split across four cards, and the software has to manage how the model is distributed between them. That's why these machines are interesting for local AI. The problem isn't always compute. Sometimes you just need somewhere to put the model. Would you build this today, or go for a newer GPU with more bandwidth?show more

KyzoroX
81,125 Aufrufe • vor 6 Tagen
What’s in sarah guo's bag? The ultimate robotics brain.... 🧠 NVIDIA Jetson AGX Thor, packing up to 2,070 FP4 TFLOPS of AI compute and 128 GB of memory. Power next generation humanoids and autonomous systems with real-time reasoning and server class performance, built for the lab, factory, and field. Learn more 👉show more

NVIDIA Robotics
92,795 Aufrufe • vor 1 Monat
The "I don't have enough VRAM" excuse just died.... I’m running Meta’s new 30B Muse Glimmer Q6_K_XL with a massive 130k context window on just 26GB VRAM FREE compute on Kaggle. Kaggle provides you free 2x Nvidia T4 GPUs. 30 hours usage each week! Yesterday, I showed you the violent throughput of Muse Glimmer on a single RTX 4090. Today, we are securing a Dual NVIDIA T4 GPU cluster with 32GB of total VRAM for exactly $0 and dropping the massive 24.5GB Q6_K_XL GGUF onto it. Here is the exact Kaggle workflow and benchmarking breakdown: # 1. The Storage Bypass & Setup I built a clean cell by cell script in the file. We dynamically fetch the CUDA accelerated llama.cpp binaries and use wget to stream the model directly into Kaggle's /kaggle/tmp scratch storage, which cleanly bypasses their 19.5GB output directory limit. # 2. The Multi GPU Performance With the -ngl 99 flag offloading all model layers across both T4 GPUs (32GB VRAM combined), we pushed a massive 131,072 token context window (-c 131072). The benchmark numbers: Prefill: 265.9 t/s Decode: 9.0 t/s VRAM Total: 26.5 GB # 3. The Architecture Insight The Q6_K_XL model itself is 24.5 GB. Because of Muse Glimmer's aggressive 16:1 GQA, the unquantized KV cache for a massive 130k context window only takes up 2 GB of memory. No heavily degraded Q4 KV quantization required. It just works. No compiling from source. No credit card. No OOM crashes. Zero excuses. If you’re running a single RTX 3090, 4090, or 5090, you need to experience this hyper efficient KV cache right now before the upcoming Qwen 3.8 27B drop completely steals your VRAM tomorrow. pick the Q4 or Q5 quants for 24 GB VRAM rigs. I'm dropping the Unsloth huggingface GGUF links and the free Kaggle notebook link in the replies. spin up your own instance, and show me your multi GPU benchmarks.show more

Alok
19,089 Aufrufe • vor 29 Tagen
> 8 GPUs in one server rig > dude... went homeless to build it > electrical bill costs more than rent now > while everyone else pays $400/month to openai > a 2 GPU desktop kills the api bill forever > rtx 4080 super + rtx 5060 ti = 32gb vram > runs qwen 3.6 with 100k context locally > no rate limits, no api keys, no data leaving the room > agents loop 400 times for free > claude opus still wins on hard reasoning > but local handles 90% of daily work > $1,200 setup pays itself off in 4 months > bookmark this and read the article belowshow more

starmex
167,058 Aufrufe • vor 3 Monaten
🔥🔥🔥We’ve been listening to your feedback! Our latest world... model HY-World 1.5 just got a major upgrade to make world generation more accessible than ever: 🛠️ Open Training Code: Fully customizable code for building and training your own models. ⚡ Accelerated Inference: Turbocharged speed and optimized VRAM for real-time interaction. 📉 Lite 5B Model: A new lightweight model that fits into small-VRAM GPUs. 🙌 Zero Waitlist: Our online app is now fully open to everyone—no application required. This is just the beginning. HY-World is building the future of spatial intelligence—open, accessible, and community-driven. 🕹️ Play now: ⭐ GitHub:show more

Tencent Hy
20,581 Aufrufe • vor 8 Monaten
⬛️ We are currently accelerating the incubation of GPU... Nodes into the infraX Network, with 12 H100’s currently available for operation. Despite the incubation of such immense GPU power, the infraX Platform is optimally designed to run on the least amount of computational power possible, meaning a lot of our available GPU nodes are currently sitting idle. Currently, we're utilising a single gigantic NVIDIA H100 server with 80GB of VRAM and over 220GB of RAM to run our Platform. To put that in perspective, it rivals the computational power of an adult human brain. This setup enables us to handle immense computational load and deliver high-quality AI content to our users, however we have much more in store. Our remaining, immense network of GPU units is currently being prepared for rental operations as we look to transform the corporate GPU lending sphere through our corporate GPU lending protocol. We already have many high tier Web3 Players ready for technical integration, with more approaching us daily. Through our V3 DApp we look to make these integrations publicly viewable with real time usage graphs integrated directly into our Platform, allowing for exceedingly unique viewing opportunities. $INFRAshow more

infraX | $INFRA
42,843 Aufrufe • vor 1 Jahr
Ball Swap! 🏀🔥 Groups start with one more ball... than the amount of people. Challenge groups to make the most swaps in a row! For more information on educational visits, CPD and much more, visit and get in touch!show more

WannaTeachPE
22,377 Aufrufe • vor 1 Jahr
⛓️ Aethir - the decentralized #GPU powerhouse reshaping #AI... & #gaming! Aethir is building the future of high-performance computing with a global #DePIN network of 400,000+ enterprise-grade GPU containers (including #NVIDIA H100s, H200s & more) spanning 90+ countries. 📍 Two flagship products: • Aethir Earth – Bare-metal GPU cloud delivering raw power for AI training, fine-tuning & inference with zero virtualization overhead. • Aethir Atmosphere – Low-latency cloud gaming rendering that streams high-quality experiences to any device. ☁️ Cloud Hosts monetize idle GPUs and earn $ATH rewards, while customers get scalable, cost-efficient compute (up to 80% cheaper than traditional clouds), ultra-low latency, and 95%+ utilization rates. No massive CapEx, no vendor lock-in – just on-demand access closer to the edge. From AI model training to real-time cloud gaming and beyond, Aethir is democratizing enterprise GPU power and powering the next generation of innovation. 🦾 Axe Compute’s $317M in customer prepayments. That single number reframes how AI data centers get built in 2026 🧵 The decentralized cloud is here. Are you ready? 🌐 #Aethir #DePIN #AI #GPUCloud #Web3show more

Crypto Holding™ 💎
229,806 Aufrufe • vor 16 Tagen
13 months after we reached 500k, we’ve just hit... 600k! Honestly can say I am enjoying what we’re producing more than ever, the team behind the scenes is the strongest it’s ever been and I get a good feeling about what’s next for United. Thank you all for helping me make the best community I could have ever dreamed of xshow more

Sam Peoples
41,173 Aufrufe • vor 3 Monaten
50% more context unlocked for Qwen 3.8 27b Q4_K_XL... dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?show more

Alok
39,189 Aufrufe • vor 21 Tagen
We just closed the largest AI agent pre-sale ever.... 76,831 Sol in under 36 hours. This historic raise shows Web3 needs proven builders. With a track record of creating virtual humans since 2018, delivering for top-tier brands and sustaining real revenue, we can’t wait to continue delivering on this vision. We’re grateful for the support. It’s more than we ever expected and more than we need. That’s why we’re returning 50% to everyone who contributed. • Refunds will be issued within 24 hours. • Launch this week. We will use the remaining 50% to support the launch, provide liquidity, pay listing fees (if any), and continue to build this project out long term. As we finalize next steps, we’re proud to share an early look at what we’ve been building:show more

MIRAI
252,236 Aufrufe • vor 1 Jahr
Run Gemma 4 26B MoE on 8GB VRAM with... 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware. Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card. The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?" Today, I’m delivering exactly that. I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!. If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed. The performance metrics are astonishing: - 20 tokens/sec flat decode throughput. - Stable, flat decode speed even with massive prompts. - I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame. # What about prefill? Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable. And this is running completely without Multi Token Prediction (MTP) active. How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4. The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse. # The Test Setup: CPU: Intel Core i7 RAM: 16GB System RAM GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) # The Secret Sauce (The -cmoe Flag) To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp. This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache. It prevents VRAM spillage and holds the throughput rock solid. # The flags: -m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking. Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the repliesshow more

Alok
292,770 Aufrufe • vor 3 Monaten
A week ago, we launched the OptimAI Core Node.... Today, the numbers speak for themselves; but the story behind them speaks even louder. More than 3,000 active Core Nodes now power the OptimAI Network: + 43,831 CPU cores + 75,289 GB RAM + Over 1,000 GPUs spanning NVIDIA, AMD, Intel, Apple, and more Together, they form a distributed AI supercomputer - built not by a corporation, but by the community. This is how the future of Agentic AI begins: AI that learns from live data, reasons across the open web, and evolves through human collaboration. We’re not just running nodes. We’re building the backbone of decentralized intelligence. 👉Join the movement:show more

OptimAI Network
20,649 Aufrufe • vor 10 Monaten
There is nothing more valuable than your time. And... yet it is time that most goes to waste. Money can’t give you more time. But money can get you the time of others. And that time is worth much much more than most folks are being compensated in exchange for it. You can’t get that time back. The majority believing there is little value for their time is part of how those who do assign high value to their own time keep their control. Time is a real resource. Money is a social construct. One of the two is much more valuable. But what do I know, I’m just a silly girl with hand puppets.show more

Liz Katz
76,299 Aufrufe • vor 1 Monat
Run Gemma 4 26b MTP on 8 GB VRAM... GPUs at 25+ tokens/second. Flags included! local llm space is moving at terminal velocity. only 3 days ago google released gemma 4 26b a4b qat quants. more efficient than before, ran on 8gb vram at 20 tok/sec. and now just a few hours ago, mainline llama.cpp merged a massive update and we just shattered our own record. decode throughput went 25-40% up on the same 8 GB VRAM setup! Before MTP: 20 tps -> After MTP: 28 tps! llama.cpp just officially merged PR #23398 ("add Gemma4 MTP"), bringing native Multi-Token Prediction (MTP) support to Gemma 4 models. By running speculative drafting on the same 8GB VRAM RTX 4060 setup, my decode throughput on a 64k context instantly leaped to a blistering 25–27 tokens/sec thats 25-30% increase with the same hardware. Here is the architectural catch you need to know: Unlike the Qwen 3.5 and 3.6 series, which bake the MTP heads directly into the base GGUF, the Gemma 4 MTP head is not built in. You must download a separate, specialized MTP drafter GGUF (the assistant model) to act as the speculator. (I've dropped the download link in the replies). copy and try the exact flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.7 --spec-draft-model gemma-4-26b-A4B-it-assistant-Q4_0.gguf -c 64000 -v n-max 4 and p-min 0.7 is also worth checking out. benchmark on your setup and workflow. if you have a single 8 gb vram nvidia rtx 4060, 3060, 3070, 2080, 2070, grab the MTP drafter GGUF link in the comments and try it yourself. Check it out even if you have asmaller or a larger gpu, such as a single rtx 3090, 4090, 3060, 2060. MTP works for all gemma 4 sizes such as gemma 4 12b, gemma 4 31b etc. but remember to grab the correct mtp draft assistant models respectively. what are you benchmarking todayshow more

Alok
200,913 Aufrufe • vor 3 Monaten