Qwen3-Coder 30B running locally in the @GitHub Copilot app... (ollama) spitting ~85 tokens per second on 96GB VRAM. It's not Sol, but we're getting somewhere.show more

Burke Holland
24,098 次观看 • 21 天前
Did I tell you that I vibe coded a... Tamagotchi style TUI for my Moltbot running on a 1400 EUR beelink AMD Ryzen AI 9 HX 370 with 64GB LPDDR5X and Ollama serving a GLM-4.7-Flash 30B at 20 tokens per secondshow more

Florian Roth ⚡️
87,745 次观看 • 7 个月前
GEMMA 4 26B ON AN RTX 4060 WITH A... 248K TOKEN CONTEXT WINDOW 20 tokens per second and a context window so large you can feed it entire codebases, books and research papers in a single prompt this is not a cloud api and not a server rack, this is a regular consumer gpu running locally with llama.cpp and q4_k_xl quantization 248k context on an 8gb vram card was not supposed to be possible and here it is just running on someone’s desk the article below covers exactly which tools and configs make this kind of setup work in 2026 ↓show more

leopardracer
56,772 次观看 • 2 个月前
It's over. OpenAI just crushed it. We have their... o3-level open-source model running on Groq Inc at 500 tokens per second. Watch it build an entire SaaS app in just a few seconds. This is the new standard. Why the hell would you use anything else??show more

Matt Shumer
498,727 次观看 • 1 年前
Right now, you may not have access to models... like GPT‑5.6 Sol, GPT‑4.6 Terra, GPT‑5.6 Luna, Claude Mythos 5, or Claude Fable 5. But you can run something surprisingly powerful today, locally, and completely free. in the next 10 mins on your 8 GB VRAM gaming laptop. Gemma 4 26B A4B QAT (MoE) delivers strong performance on a standard 8 GB VRAM GPU using Ollama, with no API, no usage limits, and no external dependencies. Out of the box, it reaches around 20 tokens per second without any optimizations. Only one command in your terminal: Ollama run gemma4:26b This means: Full offline capability (privacy by default) Zero recurring cost Competitive performance for many real world tasks Fast enough for interactive use on cheap consumer hardware If you're waiting for cutting edge cloud models, you're missing what is already practical today: a capable, local LLM that runs entirely on your own machine.show more

Alok
65,387 次观看 • 2 个月前
We are thrilled to be a launch partner for... Meta Llama 3. Experience Llama 3 now at up to 350 tokens per second for Llama 3 8B and up to 150 tokens per second for Llama 3 70B, running in full FP16 precision on the Together API! 🤯show more

Together AI
88,229 次观看 • 2 年前
A GPT-4 level chatbot, available to use completely free,... running at over 800 tokens per second on Groq. I'm genuinely mindblown by LlaMA 3. Try it with the link in the next tweet.show more

Rowan Cheung
1,242,164 次观看 • 2 年前
AN AWS ENGINEER QUIETLY BUILT A 2 PETABYTE HOME... SERVER FOR $9/MONTH THAT KILLS A $3,400/MONTH CLOUD STORAGE BILL the lenovo thinkstation pgx ships nvidia's gb10 grace blackwell superchip and 128gb of unified memory in a box the size of a mac mini at 1.2kg it runs an 80b qwen3 coder model at 25 to 40 tokens per second and a 196b step-3.5-flash moe model at 20 tokens per second locally the gb10 packs 6,144 cuda cores, 192 fifth-generation tensor cores and rates at 1 petaflop of fp4 with sparsity from a single 240 watt usb-c power supply fine tuning qwen 2.5 7b with lora took 18 minutes and 41gb of unified memory while the gpu pulled 65 watts and peaked at 77 degrees the box pulls a docker container from nvidia's registry and serves a frontier model on your local network with tool calling and zero data leaving your desk bookmark this and read the article belowshow more

starmex
193,226 次观看 • 2 个月前
They are not online or poor network? #Diaspora Call... 🇺🇿 Nigeria at $0.00167 per second. Stop paying for minute you didn’t use. Make the switch now & save 90% on your calls to Nigeria Get it App now! Living in the diaspora, you know how it goes—you need to check in on a project or family back home, but the person on the other end is offline, or the data connection just isn't holding up. SoftTalk Messenger is the solution An African super calling app - charges by per second — not per minute. • ✅ Rate $0.00167 / secshow more

RAZOR BLADE
360,591 次观看 • 1 个月前
Aaron Boone was asked if he thinks the Yankees... are "running out of time" to get themselves on track against teams above .500: "No, we're not running out of time, but if we don't do better, then it's going to fizzle out and we're not going to get to where we want to be."show more

SNY Yankees
231,772 次观看 • 1 年前
Fable 5 absolutely crushed the HTML5 physics contest, but... cost 6x more than Opus 4.8 and 39× more than GLM 5.2 in that test. Test was done on atomic[.]chat, a desktop app that runs LLMs locally. The test asked 4 models to generate self-contained canvas demos with believable motion and collisions. The scenes were not simple animations because every crash needed gravity, force, timing, and contact handling. Outputs: - Fable 5: 62,158 tokens, $3.12 - GPT 5.5: 37,753 tokens, $1.14 - Opus 4.8: 22,280 tokens, $0.56 - GLM 5.2: 36,246 tokens, $0.08show more

Rohan Paul
205,771 次观看 • 2 个月前
GROK CODE IS WIPING THE FLOOR WITH EVERYONE It... launched 5 months ago and it's already the most-used coding model out there. On the Kilo Code leaderboard, it's getting used twice as much as the second-place model in almost every category. In 2025, it’s way ahead on OpenRouter too, with over 16 trillion tokens processed. Nothing else is even close. Source: X Freezeshow more

Mario Nawfal
65,325 次观看 • 7 个月前
This is running in my web browser on my... laptop. It's deterministic and running using webgpu Tonight; I have a /goal going where it is going to set up one of my 4090s as the authority server, and then test all the rollback code. Goal is 50k boxes simmed per user, up to 512 users. Massively multiplayer massively rigid body physics online games inside your web browser. Why the hell not, we've got the tokens to spare!show more

kache
39,217 次观看 • 1 个月前
Jaylen Brown was doubled on 34.7% of his touches... last night, the most in any game of his career where he had 70+ touches. It worked out for OKC in the 1st half, with the Celts getting just 0.25 points per direct touch when JB was doubled. In the second half? Not so much, with BOS getting 1.7 points per direct touch when JB saw twoshow more

ALL NBA Podcast
39,966 次观看 • 5 个月前
50% more context unlocked for Qwen 3.8 27b Q4_K_XL... dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?show more

Alok
39,189 次观看 • 12 天前
Sightings have been reported of a Chaos Emerald making... landfall somewhere in San Diego on Saturday July 15th. We're almost there but the job's not done… Be ready to secure each and every Emerald so Sonic can power Club Chaos! Join the Chaos Hunt:show more

Sonic the Hedgehog
252,510 次观看 • 1 个月前
Been using Qwen 3.8 27B (Q4) locally on 64GB... of VRAM. Here is the verdict: SLOW 18 tps with ZERO system prompt to process and that degrades significantly with a harness system prompt and as the context window grows. RIP if you have to compact. I had it implement this PRD and it's been running for 6 hours. By comparison Grok 4.6 and Kimi K3 hosted finished in about ~30 minutes. High hopes, but these 27B variants are too dense. This is not a consumer grade local model - and I consider consumer grade to be anything up to $5000.show more

Burke Holland
121,540 次观看 • 18 天前
Run Gemma 4 26B MoE on 8GB VRAM with... 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware. Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card. The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?" Today, I’m delivering exactly that. I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!. If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed. The performance metrics are astonishing: - 20 tokens/sec flat decode throughput. - Stable, flat decode speed even with massive prompts. - I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame. # What about prefill? Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable. And this is running completely without Multi Token Prediction (MTP) active. How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4. The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse. # The Test Setup: CPU: Intel Core i7 RAM: 16GB System RAM GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) # The Secret Sauce (The -cmoe Flag) To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp. This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache. It prevents VRAM spillage and holds the throughput rock solid. # The flags: -m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking. Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the repliesshow more

Alok
292,770 次观看 • 2 个月前
Trump on NATO members declining to assist in the... Iran operation: "I just think that it's not good for a partnership, when they say, 'What you're doing is a great thing, but we're not going to help.'"show more

Open Source Intel
33,955 次观看 • 5 个月前
single RTX 3090. 24 GB VRAM. Qwen3.5-35B-A3B. 4-bit quant,... 113 tokens per second at full 262K context harnessing Claude Code locally with no API, no subscription, no proxy. told it what it is. 30 Mamba2 layers, 10 attention, 256 experts, 8 active per token. said "build something that shows off what you can do." it visualized its own architecture. interactive. tokens flowing through layers. 256 experts lighting up on routing. served in the browser from the same GPU running inference. single prompt. then i said level up. 3D. Three.js. separate files. flythrough camera. clickable layers. it planned first, scaffolded 6 files, hit one API bug, fixed it itself, then optimized for smooth framerate. two iterations to a working 3D neural network explorer. llama.cpp just merged a native Anthropic endpoint. Claude Code points at localhost. the whole setup is two commands. no LiteLLM. no proxy config. the open source models coming out of china right now are genuinely changing what's possible on consumer hardware. respect to the Qwen team. this is acceleration.show more

Sudo su
110,206 次观看 • 6 个月前