Loading video...

Video Failed to Load

Go Home

Here is Ternary Bonsai 27B running an end-to-end agentic workflow locally with Hermes on an NVIDIA GeForce RTX 5090 GPU. The model reasons, calls tools, reads outputs, modifies files, and surfaces insights - all on consumer hardware, while all private files, intermediate states, and iterations remain local.

154,473 views • 1 month ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

#Battlefield6 NVIDIA DLSS 4 Reveal Trailer 📽️ The game includes support for DLSS 4 with Multi Frame Generation, DLSS Frame Generation, DLSS Super Resolution, DLAA, and NVIDIA Reflex. Desktop GPUs Performance At 4K, Ultra settings, DLSS 4 with Multi Frame Generation and DLSS Super Resolution multiply Battlefield 6's GeForce RTX 50 Series frame rates by an average of 3.8X. ▪️ GeForce RTX 5090 performance rockets to over 470 FPS. ▪️ GeForce RTX 5080 exceeds 330 FPS. ▪️ GeForce RTX 5070 Ti is in touching distance of 300 FPS. ▪️ GeForce RTX 5070 surpasses 230 FPS. At 2560x1440, Ultra settings, DLSS 4 with Multi Frame Generation and DLSS Super Resolution increase Battlefield 6 frame rates by an average of 3X. ▪️ GeForce RTX 5090 runs at almost 600 FPS. ▪️ GeForce RTX 5060 Ti over 240 FPS. At 1920x1080, Ultra settings, the combination of DLSS 4 with Multi Frame Generation and DLSS Super Resolution boost Battlefield 6’s GeForce RTX 50 Series frame rates by an average of 2.9X. ▪️ Allowing every GPU in NVIDIA's line-up to play at over 230 FPS, maxing out at over 740 FPS on the GeForce RTX 5090. Laptop GPUs Performance GeForce RTX 50 Series Laptop GPUs benefit similarly from DLSS 4 with Multi Frame Generation and DLSS Super Resolution, with Battlefield 6 frame rates multiplied by 3X on average at 2560x1600, with Ultra settings, enabling performance of up to 360 FPS. At 1920x1080, Ultra settings, GeForce RTX 50 Series Laptop GPU performance increases by 2.8X on average, enabling Battlefield 6 frame rates to surpass 480 FPS, and the range to run at over 200 frames per second.

Battlefield Bulletin

30,559 views • 10 months ago

this is what 12 gigs of VRAM built in 2026. a 9 billion parameter model running on a 5 year old RTX 3060 wrote a full space shooter from a single prompt. blank screen on first try. i came back with a bug list and the same model on the same card fixed every issue across 11 files without touching a single line myself. enemies still looked wrong so i pushed another iteration and now the game has pixel art octopi, particle effects, screen shake, projectile physics and a combo system. all running locally on a card that was designed to play fortnite. three iterations. zero cloud. zero API calls. every token generated on hardware sitting under my desk. the model reads its own code, finds what's broken, patches it, validates syntax and restarts the server. i just describe what's wrong and it handles the rest. people are paying monthly subscriptions to type into a browser tab and wait for a server farm to respond. meanwhile a GPU you can find used on ebay is running a full autonomous hermes agent framework with 31 tools, 128K context window and thinking mode generating at 29 tokens per second nonstop. the game still needs work. level upgrades don't trigger and boss fights need tuning. but the fact that i'm iterating on gameplay balance instead of debugging whether the code runs at all tells you where this is headed. every iteration the game gets better on the same hardware. same 12 gigs. same 9 billion parameters. same RTX 3060 from 5 years ago your GPU is not a gaming card anymore. it's a local AI lab that never sends your data anywhere.

Sudo su

170,848 views • 5 months ago

six months ago this wasn't happening on 8gb vram. running unsloth's Q4_K_XL quant of gemma 4 26b-a4b-it-qat, a sparse MoE model with only 4b active params on a single rtx 4060 laptop gpu, 8gb vram, 20+ tok/s decode. no cloud, no api, no offload hacks. just a gaming laptop on battery. what makes it fit: google's QAT (quantization aware training), plus MTP (multi token prediction) support in the latest llama.cpp builds. that combo is the single biggest unlock for local inference on low vram. rtx 3060, rtx 3070, gtx 1070, gtx 1080, rtx 4050, rtx 4060, rtx 5050, rtx 5060 — any 6-8gb consumer gpu, old or new — this model runs on it. world cup season, so i told it to build a soccer themed flappy bird clone. one shot, zero iteration, fully playable. six months ago an 8gb model could barely clone vanilla flappy bird. now it's shipping a themed game from a sparse MoE model running locally on a laptop battery. inference benchmarks: - decode throughput: 30 tok/s - context: 64k. this is the real unlock. 64k ctx is what makes a hermes agent loop viable locally on this model, not just single-turn chat. llama.cpp flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 64000 -cmoe --port 8080 game's deployed on my own site, built and shipped end to end with open source llm, zero closed source api dependency in the pipeline. link in the description. gguf weights on huggingface, link in the comments. pull it down, run it on whatever 8gb card is sitting in your rig. try the game and tell me your score and what you want in v2. local llms on consumer gpus stopped being a meme.

Alok

61,660 views • 2 months ago