Loading video...

Video Failed to Load

Go Home

Kept you waiting, huh? 100% local and open source. Powered by Mario Zechner' pi coupled with qwen3 tts, parakeet.cpp and gemma-4-26b-a4b You'll need to provide the game iso but there is a script that rips all required assets automatically.

17,160 views • 3 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

💥 The Future Is Now: Pay Your Bills with Pi Using PrimePi Pay 💥 Powered by Pi. Built for the People. In a world racing toward decentralization and digital empowerment, one question still echoes for everyday people: When will crypto solve real problems? That time is now — and the answer is Pi Network. Introducing a revolutionary leap in the Pi ecosystem: a bold new app that finally lets you pay your real-world bills using Pi Coin (𝛑) — securely, instantly, and without relying on banks or middlemen. Welcome to PrimePi Pay — the bridge between blockchain freedom and the real-world responsibilities we all carry. 🔑 Why PrimePi Pay Matters Too many people are still stuck in a financial system that limits access, adds fees, and delays payments. Meanwhile, millions of Pioneers around the world have been quietly building a new financial layer — one mined on trust, time, and vision. Now it’s time to activate that vision. With PrimePi Pay, you’ll be able to: •Pay electricity, phone, internet, rent, and more using Pi •Scan bills and verify payment details with built-in AI tools •Send Pi directly to official businesses or trusted local agents •Track every payment inside your Pi wallet — fully transparent and secure ⚡ Real Utility. Real Adoption. Real Pi. This isn’t about hype. It’s about empowerment. You don’t need to convert to fiat. You don’t need to wait on banks. You don’t need permission. All you need is your Pi — and now, it can take care of your life’s most essential needs. PrimePi Pay is proudly powered by Pi — the people’s digital currency. 🧠 Powered by GenAI. Built by Pioneers. Using GenAI and Pi-native tools like Pi App Studio and Firebase, PrimePi Pay was created by Pioneers, for Pioneers. It’s simple. It’s powerful. And it’s laser-focused on solving real-world financial problems. It’s more than an app — it’s a global movement. You can even participate as a Prime Agent, helping users in your community pay bills while building a reputation inside the Pi economy. 🚀 PrimePi Pay: Just the Beginning As Pi Network continues its Open Mainnet expansion, PrimePi Pay will unlock: •Partnerships with major billers and utility companies •Mobile top-ups and rent payments in emerging markets •Local-to-global remittances, powered by trust and decentralization And guess what? It all starts with you. Your Pi. Your bills. Your power. 💬 Final Word: “One day, you’ll stop asking what Pi is worth. Instead, you’ll ask what you can do with it.” – A Pioneer of the New Economy Let’s make history. Let’s pay bills with PrimePi Pay. Powered by Pi. Designed for a new world. 💜🔌📲 #PrimePiPay #PoweredByPi #PiNetwork #PayWithPi #DecentralizeLife Pi Network Nicolas Kokkalis Chengdiao Fan

Mr Spock 𝛑

15,731 views • 1 year ago

Run Updated Gemma 4 26B A4B QAT (MoE) with Vision at 25 tokens/sec and massive 120k context window on a single RTX 4060 (8 GB VRAM + 16 GB RAM Only!!) Yesterday I pushed Gemma 4 26B A4B QAT to 250k context on a single RTX 4060 using nothing but Q8 KV cache and optimized -b and -ub flags for higher prefill throughput. Today I stacked Multi Token Prediction (MTP) self speculative decoding AND the vision projector (mmproj) on top of that same card, same batch size optimization, same $250 GPU and pushed it until it broke, then found the fix. All text only runs consist of a 28k prompt. vision runs consist of 28k text prompt + an image. # 1. MTP alone. near free decode speed, no catch MTP draft assistant is a separate small model (MTP heads are backed into the main model itself for the qwen 3.5+ models but its a separate small model for gemma 4 series), 240 MB gguf 80k ctx: Prefill 510 t/s | Decode 29.5 t/s 120k ctx: Prefill 433 t/s | Decode 29 t/s 180k ctx: Prefill 240 t/s | Decode 24.9 t/s 250k ctx: Prefill 63 t/s | Decode 13 t/s llama.cpp flags: m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp -md mtp-gemma-4-26B-A4B-it.gguf-c 180000 -b 1024 -ub 1024 --spec-draft-n-max 6 --spec-draft-p-min 0.7 -ctk q8_0 -ctv q8_0 # 2. Add vision on top. the tax you actually pay the vision projector gguf is about 1.1 GBs 80k ctx: Prefill 360 t/s | Decode 25.4 t/s 120k ctx: Prefill 230 t/s | Decode 23.8 t/s 180k ctx (Q8 KV): Prefill 75 t/s | Decode 12.5 t/s - cliff flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp -md mtp-gemma-4-26B-A4B-it.gguf -c 80000 --port 8080 -b 1024 -ub 1024 --spec-draft-n-max 6 --spec-draft-p-min 0.7 -ctk q8_0 -ctv q8_0 --mmproj mmproj-F16.gguf # 3. The fix if you want to run vision over 120k context: swap Q8 KV for Q4 KV past 120k Stack MTP + vision + Q8 KV past 120k context and you hit a wall. draft model overhead plus KV pressure tanks everything. Drop to Q4 KV and the wall disappears: 180k ctx (Q4 KV): Prefill 220 t/s | Decode 25.5 t/s -ctk q4_0 -ctv q4_0 --mmproj mmproj-F16.gguf (rest same as above) Bottom line: MTP gives you a near free +20-30% decode boost up to 120k context. Past that, it's fighting your VRAM, not helping and if vision is loaded too, Q4 KV isn't optional past 120k, it's mandatory. 30% boost is model and card specific, MTP boosted decode 2x for gemma 4 31b on a single rtx 4090. Same 8GB card. Same $250 GPU. Multimodal, speculative decoding, 180k usable context, zero upgrades. You gotta try this if you have a single NVIDIA RTX 3050, 3060, 3070, 4050, 4060, 5050 or 5060. You can try it with a 6 GB VRAM card as well but you will have to lower the context window. Hugging Face links to the updated Unsloth's QAT quants and performance graph are in the replies below. Which models are you running on your 6/8/12GB cards with MTP?

Alok

16,405 views • 2 months ago