Загрузка видео...

Не удалось загрузить видео

На главную

Easter loading… but the data is already here. Start your weekend early with #GagaWeekend: UGX 6,000 4.5GB + 20 MTN mins UGX 8,000 6GB + 20 MTN mins UGX 10,000 8GB + 20 MTN mins Dial *100*0# or use MyMTN App.

14,643 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Run Gemma 4 26B MoE on 8GB VRAM with 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware. Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card. The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?" Today, I’m delivering exactly that. I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!. If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed. The performance metrics are astonishing: - 20 tokens/sec flat decode throughput. - Stable, flat decode speed even with massive prompts. - I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame. # What about prefill? Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable. And this is running completely without Multi Token Prediction (MTP) active. How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4. The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse. # The Test Setup: CPU: Intel Core i7 RAM: 16GB System RAM GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) # The Secret Sauce (The -cmoe Flag) To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp. This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache. It prevents VRAM spillage and holds the throughput rock solid. # The flags: -m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking. Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the replies

Alok

292,770 просмотров • 2 месяцев назад

you're paying $20/mo for something your $500 GPU can already do. Gemma 4 26B A4B QAT MoE + Hermes Agent running on a single RTX 4060 (8GB VRAM). Built a vision capable, 100% free, 100% local, private AI assistant that lives in my Chrome browser. No API keys. No cloud. No subscriptions. 100% vibe coded. 0% handholding. It has full context of whatever's on my screen can answer questions, summarize pages, extract data, and see images. Same local model handles everything, no external calls, ever. keep reading for the model and hermes agent tips i learnt while building this locally. Here's the exact setup for anyone running local LLMs on 6-8 GB VRAM: llama.cpp server flags (on my NVIDIA RTX 4060 8gb VRAM): -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --cache-type-k q8_0 --cache-type-v q8_0 -c 150000 --port 8080 Throughput with quantization: Prefill: 200-250 tokens/sec Decode: 20-25 tokens/sec reduce context if oom on 6 gb vram card. Key learnings: - Quantize KV cache to q8 for faster prefill/decode. Prefill goes from 100-150 (unquantized) to 200-250 tok/s (q8). - But watch out, once actual context grows past ~50k tokens on high entropy workloads, q8 KV quantization can cause hallucinations. Low entropy workloads are mostly unaffected. If you see it happening, drop the quantization. This is common across all local models. - In Hermes Agent settings -> Memory & Context, bump compression threshold from default 0.5 to 0.7. Default triggers way too frequent context compression and eats time. Up next: add persistent memory, web search, tool calling, streaming output and whatever you suggest. Running a 26B MoE with vision + 150k context window on 8GB VRAM would've sounded impossible 6 months ago. Works the same on the NVIDIA RTX 3060 Ti, 3070, 4060 Ti, 5060, 2080, or any 8GB card. VRAM is the only requirement. Local AI agents are closer than people think. You just need to know where the knobs are. Model's Unsloth quant hugging face link in the comments. Have you tried Hermes agent by Nous Research yet? What are you building with local LLMs? Drop it below, let's see what this community is shipping.

Alok

36,031 просмотров • 1 месяц назад

🚨🎙️İlkay Gundogan after Galatasaray 0-3 Venezia: 🗣️ “Look, these are friendlies. I know that. But 0-3 against Venezia, after we already lost to Monza, is not something we can just brush off as ‘pre-season’. Venezia are a solid Serie A side, yes, but they are not a European powerhouse. Getting completely dominated by them while the fact this squad still has players of world-class calibre suggests something beyond individual quality is missing. We have quality in individual positions. That is not the issue. The issue is depth and completeness. One signing so far is not a squad rebuild. Midfield is thin the moment someone is rotated or tired. The defensive line looks fragile when the intensity drops. I have played in systems at the very highest level —systems that demand 18–20 ready players, not 13 or 14 and a prayer. Right now it feels like we are waiting for the perfect deal while the team is already being exposed. Management keeps repeating the same line: ‘No panic. We will wait for the right players. We will not overpay.’ I understand the principle. But while we wait for the perfect transfer, we are shipping goals in Austria against teams that will not even be fighting for European places. Last season the same ‘patience’ talk was easier to accept because the squad still had more options. This summer the cupboard looks emptier, and the results are already reflecting it. The fans have been restless for weeks. ‘Transferler nerede?’ is not just noise. When you win four titles in a row, expectation becomes the standard. If we walk into the new season with the same holes we showed today, those titles start looking like the past instead of the baseline. The dressing room sees it too. You cannot keep asking players to cover structural gaps indefinitely. I did not come to Galatasaray to manage decline or to protect a dynasty with polite waiting. Either the reinforcements arrive quickly in the positions we clearly need, or the anger that is already building among the supporters will stop being about friendlies and start being about the season itself. And at that point, the players will have very little room left to hide behind ‘it’s only pre-season.”

Vfynn_🥷🏼 𐙚

15,568 просмотров • 15 дней назад

6 months ago we were dropping a new app every week No one cared We’d randomly get 10k users on an app But they came for the app, not the person The thing is, I don’t care about making a retentive app for one audience I want to be a retentive person ~ for my audience What if I was the app? Not a single Paul Thomas Anderson movie is the same Different subject matter, different genres, so you’d assume different audiences However the same people that went to see his last film, came to see One Battle After Another So the demographic isn't dependent on the subject matter The demographic is just, Paul Thomas Anderson fans The software industry has long told that you need to work on one thing, for the rest of your life That’s not how art works tho, is it? Can you imagine telling Jay Z “Great job on the blueprint, now iterate on that same album for the next decade” The landscape of tech haas been stifling the growth of creators by not allowing them to explore other interests 6 months ago I said no to this "requirement", despite what everyone told me, and continued to drop what I liked every week The second a trend was happening on Tiktok, I had the app out that week Somehow 6 months later, the world is conforming to this ideology Instead of software creators limited to making apps for one audience and one niche, there’s a new world of ephemerality and expression What if instead of optimizing for users, we optimized for fans Making apps that are expressive of your life, your commentary, your heartbreak Garnering an audience that will follow you through each step of your story Each of those steps being its own app Why shoot for daily active users when you can get daily loving fans When fans use your app, it’s not just about resonating with the story, the app places them IN THEIR OWN story Here’s an example You’re a 20 year old girl who’s at UMiami You scroll through Tiktoks in your dorm room about “mogging”, a trend to outshine your friend in a photo You laugh and share videos seeing celebrities mog each other, but that’s the extent of it Then at Danger Testing we make an app called mog or not, where you and your friend can upload a photo and AI tells you who’s mogging Now you’re at the bar with your sorority sisters, playing all night, whether your winning or losing it’s the time of your life cause something is finally about YOU ENOUGH OF WATCHING MOVIES LET’S MAKE YOU THE MOVIE LET’S MAKE YOU THE STAR AN APPSTAR

los (appstar)

14,257 просмотров • 10 месяцев назад

$25K+ profit daily from 1 wallet, with OpenClaw. I have the exact step-by-step guide, giving it free for 24 hours. To get it: 1. Comment "OpenClaw" 2. Like and Retweet. 3. Follow me Himanshu Kumar ( So, i can send you DM) I ran a simple script last night with Claude Code. Pull on-chain data from Polymarket, sort by win rate on 15 minute BTC markets. 20 minutes later, 100s of wallets showed up. Most were losing money or barely breaking even. Then I spotted 1 address. 200+ trades daily, every single week profitable, timing so precise it looked robotic. Because it is. I fed the wallet address back into Claude Code. Asked it to reverse engineer the strategy. 20 mins later the full breakdown appeared on my screen. Here is how it works: Bot monitors Binance and Bybit every 100ms. Waiting for BTC volatility compression to drop below 0.08%. When it hits that level, it buys both Up and Down contracts at 25 to 35 cents each. Classic straddle play. 1 contract loses, the other rockets to a dollar. Entry at 30 cents means 3x to 4x return every time. Repeats dozens of times per day. Result: $13K to $25K profit daily from 1 wallet. No human intuition, no insider tips. Just an algorithm exploiting a gap in market mechanics. I searched to see if anyone else found this wallet. Turns out yes. There is a Telegram bot that auto-copies trades from wallets like this. I connected it to the same address. Every entry matched what my terminal showed. You can now copy-trade an algorithm in real time. That capability did not exist 12 months ago. Comment "OpenClaw" and I will send you everything. Must Follow me Himanshu Kumar to get the DM.

Himanshu Kumar

13,188 просмотров • 5 месяцев назад