
AJ
@ItsmeAjayKV • 2,428 subscribers
Bullish on local AI, llm finetuning, abliteration, llama.cpp always open source GPUmaxxing ( 1x 3060, 1x 3090, ++) ML, GenAI, AIAgents, CUDA, N4 (日本語勉強中) 🎌
Videos

Here's a comparison i never thought i would make Qwen-3.8-27b vs Gemini-3.7-flash Yes, you read it right! I don't know what to say, this is soo damn impressive. the difference is hugeee🤯 Take a look at this result 😍 (watch it in highest quality) Same prompt, both one shot on web-chat ui Qwen-3.8-27b produces consistent beautiful, visually stunning results. This is one of the best water surface simulations i have seen across every local models i have run and at the same time beating Gemini-3.7-flash result. And this is running on my 3090. And yes its Q5, not even f8.
AJ168,072 просмотров • 3 дней назад

Every big MoE model I test on my 3090 is basically to see how good it is compared to Qwen3.6-27B. "How good is it compared to Qwen3.6-27B?" is the question I get asked the most, and that shows how much people love and, more importantly, trust the 27B 27B is battle tested. It has proven its worth and has nothing left to prove. Now it just needs to make way for the next 27B. Here’s an example video: a Three.js task, same prompt, not oneshot but two shots for all models. All ran on my 3090 + 64GB VRAM Qwen-3.6-27b fully on vram while others have experts offloaded to system ram. Pinned on the left: Qwen3.6-27B-Q5 On the right, i'm switching between Inkling-Small-Q4_K_M DSV4-Flash-0731-UD-IQ1 Ling-3.0-Flash-Q4_K_M Qwen3.6-27B nailed it. The effects and visuals are just spot on.The other models got some things right, but not everything. Second place would have to go to Inkling-Small while DSV4-Flash-0731 and Ling-3.0-Flash failed to capture the full requirements. This is also why we’re all excited for Qwen Qwen3.8-27B!
AJ10,669 просмотров • 9 дней назад

🤯Laguna S-2.1 UD-IQ3_S on a single 3090 ~30-35t/s decode ~280-290t/s prefill ~23GB VRAM @ 64k context, f16 KV llama.cpp: -ngl 999 + --n-cpu-moe 32 + fa on 118B MoE, ~8B active, did we get a decent mid-range model that doesn't need multi gpu and can run with good enough speeds for long-context coding tasks? Might actually stick with this one. Been waiting for a mid-range model that fits on a single 3090, this is it. Usual three.js tests underway. Update soon.
AJ20,031 просмотров • 26 дней назад

Y’all thought I stopped at Q4_K_M 😭 Nope, q3 and q4 didn't cut it, quality was not at all good on my three js tests. So now running Poolside Laguna S-2.1 UD-Q5_K_S same 3090 + 64GB RAM same hybrid offload q8_0 KV 166k ctx Q4 was ~73 GB. Q5 is ~83 GB. More bits on disk. Hopefully this one is better, We’ll see if it’s worth it.
AJ13,341 просмотров • 23 дней назад
Больше нет контента для загрузки