Загрузка видео...

Не удалось загрузить видео

На главную

Here's a comparison i never thought i would make Qwen-3.8-27b vs Gemini-3.7-flash Yes, you read it right! I don't know what to say, this is soo damn impressive. the difference is hugeee🤯 Take a look at this result 😍 (watch it in highest quality) Same prompt, both one shot...

168,072 просмотров • 3 дней назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

This happens all the time! This is why I feel like I'm not worth myself! This is why I feel like I cant grow! This is twitch actively nerfing my account every time I go live! The audience didn't leave they were still in the chat! When I asked twitch for help, I got an automated incomplete response that said "everything looks" . I am so tired of fighting this and have pleaded with twitch, Daniel Clancy and have told the VP FACE TO FACE about this issue and still nothing has been done! Why? I'm small potatoes no one cares and no one will even put an iota of time into solving this issue! This has been a constant every day for over 2 years and I cannot express how defeated I absolutely feel every time I press start to immediately have my audience ignored by twitch. I know some of you are going to say its not all about the numbers and if you enjoy it, it shouldn't matter. Unfortunately the ultimate truth is we are heavily judged on our CPM by everyone, potential partners, sponsors, collaborative opportunities and even by twitch themselves. So yes, it does matter. It matters a lot and worst part is no one, will take this issue seriously! I have literally 120 days of information from twitches 9th longest subathon (unrecognized of course as again small potatoes) that would absolutely show this being a consistent issue. I don't know what to do, I am at wits end here and I every time I see this happen to me I just want to curl up into a ball and let the world wash over me. I need help, i need someone to look at this, because I have dedicated my life to entertaining individuals and I do this full time, I dedicate 8 hours a day every day for the past 2 years to this company and the only thing I get is ignored. I know this is a lot and it looks like just some random person complaining, but I need some way to get this in the right hands and not the automated jarble that I've been getting. Its even been acknowledged in the past by a tech for twitch that something might be up. Again, I know I'm not big but were all working under the same logo and should all be taken care of equally! Twitch

Frosted Fricks 💣🐦‍⬛

21,589 просмотров • 1 год назад

hey here is the final result of octopus invaders on nvidia's flagship at full precision. nemotron super 120B on 2x H200 NVL. BF16 unquantized. 287GB of VRAM. hermes agent as the harness. 60 tok/s. first try it autonomously coded for 6 minutes straight. created 11 files. correct project structure. correct load order. started the server. i opened the browser and the result was a blank screen. i did not give up. second try i gave it a precise list of bugs and things to fix. it went back in for another 3 minutes. patched the code. served it again. still blank. so i did what any sane person would do. third try i just said the screen is blank, test it and fix it yourself. and this is where nemotron showed what it actually is. it became a debugger. you can see it in the video. realtime CSS test squares, red screen flashes, hermes agent browser tools, inspecting its own output. it built the parallax background with planets and comets. it rendered a rocket ship that tracks your mouse with fire and bullet physics. the aesthetic is real. but no enemies spawn. no collision. not playable. what surprised me is qwen 27B one shotted this exact game on a single RTX 3090 at Q4 quant. and here is nvidia's flagship at full precision on enterprise hardware needing 3 tries and still not getting there. that makes my hope high for the undisputed qwen 122B which is about to face the same test next. same hardware. same prompt and same harness. lets see if it one shots or not. full session in the video. no cuts. 5x speed.

Sudo su

11,011 просмотров • 4 месяцев назад

hey if you're thinking about running qwopus (the claude opus distilled qwen 3.5 27B) as a coding agent, this might save you a few hours. i tested both the base and the distilled version on the same hardware. single RTX 3090. same prompt. same context. same everything. the only variable was the model weights. base qwen 3.5 27B built octopus invaders in 13 minutes. 1,827 lines across 11 files. zero steering. one scope bug that took 2 lines to fix. game ran. qwopus couldn't finish the same task. enemies overlapping on screen. bullets not firing. controls worked but the game was broken. i had to steer it multiple times and it still didn't produce a playable result. both run at 35 tok/s. both use thinking mode. the distilled version actually has better jinja compatibility and doesn't stall midtask like base does on claude code. for conversation and reasoning it feels sharper. but for multifile autonomous coding where the model needs to coordinate 10+ files without losing track, base wins and it's not close. distillation compresses reasoning patterns but seems to lose precision on complex coordination. the model "thinks" well but can't hold the full picture across files the way base can. tested on opencode (base) and claude code (both). next up is hermes agent framework on base. same hardware. same prompt. comparing agents now, not just models. video below. first half is the distilled model's broken game. second half is what base built on the same 3090. judge for yourself.

Sudo su

45,052 просмотров • 5 месяцев назад

i watched gemma 4 12b build something genuinely impressive today, and then loop itself to death right in front of me. the full run is in the video, sped up but completely uncut, watch it to the end and you will catch the exact moment it stops building and starts looping right in the middle of the work. the task was clean, build a single file gravity simulator, n-body physics, orbits, collisions, running locally on one 3090 through an agent. and for ten minutes it was a joy to watch. it reached for a symplectic integrator on its own, the correct one, the kind that keeps orbits stable instead of spiralling out. real gravity with softening, proper orbital velocities, momentum conserved on collision. the physics was right. the thing actually worked. then on the very last step, writing a few tests to prove its own code, it fell into a loop. not a crash, a loop. it started repeating itself and would not stop. ten more minutes, thirty four thousand tokens into a single answer, the same fragments over and over, until i killed it myself. so it's not that gemma can't code. it did the hard part beautifully. it cannot finish. it cannot hold a long task together without unravelling, and finishing is the entire job in agentic work. here's the part that stings. i run this exact task, same harness, same card, on the chinese open models, qwen especially, and i never see this. they build it, they test it, they stop. every single time. google has the raw capability, you can see it sitting right there in the code, and then the model loops itself to death on a task a 27b from alibaba finishes clean. open weights, apache 2.0, so much to love on paper. i just need it to know when to stop talking.

Sudo su

39,574 просмотров • 2 месяцев назад