
OmedTheVibeCoder
@OmedVibeCodes • 1,441 subscribers
🇩🇪 Software Developer | Extreme AI model testing, honest results, no hype | Building something that could change game development 🎮
Shorts
Videos

Just tested Qwen3.8-Max-Preview on my benchmark. It got absolutely cooked. Not Fable 5 level. Not even remotely close. It doesn't come close to GPT-5.6 or Opus 4.8 Max either. Honestly, marketing it at that level feels incredibly misleading. I wasted my time and money testing this. Qwen, refund me.
OmedTheVibeCoder183,804 Aufrufe • vor 2 Monaten

Nah, this can’t be real. Same GPT-5.6 Sol. Same task. Same thinking level. I switched from regular Codex to the Claude Code harness, and the result instantly became way better. Claude Code is somehow making every model cook. GPT-5.6 SOL WAS ALREADY TOP TIER!! 😭
OmedTheVibeCoder127,040 Aufrufe • vor 2 Monaten

Muse Spark 1.3 is not that good, and Gemini Flash 3.8 is just pure benchmark-maxxing. Gemini even choked on the first run and needed a second attempt just to produce a working scene. By stripping colliders and enforcing rigorous backend-level logic in 3D space, this benchmark immediately separates true reasoning from trained test shortcuts.
OmedTheVibeCoder29,403 Aufrufe • vor 19 Tagen

This is probably my most important post about DeepSeek V4 Flash. Do NOT judge this model after testing only one harness. Pi, Claude Code, and OpenCode produced dramatically different results—and OpenCode completely changed my verdict. You really need to read the details below.
OmedTheVibeCoder68,681 Aufrufe • vor 1 Monat

This is why my benchmark is not benchmaxxable and is one of the best on twitter Put Gemini 3.8 High against Opus 5 xHigh and the hype does not hold up. It got mogged hard. No real leap here, just obvious benchmark-maxxing I have heavy constraints in this test, so it clearly shows, if a model will be able to still mantain good quality, while having to follow architectural patterns
OmedTheVibeCoder18,850 Aufrufe • vor 20 Tagen

Opus 5 High might be the most ridiculous model I’ve ever tested. It produced this result in around ONE HOUR. Fable 5 needed almost 3 hours. GPT-5.6 Sol worked for over an hour and still created almost nothing compared to this. Opus 5 didn’t just win. It completely OUTCLASSED them.
OmedTheVibeCoder39,729 Aufrufe • vor 1 Monat

I think I accidentally created a completely new game genre. I’ve been working on this nonstop for months, and I’m finally ready to show just a tiny fraction of the gameplay. No traditional 3D models. Everything is generated in Three.js through pure math. This is completely insane guys. The actual main feature is still hidden—and it changes EVERYTHING. #threejs
OmedTheVibeCoder37,476 Aufrufe • vor 1 Monat

Guys, stop underestimating the harness. GPT-5.6 Sol through Claude Code vs Codex produced a GIGANTIC difference in my benchmark. But Opus 5 is the real innovation: High, xHigh, and Max don’t just feel like different thinking levels—they feel like COMPLETELY DIFFERENT MODELS. It’s basically like getting three or four Opus models in one.
OmedTheVibeCoder36,074 Aufrufe • vor 1 Monat

This completely changes how I look at AI benchmarks. Qwen 3.8 Max Preview felt nearly unusable in OpenCode and Qwen Code. With the Claude Code harness, it suddenly came surprisingly close to Kimi K3 and Fable 5. Same model. Same task. Different harness. Wildly different result. How is this even possible?
OmedTheVibeCoder30,223 Aufrufe • vor 2 Monaten

DeepSeek V4 Flash completely failed this test. Compared to Opus 4.8, GPT-5.6 Sol, Kimi K3, and the other current top models, it’s honestly not good. The gap is massive. Against Luna, however, it performs similarly well. Both have different strengths and weaknesses, so I still need to test them more. The weird part: the Pi harness performed worse than OpenCode for DeepSeek. More tests are coming, guys. Don’t worry
OmedTheVibeCoder26,166 Aufrufe • vor 1 Monat
Keine weiteren Inhalte verfügbar