
OmedTheVibeCoder
@OmedVibeCodes • 1,221 subscribers
🇩🇪 Software Developer | Extreme AI model testing, honest results, no hype | Building something that could change game development 🎮
Shorts
Videos

DeepSeek V4 Flash completely failed this test. Compared to Opus 4.8, GPT-5.6 Sol, Kimi K3, and the other current top models, it’s honestly not good. The gap is massive. Against Luna, however, it performs similarly well. Both have different strengths and weaknesses, so I still need to test them more. The weird part: the Pi harness performed worse than OpenCode for DeepSeek. More tests are coming, guys. Don’t worry
OmedTheVibeCoder25,844 Aufrufe • vor 7 Tagen

Guys, stop underestimating the harness. GPT-5.6 Sol through Claude Code vs Codex produced a GIGANTIC difference in my benchmark. But Opus 5 is the real innovation: High, xHigh, and Max don’t just feel like different thinking levels—they feel like COMPLETELY DIFFERENT MODELS. It’s basically like getting three or four Opus models in one.
OmedTheVibeCoder35,970 Aufrufe • vor 13 Tagen
Keine weiteren Inhalte verfügbar