Загрузка видео...
Не удалось загрузить видео
Claude Opus 5.5 completely MOGS GPT 6 Sol. Same prompt. Same task. Look at the results. Opus 5.5 has way better taste and design. It is not subtle. You can see it immediately. GPT 6 Sol is a step up from GPT 5.6 Sol. But next to Opus 5.5... show more
55,238 просмотров • 7 дней назад •via X (Twitter)
Комментарии: 42

GPT 6 Sol is half the price of Opus 5.5. You never mentioned that, why not? One is literally half the cost of the other, whilst offering comparable output.

Fair point, cost matters too

Man, since I'm not creating a FPS game and I need things done with a good ratio of price and performence I prefer gpt 6 sol.

Devastating to say the least…

Sol looks better imo, the only difference is Sol had a pistol and Opus had an automatic. Sol is also half the price, so this doesn’t seem like a great comparison, but it looked like Sol won to me.

Opus costs 5x more and took 6x longer, so this comparison doesn’t really make sense. Opus 5.5 costs roughly 2× more per task than Astra, so that would be a much more reasonable comparison.

Fair, speed and cost count too

Curious as to why you use GPT 6 Sol and not GPT 6 Astra? And what effort level for each (max or xtra high?), the total cost, total tokens and total time it took to create. Then I can make a better determination. Details please. Thanks.

What about Opus 5.5 vs Gpt 6 Astra?

I think if you take into account speed/cost, they both have a place in your work flow. I do plenty of boring work, and Sol will satisfy that end much better than Opus.

Agreed, both have their lane

Opus cooked as always

It really did on this one

Anthropic still holds the one-shot crown. 👑

For now at least

The way Anthropic plays with OpenAI is funny💀

All we care about is improvement and cost. We got both today, that’s a W. We are at the lowest point AI will ever be. Think about that.

Big W for everyone today

yeah, the shit yellow crab absolutely smoked GPT here. frontend taste carries through everything, the overall style, the detail in the models, the textures, even that “soul” artists keep going on about.

This side-by-side is so helpful. Taste is hard to quantify, but you can feel it right away. Curious if it stays good on messy real-world prompts too.

Hard to benchmark, easy to see

Bro Opus 5.5 is the best. I hope in subscription it's good usage as well and doesn't feel like fable.

Same, usage limits will decide it

True

El mismo prompt no mide el gusto. Mide qué modelo acierta más con el gusto de quien escribió ese prompt. La comparación útil empieza cuando defines criterios antes de ver el resultado.

I’m rooting for OpenAI, but let’s be honest—they fell flat on their face today.

Test same level model lol 😆

Sol is not even close to opus

Slop Bench at it again

Same prompt, same task, same benchmark vendor incentives. The real test is not which model wins the demo, it is which one is still the cheapest per correct answer six months from now.

Try this on Astra Ultra Max PLEASE

On average, GPT-6 Sol is 5.6x cheaper than Opus 5.5 for the task based on AA

unblock @argofowl, he comes in peace

Same prompt is a good control. I'd want to see it hold across a few different briefs before calling it taste — one task is where that usually falls apart.

the show case yu do are built in web game/ simulation which will never be used for big project, gpt 6 for exemple is so much better at using UE5 right now than opus 5.5

mind sharing prompt and resources you used to create it?

What is the prompt

Opus 5.5 is only "good" at delivering UI stuff. It's like a mid-senior level engineer. The real deal is GPT Sol/Astra. You can get a lot more done with it.

@bridgemindai Does it still feel lazy and slop?

Not the same price idiot 🤦♂️

A visual comparison is useful when the rubric is fixed before the outputs are seen. Measure hierarchy, spacing, typography, responsive behavior, and edit distance across the same prompt and assets; “better taste” becomes a reproducible result instead of a creator preference.

Was the same amount of money used for both ?
