Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🚨 codex users!! I tested GPT-6.1 Sol at max and.... it is BAD I gave it a prompt to create a motion graphics video it used: - 2% weekly limit (15% 5 hr limit) - 1.7 Million tokens - 58 minutes of work - $1.8 API usage So... the...

28,875 Aufrufe • vor 1 Tag •via X (Twitter)

90 Kommentare

Profilbild von shownotover
shownotovervor 1 Tag

If you wanna grow on X

Profilbild von Harshith
Harshithvor 1 Tag

just $360 we get nearly $1k with Claude pro

Profilbild von shownotover
shownotovervor 1 Tag

Yeah I know that... I made 4 posts about it...

Profilbild von John Rood
John Roodvor 1 Tag

the trap is a task with no verifier: nothing can fail, so nothing ends the loop and it just keeps polishing. the runs in my fleet that burn budget always look like this.

Profilbild von shownotover
shownotovervor 1 Tag

Bro, it has to verify itself. It was literally written inside the instructions to watch the video before shipping it. This is the final output after an hour of work. Opus did so much better than this.

Profilbild von :
:vor 1 Tag

@johnroodepic where is the opus video for comparison? i smell bullshit

Profilbild von shownotover
shownotovervor 1 Tag

@johnroodepic Is this bullshit by Opus 5.5??

Profilbild von Anupam Haldkar 
Anupam Haldkar vor 1 Tag

Sol is cheap compute for agents, not a director, $1.8 and 58 min for that clip is the actual product review.

Profilbild von Ravi Kushwaha
Ravi Kushwahavor 1 Tag

hahahahhaa

Profilbild von shownotover
shownotovervor 1 Tag

🫠

Profilbild von Sahil Panhotra | Indie Builder
Sahil Panhotra | Indie Buildervor 1 Tag

opus is better

Profilbild von shownotover
shownotovervor 1 Tag

It is not just better, it's so much better 🥲

Profilbild von Yipeng | Daily AI
Yipeng | Daily AIvor 1 Tag

This is a useful distinction: low quota burn doesn’t help much if the output needs a full redo. I’d love to see the same prompt and workflow run through Sol, Astra, and Opus side by side—quality per usable result is the number that matters.

Profilbild von shownotover
shownotovervor 1 Tag

All of them are on my profile... You can have a look and then decide

Profilbild von Yipeng | Daily AI
Yipeng | Daily AIvor 1 Tag

Thanks, I’ll take a look. The side-by-side outputs are exactly what I was hoping to see. Were Sol, Astra, and Opus given the same prompt and workflow?

Profilbild von shownotover
shownotovervor 1 Tag

Yes. Exactly same prompts, same env and everything was done under keen observation by yours truely 🙂

Profilbild von Yipeng | Daily AI
Yipeng | Daily AIvor 1 Tag

That makes the comparison much more useful. Thanks for controlling the prompt and environment 🙂 Did you measure how much human cleanup each final video needed? That would help put the cost and quality differences in context.

Profilbild von shownotover
shownotovervor 1 Tag

Well I didn't control anything at all. I just took the output and posted on the post.

Profilbild von Prompt Heat
Prompt Heatvor 1 Tag

claude 20 dolar > gpt 20 dolar acc

Profilbild von shownotover
shownotovervor 1 Tag

100% agree with that bro.

Profilbild von Aden
Adenvor 1 Tag

That ribbon looks great for 58 minutes and $1.8

Profilbild von shownotover
shownotovervor 1 Tag

Nah... Check Claude $20 ... This usage is crap compared to Claude $20

Profilbild von Jack Dev
Jack Devvor 1 Tag

Try at xhigh, I've been running planner xhigh and reviewer xhigh with the implementation as medium/low

Profilbild von shownotover
shownotovervor 1 Tag

I know the usage limit is very, very good. The only problem I have right now is the quality that you are getting for the price you are paying on Codex and on Claude.

Profilbild von Jack Dev
Jack Devvor 1 Tag

Yeah, OpenAI seems to be focusing more on cost now. which is great for users who doesn't have a lot of money like me, switching between the two works too but every week another model beat another model, back and forth so i stay with openai

Profilbild von Miro Jomaa
Miro Jomaavor 1 Tag

I prefer this over opus. Opus felt like adhd for the sake of something moving

Profilbild von shownotover
shownotovervor 1 Tag

I mean, it's your choice but check the results opus got me... They are crazy

Profilbild von Miro Jomaa
Miro Jomaavor 1 Tag

yeah opus one is better, just a bit less adhd. I also noticed codex basicly does the same design. small text over the headline text and the end of the headline text is colored.

Profilbild von Silicon Rumors
Silicon Rumorsvor 1 Tag

I don't think this accurately shows performance of the model, it only shows that performance in making these videos.

Profilbild von shownotover
shownotovervor 1 Tag

What will show accurate performance then? I'll test that out

Profilbild von Silicon Rumors
Silicon Rumorsvor 1 Tag

I think complex coding and tasks will show more accurate performance because i believe its more tuned for that anyways rather than making graphic videos like these

Profilbild von shownotover
shownotovervor 1 Tag

Okay explain to me one: what should I test and what exactly should the output feel like? I have to post these results and if I don't post the results in quality (and/or if it's difficult for a normal person to see the difference), then I believe I cannot conduct the tests because that wouldn't actually make sense to anybody. I want you to come up with a test that you think would best fit every model that I test, in the same environment with the same prompt. I'll test these out for you and then paste the results right here along with the usages and everything.

Profilbild von Thanh Nam
Thanh Namvor 1 Tag

Opus

Profilbild von shownotover
shownotovervor 1 Tag

Good choice

Profilbild von Anima
Animavor 1 Tag

Astra 6.1 got delayed unfortunately So no good motion design OAI users

Profilbild von shownotover
shownotovervor 1 Tag

It's delayed until late October so... We are gonna have plenty of time to use Claude

Profilbild von Darhan
Darhanvor 1 Tag

The shit test

Profilbild von shownotover
shownotovervor 1 Tag

What do you mean 😭

Profilbild von CauseWhyNot
CauseWhyNotvor 1 Tag

Better than anything I can do but pretty sad compared to opus 🥲

Profilbild von shownotover
shownotovervor 1 Tag

Vibe coders can't do this 😔

Profilbild von CauseWhyNot
CauseWhyNotvor 1 Tag

😭 ouch 😂

Profilbild von Victor Pimentel
Victor Pimentelvor 1 Tag

Horrível.

Profilbild von shownotover
shownotovervor 1 Tag

They really gotta improve ngl

Profilbild von Eric
Ericvor 1 Tag

“I gave ai 1 task out of millions of tasks it can do and based on this 1 task you shouldn’t buy this “

Profilbild von shownotover
shownotovervor 1 Tag

Well it's just not one task. This is the second task it's failing on... But yeah you can stay on Codex. Don't worry about that.

Profilbild von Gabriele Bolognese | SaaS founder
Gabriele Bolognese | SaaS foundervor 1 Tag

claude is ALWAYS one step above

Profilbild von shownotover
shownotovervor 1 Tag

💯 agreed

Profilbild von JP
JPvor 1 Tag

Can’t compare it with Opus 5.5 or sonnet 5:5

Profilbild von Arnold
Arnoldvor 1 Tag

Thanks for doing thi Tests for us and being honest 🤝

Profilbild von shownotover
shownotovervor 1 Tag

No problem... More tests coming soon

Profilbild von HASSAN
HASSANvor 1 Tag

Scamgpt 6.1 sol is not eve close to sonnet 5.

Profilbild von shownotover
shownotovervor 1 Tag

I know and this is very sad 😭

Profilbild von Josh
Joshvor 1 Tag

We would not call it bad if it wasn’t for opus.

Profilbild von shownotover
shownotovervor 1 Tag

This is their latest 2nd gen model 🥲

Profilbild von AnimeVibeX
AnimeVibeXvor 1 Tag

Which Plan?

Profilbild von shownotover
shownotovervor 1 Tag

$20 codex plan

Profilbild von AnimeVibeX
AnimeVibeXvor 1 Tag

I bought Claude );

Profilbild von shownotover
shownotovervor 1 Tag

I have both and I test them every day or two.

Profilbild von AnimeVibeX
AnimeVibeXvor 1 Tag

Opus 5.5 still Beats every ai Model

Profilbild von Shahzeb Naveed
Shahzeb Naveedvor 1 Tag

$20 codex still ain't worth it 😢

Profilbild von shownotover
shownotovervor 1 Tag

Better buy a Claude $20 account

Profilbild von hisoka
hisokavor 1 Tag

It’s amazing. So good and so efficient.

Profilbild von shownotover
shownotovervor 1 Tag

And painfully slow

Profilbild von hisoka
hisokavor 1 Tag

Not for me

Profilbild von avhq
avhqvor 1 Tag

Space bunny...

Profilbild von shownotover
shownotovervor 1 Tag

I haven't tested it but it's free and very good... Free users should definitely use that 🙂

Profilbild von avhq
avhqvor 1 Tag

No... It's better at ui and definitely motion that chatgpt. There's so many stunning posts on x

Profilbild von Peter
Petervor 1 Tag

Happy to see your initial feelings changed after you tested; seemingly, the big SELLING point is the cost-effectiveness. I am super happy that we finally have many options to choose from (and all are OK or awesome quality)

Profilbild von shownotover
shownotovervor 1 Tag

Yeah ... Now it's about if you want more quality for more cost or wanna be cost effective 🥲

Profilbild von Peter
Petervor 1 Tag

yap, but AT least you have choices!

Profilbild von Linus
Linusvor 1 Tag

A final visual check should be non-negotiable here.

Profilbild von shownotover
shownotovervor 1 Tag

That's on you bro 🫵

Profilbild von cxan
cxanvor 1 Tag

You remember long the other models took to make the video? Have been using sol 6.1 this morning and indeed it felt very slow. Usage seems decent but slowing down tokens is also a way to effectively reduce it....

Profilbild von shownotover
shownotovervor 1 Tag

You know what? I feel like we are "feeling" that the usage is great because they have slowed down the models. If the model was fast enough, the usage would have drained faster. To be honest, this is what I feel.

Profilbild von cxan
cxanvor 1 Tag

yes honestly I really really hope this isn't the case but I also thought about that. Honestly, the only thing thats keeping me from cancelling my $100 plan right now is my hopes in 6.1 sol being a decent daily-driver with good usage. If that now turns out false I'm back to claude

Profilbild von dhru
dhruvor 1 Tag

i bought it few days before claude launched opus 5.5

Profilbild von shownotover
shownotovervor 1 Tag

That hurts bro. 😭

Profilbild von dhru
dhruvor 1 Tag

i know, will switch next month now 😭

Profilbild von Kekius Pisimus
Kekius Pisimusvor 1 Tag

Why the hell bro like GPT model so much???

Profilbild von shownotover
shownotovervor 1 Tag

Where did I say I like it? I'm just giving you guys information about whether or not you should buy a Codex $20 account.

Profilbild von Kekius Pisimus
Kekius Pisimusvor 1 Tag

You said you will delete codex and now you're using it 😂

Profilbild von shownotover
shownotovervor 1 Tag

I have a $20 sub and I don't wanna let my $20 go to waste just because I dont like it

Profilbild von Kekius Pisimus
Kekius Pisimusvor 1 Tag

I think I said to you a long time ago... and you're still using codex. 😂

Profilbild von shownotover
shownotovervor 1 Tag

it may have been a week ago bro.

Profilbild von Nabendu Biswas
Nabendu Biswasvor 1 Tag

I am testing now and gave the same prompt as I gave to Sol 6, Astra, Opus 5.5 and Sonnet 5.5. It is to create a Subway surfer game without Blender. Let's see what it does and will post tomorrow

Profilbild von shownotover
shownotovervor 1 Tag

Cool... Lemme know about that

Profilbild von James Malsawm
James Malsawmvor 1 Tag

If Opus 5.5 don't come out, Sol 6.1 definitely is awesome model.

Profilbild von shownotover
shownotovervor 1 Tag

What? Opus 5.5 is already out

Profilbild von Cuong Quach
Cuong Quachvor 1 Tag

Why not build a page to track those numbers? It’d make them much easier to compare.

Profilbild von shownotover
shownotovervor 1 Tag

I'd love to... I should definitely build one 🫠

Ähnliche Videos

I tested Claude Code on a fresh account - 1,500 lines of HTML cost me 50% of my window. Full video and summary is here.. I just ran a recorded test on Claude Code with a fresh account (Pro, not Max - my main account was 20x Max) , and the result is honestly insane. The task was trivial: create 3 simple demo HTML pages, around 500 lines each. Roughly 1,500 lines of code total. Nothing massive. Nothing enterprise-grade. Nothing that should meaningfully stress a premium coding product. And yet Claude Code burned through 40% of my 5-hour window almost immediately. I ran the exact same test with Codex, and it consumed only 2%. Then it got even worse: after the session ended, I did absolutely nothing for 15 minutes, and Claude still ate another 10%. Total: 50% of the 5-hour window gone for a tiny HTML demo. My weekly usage had already started at 2% before I even really used it, and after this tiny test it jumped to 8%. Now let us be generous and assume this entire run used around 30k tokens total. If 30k tokens represents 10% of weekly usage, that implies around 300k tokens per week. That is roughly 1.2M-1.3M tokens per month, and even if you round up aggressively, you are still in the 1.5M token range. Using the Sonnet 4.6 pricing you list: $3 per 1M input tokens $15 per 1M output tokens How exactly is this supposed to make sense for a paid coding product? Because from the user side, this no longer looks like "premium usage protection." It looks like a quota system that is either wildly inefficient, badly broken, or being accounted in a way users are not being told about. And that is before I even get to my main account: my $200 Max plan now dies in a single day. Just a few months ago, similar or heavier usage would last me about a week. So no, I do not buy the "maybe you just used it more" excuse anymore. Something is clearly broken in Claude Code. Either token accounting is broken, context handling is broken, background consumption is broken, or all three. Alex Albert is this really the experience you want users to pay for? Just watch the video. I tried to be very transparent and clear for your team! I was fan of Claude but just disappointed! And if you want, send me the detailed token accounting for this session and let us inspect it together publicly. Because from where I am standing, this is no longer a small pricing annoyance. It looks like something seriously wrong is happening, and users deserve a real explanation.

Hayrettin Tüzel

26,854 Aufrufe • vor 6 Monaten

I just compared Claude Code vs Codex vs Cursor CLI The task was to build a Next.js app with Tailwind 4 and shadcn components to collect customer feedback and showcase it with a widget. I gave all three the same prompt and let them go for 30 minutes to see what they came up with. Claude Code with Opus 4.1 Even though I told it to set up the app in the existing project folder, it tried to create a directory for it. After I interrupted and told it not to do that, it built a demo form and landing page with no errors. I had to ask it to make the demo interactive so users could submit a testimonial and preview it. The landing page looked like AI and was pretty basic, but it worked and it was done in a fraction of the time of the others. Total tokens used: 33k Codex with GPT-5 At the end of the 30 minutes I just could not get Codex to produce a working app. It got stuck in a loop of not being able to set up Tailwind 4 and despite many, MANY, attempts, I ended up with a "failed to compile" error. Total tokens used: 102k Cursor Agent with GPT-5 This was the slowest agent by far and a couple of times I actually thought it got stuck in a loop and was close to Ctrl+C'ing to cancel it. The TUI is really nice though, especially how it shows diffs and it did eventually build a working app (after one or two slight errors that needed fixing) The demo was interactive and it had a very minimal design that looked bare but also a lot less like an "AI generated" app than the Opus 4.1 design. It also wasn't too chatty and just did what it needed to do! Code quality was on a par with Opus 4.1, but it did use 5.5x as many tokens to get there. Still cheaper than Opus on a direct comparison but not when you factor in a Claude Code Max subscription. Total tokens: 188k I'll be able to do a proper comparison and record some videos when I'm back from holiday but for now, Opus is still the more capable model out of the box and Claude Code is the more complete CLI product. It will be interesting to see how Cursor evolve their CLI though with commands and subagents because I think with GPT-5 they have a real shot at providing competition for Claude Code if they can optimise output to get similar quality with less tokens. Jump to 0:40 in the video to see the two apps. Which do you think is which? ;)

Ian Nuttall

195,173 Aufrufe • vor 1 Jahr