正在加载视频...

视频加载失败

🚨codex users!!!! I tested GPT-6 Sol on $20 plan I set the effort to "MAX" I am NOT happy with the results - 15 Million tokens - 7% weekly limit used - $4.21 API cost - 35 minutes of work So... you are getting 857 Million tokens of GPT-6...

82,418 次观看 • 4 天前 •via X (Twitter)

83 条评论

shownotover 的头像
shownotover4 天前

This is the config for anyone who wanna dive deep

Usman Bashir 的头像
Usman Bashir4 天前

I understand setting Sol to max effort but don't set Opus 5.5 to max effort. It actually gives better results at medium effort than any other reasoning. Use Opus 5.5 on medium on your $20 plan and you won't need to buy Claude Max plan at all.

shownotover 的头像
shownotover4 天前

Yeah I know that but I set these efforts to the Max plan so that I can test out the maximum usage that you can get on each $20 plan. That's all.

Usman Bashir 的头像
Usman Bashir4 天前

Absolutely. I think that Claude's usage might be a bit higher than Codex's -- that's what you found out too.

RobIW 的头像
RobIW4 天前

You put it on token burner mode and complain that it burns tokens Try anything less than max Not saying it’s good or anything but you can’t complain that it burns tokens on token burner mode

shownotover 的头像
shownotover4 天前

It's only for tests bro... I am only calculating the amount of work you get for the price you pay

RobIW 的头像
RobIW4 天前

Alright my mistake I interpreted it as complaining

shownotover 的头像
shownotover4 天前

Guys... I am gonna test out every provider's limits now to make it MORE clear It takes effort so... Thanks!

Arbaz 的头像
Arbaz4 天前

cheap sol. still burns the week

shownotover 的头像
shownotover4 天前

Yeah man, I am fed up with all this

Rayistech 的头像
Rayistech4 天前

All it took was 1 prompt using GPT6 Sol on medium.

shownotover 的头像
shownotover4 天前

Sorry bro 😐 Switch to Claude maybe

Mochi 的头像
Mochi4 天前

yeah it's time to move back to claude,i miss the old opus 4.6 and i think 5.5 has the same feel

shownotover 的头像
shownotover4 天前

It DOES have the same feel bro

Tejas Arlimatti 的头像
Tejas Arlimatti4 天前

Wow curious how you arrived at 1100$ with Opus 5.5?

shownotover 的头像
shownotover4 天前

I didn't test Opus 5.5... only Opus 5. Its also posted on my account... you'll love to see it

Nish 的头像
Nish4 天前

they are not only shit about there usage but there support team is full of AI, even being 20x pro user i regret having it. despite 20x plan is in paused

shownotover 的头像
shownotover4 天前

OpenAI is making so many worse decisions ngl

Nish 的头像
Nish4 天前

month ago i was huge huge fan of @thsottiaux and @sama every single user response is genuinely heard. right now they even scam the banked reset. we are getting 30/40% less usage in banked reset

shownotover 的头像
shownotover4 天前

@thsottiaux @sama Exactly bro... I was a fan too but now... meh

Uncle cat bangkok 的头像
Uncle cat bangkok4 天前

I know!! astra cooked but sol just like yea imma head out and say i cant do it

shownotover 的头像
shownotover4 天前

Sol came later so at least it should have been better than Astra but nah...

Uncle cat bangkok 的头像
Uncle cat bangkok4 天前

I guess Sammy wants to address the astra huge consumption complaint, we Plus users can stay in the loop. I guess any user count at this stage. Astra is amazing though

Absurd 的头像
Absurd4 天前

Yeah it’s exactly the same usage as the old one, it’s rubbish

shownotover 的头像
shownotover4 天前

Honestly... I can't see how I will use Codex in future

VulKan 的头像
VulKan4 天前

yep, how crazy that Anthropic has a much better value for $20 now 😭

shownotover 的头像
shownotover4 天前

I know man... I was a fan of codex just 1 month ago... now its disappointing

Jordan 的头像
Jordan4 天前

My Claude Max 20x already works out to around $1,500/week in API-equivalent usage. And you’re telling me Pro gets $1,100/month? Doesn’t that make it pretty obvious that something is wrong with the calculation?

shownotover 的头像
shownotover4 天前

20x plan gets around $8k worth of usage... And it's not 20x the Pro plan that is why Anthropic got sued

Jordan 的头像
Jordan4 天前

In practice, Claude Max 20x works out to roughly $6,000/month in API-equivalent usage. Based on earlier testing, Max 5x is about half of that, so roughly $3,000/month. And Pro is about 1/5 of Max 5x, which would put it closer to $600/month. So I honestly don’t understand how you ended up measuring $1,100/month for Pro.

shownotover 的头像
shownotover4 天前

Here is the test... Go check it out... I tested with proof how we are getting $1100 of usage monthly

Sahil Panhotra | Indie Builder 的头像
Sahil Panhotra | Indie Builder4 天前

very bad 😭 but opus 5.5 is extremely slow on max

shownotover 的头像
shownotover4 天前

No way!! I'll have to test it out too

Aravind Mohan 的头像
Aravind Mohan4 天前

I agree. I am a plus user and single prompt entire 5hr quota over in 25 min and work incomplete. I used GPT-6 SOL Medium

shownotover 的头像
shownotover4 天前

I used GPT-6 SOL on max It usage is "okay-ish" if you don't code

Victor Jaro 的头像
Victor Jaro4 天前

I’ve been running sol on high and xhigh and I had a completely different experience. Feels like 10x usage compared to astra, I ran out after ~2 hours

shownotover 的头像
shownotover4 天前

I am not saying that it's a bad experience. Yes I can agree with you that GPT-6 Sol can run for 2 hours straight with high effort but the problem is that it is slow. The amount of work that you got done in 2 hours is nowhere near the amount of work that you would have gotten done with Claude Opus 5. Also the amount of usage that you get on Claude and Codex is completely different. Claude gives you a lot more usage. 32 Million tokens vs ~18 Million for same 10% weekly is crazy difference!

Victor Jaro 的头像
Victor Jaro4 天前

At least we get more resets for codex… you can’t have everything in this ai world🥲

skhlgnev 的头像
skhlgnev4 天前

it is very slow and anyway even without using astra my usage is already 60% wtf (im only using luna mostly and sol medium / high)

shownotover 的头像
shownotover4 天前

Bro, that is so rough. You might be having some kind of glitch.

skhlgnev 的头像
skhlgnev4 天前

Maybe I don't know, bro, but what's happening with Codex? I'm on Windows. To get these models, I even had to clear my cache manually. It means that even when I update an app, I don't get the models they have. So many issues that haven't been solved.

shownotover 的头像
shownotover4 天前

Yeah... The Desktop app is slow af... That's why I used CLI ... Might as well use CLI cause it's also a little bit better at caching... Personally speaking

skhlgnev 的头像
skhlgnev4 天前

OK, bro. Thanks, I've got it! I'll give it a try. I'm using the app only because I like reading the text from it, but I'll give the CLI a try again!

shownotover 的头像
shownotover4 天前

No problem... lemme know your experience

Arun 的头像
Arun4 天前

I’d think performing the exact same task/prompt/context etc on these 2 specified models at max would give us a better picture.

shownotover 的头像
shownotover4 天前

I will test it out for you... No problem... Next post is exactly that

Rajantha 的头像
Rajantha4 天前

OpenAI might have a compute problem. With Opus 5.5 being this capable they can't cutting limits like this without loosing subscribers. Thanks for the test. Real world use cases are the only benchmark we can trust right now.

shownotover 的头像
shownotover4 天前

No probs bro... I will test out Claude Opus 5.5 too

catman 的头像
catman4 天前

Tokens are fuel, not miles: a cheap, high-throughput model can still cost more in time if it doesn't finish the task.

shownotover 的头像
shownotover4 天前

That is what I am saying 😂

Harshit Arora 的头像
Harshit Arora4 天前

remember there is batch request in sol which is good for only bulk changes or scanning giant codebases.

shownotover 的头像
shownotover4 天前

Yeah but it doesn't change the fact that we are getting 1/2 usage fr same price

Harshit Arora 的头像
Harshit Arora4 天前

Yes, API prices seem more reasonable sometimes

shownotover 的头像
shownotover4 天前

but rn now they are not 😢

Md. Abu Taher 的头像
Md. Abu Taher4 天前

Just last week it was 350$ of usage. You mean its 250$ now? 😐

shownotover 的头像
shownotover4 天前

Yeah 😂 Move to Claude!!!

Qais 的头像
Qais4 天前

So the result is use Claude 🔥 it’s always the best

shownotover 的头像
shownotover4 天前

Yeah... I think we should switch to Claude for now Don't know if this will be the same for next month

Qais 的头像
Qais4 天前

I think Claude is always better in every case

shownotover 的头像
shownotover4 天前

It wasn't until August... I think they did something like this: 🤣 OpenAI did all the bad things themselves

ミ Α Ω 彡 的头像
ミ Α Ω 彡4 天前

Yea gpt 6 is fricken slow. I don't see any improvements really. Opus all the way.

shownotover 的头像
shownotover4 天前

It's only an update, rather than an upgrade 😂

ミ Α Ω 彡 的头像
ミ Α Ω 彡4 天前

Exactly. Well they said devday is huge and that's coming Tuesday next week. Time will tell.

shownotover 的头像
shownotover4 天前

Really man, I think that'll be flop too

cuandada 的头像
cuandada4 天前

Could you please check the performance of Codex 20$ 6Sol and Claude 20$ Opus 5.5?

shownotover 的头像
shownotover4 天前

Claude Opus 5.5 is on the way... I a travelling right now... Will test it when I get back.

Uriel Bitton 的头像
Uriel Bitton4 天前

last time i tested claude fable 5 vs gpt 5.6, gpt was way better.

shownotover 的头像
shownotover4 天前

Hahahahaha... GPT 5.6 Sol can't touch Fable 5 ngl... I don't know how you tested it... Lemme know about that

Nayan Surya 的头像
Nayan Surya4 天前

yikes, that's rough. speed and accuracy are everything.

shownotover 的头像
shownotover4 天前

Yeah man... I was a codex fan... now, I am not

Creepbrute 的头像
Creepbrute4 天前

Fair. But remember, chat + images also count ontop of that. Codex / Work limits are apart from Chat + Image limits. Claude has a total limit, not unified. 5.6 Sol on high - no matter how many tokens it outputs. Perfect to check your codebase in a zip... 🫢

shownotover 的头像
shownotover4 天前

The thing is simple. I don't use chat or images... Simply give me full usage of what I want... But even that won't be enough ngl

Gia 🃏 的头像
Gia 🃏4 天前

damnn thats sounds like a steal 🥀 857 million tokens for $250 and it still didn't finish the task in 35 minutes

shownotover 的头像
shownotover4 天前

Not a steal... but a scam

orvir 的头像
orvir4 天前

The denominator is doing all the work here. 35 minutes for 7% implies roughly 8.3 hours per week if the burn stays linear; 857M tokens is extrapolated throughput, not an allowance the plan guarantees.

shownotover 的头像
shownotover4 天前

💯💯💯

cxan 的头像
cxan4 天前

yeah mate, disappointed as well. Have been using it as my daily driver today at High and only about 5h of use later it's down more than 30%. I'm still on the $100 codex plan with one reset left but honestly as much as I dislike anthropic I might switch back after this month

shownotover 的头像
shownotover4 天前

Feeling sorry for that bro... Might as well try Claude... I think it's better with $20 plan... Can't say anything for $100 plan

Guido Schoonheim 的头像
Guido Schoonheim4 天前

focus 100% on cost per task not on tokens, and stop using max

shownotover 的头像
shownotover4 天前

I only tested sir... The cost per task is 2x the price of Claude if we look at the $20 we pay them (not API usage)

Human 的头像
Human4 天前

Max effort is not free intelligence. Token burn without an outcome is just an expensive thinking face.

doug houser 的头像
doug houser4 天前

Opus 5 always gave 2x more usage than sol 5.6 in every plan. Openai made up for it with the resets, I realized how bad they were when they gave a reset after 10 days and I had to use 200% of my weekly quota with a banked reset

shownotover 的头像
shownotover4 天前

They are making themselves worse like this 😐 I hope they'll learn

相关视频

I just compared Claude Code vs Codex vs Cursor CLI The task was to build a Next.js app with Tailwind 4 and shadcn components to collect customer feedback and showcase it with a widget. I gave all three the same prompt and let them go for 30 minutes to see what they came up with. Claude Code with Opus 4.1 Even though I told it to set up the app in the existing project folder, it tried to create a directory for it. After I interrupted and told it not to do that, it built a demo form and landing page with no errors. I had to ask it to make the demo interactive so users could submit a testimonial and preview it. The landing page looked like AI and was pretty basic, but it worked and it was done in a fraction of the time of the others. Total tokens used: 33k Codex with GPT-5 At the end of the 30 minutes I just could not get Codex to produce a working app. It got stuck in a loop of not being able to set up Tailwind 4 and despite many, MANY, attempts, I ended up with a "failed to compile" error. Total tokens used: 102k Cursor Agent with GPT-5 This was the slowest agent by far and a couple of times I actually thought it got stuck in a loop and was close to Ctrl+C'ing to cancel it. The TUI is really nice though, especially how it shows diffs and it did eventually build a working app (after one or two slight errors that needed fixing) The demo was interactive and it had a very minimal design that looked bare but also a lot less like an "AI generated" app than the Opus 4.1 design. It also wasn't too chatty and just did what it needed to do! Code quality was on a par with Opus 4.1, but it did use 5.5x as many tokens to get there. Still cheaper than Opus on a direct comparison but not when you factor in a Claude Code Max subscription. Total tokens: 188k I'll be able to do a proper comparison and record some videos when I'm back from holiday but for now, Opus is still the more capable model out of the box and Claude Code is the more complete CLI product. It will be interesting to see how Cursor evolve their CLI though with commands and subagents because I think with GPT-5 they have a real shot at providing competition for Claude Code if they can optimise output to get similar quality with less tokens. Jump to 0:40 in the video to see the two apps. Which do you think is which? ;)

Ian Nuttall

195,173 次观看 • 1 年前

I tested Claude Code on a fresh account - 1,500 lines of HTML cost me 50% of my window. Full video and summary is here.. I just ran a recorded test on Claude Code with a fresh account (Pro, not Max - my main account was 20x Max) , and the result is honestly insane. The task was trivial: create 3 simple demo HTML pages, around 500 lines each. Roughly 1,500 lines of code total. Nothing massive. Nothing enterprise-grade. Nothing that should meaningfully stress a premium coding product. And yet Claude Code burned through 40% of my 5-hour window almost immediately. I ran the exact same test with Codex, and it consumed only 2%. Then it got even worse: after the session ended, I did absolutely nothing for 15 minutes, and Claude still ate another 10%. Total: 50% of the 5-hour window gone for a tiny HTML demo. My weekly usage had already started at 2% before I even really used it, and after this tiny test it jumped to 8%. Now let us be generous and assume this entire run used around 30k tokens total. If 30k tokens represents 10% of weekly usage, that implies around 300k tokens per week. That is roughly 1.2M-1.3M tokens per month, and even if you round up aggressively, you are still in the 1.5M token range. Using the Sonnet 4.6 pricing you list: $3 per 1M input tokens $15 per 1M output tokens How exactly is this supposed to make sense for a paid coding product? Because from the user side, this no longer looks like "premium usage protection." It looks like a quota system that is either wildly inefficient, badly broken, or being accounted in a way users are not being told about. And that is before I even get to my main account: my $200 Max plan now dies in a single day. Just a few months ago, similar or heavier usage would last me about a week. So no, I do not buy the "maybe you just used it more" excuse anymore. Something is clearly broken in Claude Code. Either token accounting is broken, context handling is broken, background consumption is broken, or all three. Alex Albert is this really the experience you want users to pay for? Just watch the video. I tried to be very transparent and clear for your team! I was fan of Claude but just disappointed! And if you want, send me the detailed token accounting for this session and let us inspect it together publicly. Because from where I am standing, this is no longer a small pricing annoyance. It looks like something seriously wrong is happening, and users deserve a real explanation.

Hayrettin Tüzel

26,854 次观看 • 6 个月前