Loading video...

Video Failed to Load

Go Home

🚨 codex users!! PLEASE DO NOT BUY CODEX!! I have $20 codex I used GPT-6 Astra of Max I asked it: "Here is the project, create a promo video. Give me your best" and disabled all the skills Now... the results are... BAD: - 2.4 Million tokens ONLY!!!! -...

56,183 views • 4 days ago •via X (Twitter)

75 Comments

noah helms's profile picture
noah helms4 days ago

I’ve found the same thing! Why use codex anymore when they’re cutting usage and releasing terrible models over and over? They’re really struggling

shownotover's profile picture
shownotover4 days ago

same bro... Don't use it now... they were good 2 months ago

Arnold's profile picture
Arnold4 days ago

For Codex users from an EX Codex user... 75% in 11 minutes 👌🤦‍♂️🤦‍♂️🤦‍♂️

shownotover's profile picture
shownotover4 days ago

That's sad 😭 Codex used to be so good

Morris Dweck's profile picture
Morris Dweck4 days ago

Codex is still the best 20 dollar plan. not for value, but because Claude doesn't have anything that matches ma boi luna. 5x and 20x? go to claude. 20 dollar plan? Codex. For Luna.

shownotover's profile picture
shownotover4 days ago

Yeah bro 💯 If you need unlimited coding, go with Codex because of Luna... I can't argue here... no sub does that

Hussain Hashim | Building SundayBack's profile picture
Hussain Hashim | Building SundayBack4 days ago

@shownotover careful with token limits! I hit a wall when I tried maxing out features. Sometimes less really is more.

shownotover's profile picture
shownotover4 days ago

You are right but the problem is, right now, the task was entirely logical and had nothing to read about so it was a good decision to use it on Max. Otherwise it wouldn't make sense. 🥲

NiteLite's profile picture
NiteLite4 days ago

You might get better results at a lower reasoning level. Asking it to do max reasoning for such a simple prompt is likely making the model over-reasoning. This is very far from the "PHD level research" that max reasoning is designed for.

shownotover's profile picture
shownotover4 days ago

Okay sir... Will do a test for medium... Let's see the results

NiteLite's profile picture
NiteLite4 days ago

Interesting, looking forward to seeing the results.

Dmytro Shevchenko 🇺🇦's profile picture
Dmytro Shevchenko 🇺🇦4 days ago

on Pro 200$ you have approximately: 1b tokens on Astra 6b tokens on Sol (I have calculated before astra release) 14b Terra ~120b-140b Luna (not exact because it's between Astra and Sol 😅) Calculations based on usage

shownotover's profile picture
shownotover4 days ago

If I assume that's weekly... Then you get around 400 Million tokens for Opus 5.5 (better than Astra) for only 1/10th of the price... I don't think Codex is worth it now 🥲

Dmytro Shevchenko 🇺🇦's profile picture
Dmytro Shevchenko 🇺🇦4 days ago

yes, I have counted weekly limits :)

Teo Nex's profile picture
Teo Nex4 days ago

I kinda solved this problem, just wrote a huge post about it

shownotover's profile picture
shownotover4 days ago

Tag me in bro.

Teo Nex's profile picture
Teo Nex4 days ago

Here you go, bro. I did a lot of research and built this harness to get more out of my subscription:

Watzon's profile picture
Watzon4 days ago

Your mistake is using max. Don't get me wrong, Astra blows through limits, but you do not need max effort for just about anything.

shownotover's profile picture
shownotover4 days ago

Bro, this was a logical task. That's why I used max effort. I did the same with Opus 5.5 on max (same prompt) and it didn't blow up with the usage limit like the one Astra did.

Watzon's profile picture
Watzon4 days ago

Opus 5.5 also shouldn't be used on max. See Theo's Opus 5.5 video for the reasons why, they're valid. It's not even about usage. Max effort forces the models to think more, which not only leads to more token usage, but it also can lead to worse outcomes because they think themselves in circles.

Shahzeb Naveed's profile picture
Shahzeb Naveed4 days ago

Claude code 👑

shownotover's profile picture
shownotover4 days ago

Using Opus 5.5 as my main orchastrator

GamerzArtist's profile picture
GamerzArtist4 days ago

I want to switch, But can't since I'm banned from Anthropic for being Under 18

shownotover's profile picture
shownotover4 days ago

Try @cognition ... Good usage for all models

Sahil Panhotra | Indie Builder's profile picture
Sahil Panhotra | Indie Builder4 days ago

yeah the quality isn't as good as opus 5.5 but comparing api costs and tokens counts is not right way to compare codex vs claude

shownotover's profile picture
shownotover4 days ago

So... what should be the way? We are paying $20 for both... One gets better quality, more work done tbh... ain't that good?

Sahil Panhotra | Indie Builder's profile picture
Sahil Panhotra | Indie Builder4 days ago

yeah bro i agree but API costs and token counts isn't right way here 😅 codex is heavily efficient (uses way less tokens and costs way less too compared to claude) so you can't compare lemons to apples you getting it right ?

shownotover's profile picture
shownotover4 days ago

Okay I understand the basics: one is token-efficient and the other one uses a lot of tokens. At the end of the day we all care about quality and the amount of work done. Seeing Codex and Claude code in action, I can say to you and write it for you on paper that Claude does two times the work Codex does for the same amount of money.

Sahil Panhotra | Indie Builder's profile picture
Sahil Panhotra | Indie Builder4 days ago

okay bro 😅 so compare it like that make 2 branches and see which one is doing how much work and not just try to burn limits and yeah i know currently Opus 5.5 is doing more work (coz its cheaper a bit now)

yorunoken's profile picture
yorunoken4 days ago

that's really good no? on my sub Astra on low lasts like 5 minutes and dies right after

shownotover's profile picture
shownotover4 days ago

try Claude... and if not, try Cursor

Aapakari's profile picture
Aapakari4 days ago

Quality wins. $5.7 barely dents the $20, so the 29 minutes are the real bill.

shownotover's profile picture
shownotover4 days ago

bro 😭 Why you using AI

ミ Α Ω 彡's profile picture
ミ Α Ω 彡4 days ago

Yeah I have tried with Astra and man it was horrible

shownotover's profile picture
shownotover4 days ago

$20 on Astra is horribly horrible haha still, I only had $20 so... what could have I done 🤣

Shayan Banerjee's profile picture
Shayan Banerjee4 days ago

@thsottiaux @victornunez @OpenAI @sama

shownotover's profile picture
shownotover4 days ago

@thsottiaux @victornunez @OpenAI @sama Nope... they are gonna say "DevDay"

anytan's profile picture
anytan4 days ago

2.4M tokens burned for a promo that still looks mid. that weekly % hit different

shownotover's profile picture
shownotover4 days ago

Yeah bro... But tell me one thing Am I not good enough for a human being that you have to use AI to respond? If I wanted to talk to AI ... Why would I make this post?

Lee Wyatt Corp's profile picture
Lee Wyatt Corp4 days ago

Burning 2.4M tokens on a bad video hurts. We'd rather burn them on our own apps.

shownotover's profile picture
shownotover4 days ago

XFastest is my own app 🥲

Lee Wyatt Corp's profile picture
Lee Wyatt Corp4 days ago

didn't realize XFastest was yours. that 2.4M token run hits different when it's yours

Kekius Pisimus's profile picture
Kekius Pisimus4 days ago

A.I is do good now

shownotover's profile picture
shownotover4 days ago

Its gonna come after us in 10 years haha

RV's profile picture
RV4 days ago

dude

shownotover's profile picture
shownotover4 days ago

What is it sir? 😀

Garfy's profile picture
Garfy4 days ago

why are u so cringy bro....

shownotover's profile picture
shownotover4 days ago

what do you mean?

Jeveloper's profile picture
Jeveloper4 days ago

Now we are monthmaxxing depending on the model 😄

shownotover's profile picture
shownotover4 days ago

Well, gotta save $ where we can 🙂

André Felipe's profile picture
André Felipe4 days ago

Kkkkkkkkkkk você tá de sacanagem com a minha cara certo? É uma piada??? Com o plano plus apenas vc usa o astra no Max, envia um promot literalmente dizendo "crie um vídeo, faça seu melhor" recebe um resultado ruim e vem bostejando aqui? Primeiro aprenda usar ferramentas de ia

shownotover's profile picture
shownotover4 days ago

This is what Claude did with exactly the same prompt. That is why I tested Astra like this.

Thanh Nam's profile picture
Thanh Nam4 days ago

Use Astra to create then Opus fix the "taste"

shownotover's profile picture
shownotover4 days ago

That'll simply mean I am using Astra as implementor and Opus as Orchastrator

Peter's profile picture
Peter4 days ago

Good stats, Claude indeed seems way better!!!

shownotover's profile picture
shownotover4 days ago

yeah bro! I had switched to claude

Nazran F.'s profile picture
Nazran F.4 days ago

Opus 5.5 is 50x better than this for the same use case. I don’t know how OpenAI randomly dropped the bag

shownotover's profile picture
shownotover4 days ago

I even created a video on Opus 5.5 and it was so much better. OpenAI used to be a lot better.

Nazran F.'s profile picture
Nazran F.4 days ago

It’ll flip again when the new models come out I’m sure. It’s just one beating the other in a loop.

Wytze H.'s profile picture
Wytze H.4 days ago

Ok, I’m usually not the guy to comment on this, but this pissed me off a little. Most of the time a model is as good as the prompt you give it. Of course sometimes a model comes around that does better with less, however, more often then not the more specific you explain what you want the better the result will be. I did the same with GPT-6 Sol on High effort.

shownotover's profile picture
shownotover4 days ago

Same prompt and same $20 sub btw Now choice is yours if you get pissed or open your eyes

Adriel S.'s profile picture
Adriel S.4 days ago

perahps high or extra high could've been more than enough; htese models are pretty good and avoid it having it "overthink" itslef in a loop at times

shownotover's profile picture
shownotover4 days ago

Hmm... When I use Opus 5.5 on max, that doesn't hurt... but when I use Astra on Max, that starts hurting?? why is that? I want maximum capability by the Model that I can get out for one task that I want to be the best. am I missing something?

Adam's profile picture
Adam4 days ago

these really work well

shownotover's profile picture
shownotover4 days ago

🥲 I guess if it works for you, I support you bro

Adam's profile picture
Adam4 days ago

Your posting style 🤣 not codex

shownotover's profile picture
shownotover4 days ago

😁 From Social media, I have learnt that the more you have emotions in your content, the more people will relate to it. ofc yapping doesn't count so have to do it like this ☺

Adam's profile picture
Adam4 days ago

Yeah thats why I shout all the time 🤣. The more emotions you can stir the better

Mr Leefy's profile picture
Mr Leefy4 days ago

"First time?"

Javier Jiménez's profile picture
Javier Jiménez4 days ago

Tremendo lameculo de Darío y famboy de anthropic

shownotover's profile picture
shownotover4 days ago

you had some serious brainwash to not see the clear difference between the quality to cost ratio can't argue with you

Javier Jiménez's profile picture
Javier Jiménez4 days ago

A ver todo sabe que anthropic le importa sus cliente empresarial y nos a su usuario normales y lo que hace con 5.5 opus es para su IPO y cuando ya eso pase será la misma mierda que antes y openAI tiene modelo más potente intenamete ok

Conza Walker's profile picture
Conza Walker4 days ago

So let’s see if Tuesday addresses this, I like the idea that the 6th dot is for multi agent mode, and ties into the M&Ms they’re going for

Expansion's profile picture
Expansion4 days ago

no one cares

Crio Songo's profile picture
Crio Songo4 days ago

This is really unacceptable. Codex's quota is too tight compared to Claude, I'd better not buy it for now.

Related Videos

I tested Claude Code on a fresh account - 1,500 lines of HTML cost me 50% of my window. Full video and summary is here.. I just ran a recorded test on Claude Code with a fresh account (Pro, not Max - my main account was 20x Max) , and the result is honestly insane. The task was trivial: create 3 simple demo HTML pages, around 500 lines each. Roughly 1,500 lines of code total. Nothing massive. Nothing enterprise-grade. Nothing that should meaningfully stress a premium coding product. And yet Claude Code burned through 40% of my 5-hour window almost immediately. I ran the exact same test with Codex, and it consumed only 2%. Then it got even worse: after the session ended, I did absolutely nothing for 15 minutes, and Claude still ate another 10%. Total: 50% of the 5-hour window gone for a tiny HTML demo. My weekly usage had already started at 2% before I even really used it, and after this tiny test it jumped to 8%. Now let us be generous and assume this entire run used around 30k tokens total. If 30k tokens represents 10% of weekly usage, that implies around 300k tokens per week. That is roughly 1.2M-1.3M tokens per month, and even if you round up aggressively, you are still in the 1.5M token range. Using the Sonnet 4.6 pricing you list: $3 per 1M input tokens $15 per 1M output tokens How exactly is this supposed to make sense for a paid coding product? Because from the user side, this no longer looks like "premium usage protection." It looks like a quota system that is either wildly inefficient, badly broken, or being accounted in a way users are not being told about. And that is before I even get to my main account: my $200 Max plan now dies in a single day. Just a few months ago, similar or heavier usage would last me about a week. So no, I do not buy the "maybe you just used it more" excuse anymore. Something is clearly broken in Claude Code. Either token accounting is broken, context handling is broken, background consumption is broken, or all three. Alex Albert is this really the experience you want users to pay for? Just watch the video. I tried to be very transparent and clear for your team! I was fan of Claude but just disappointed! And if you want, send me the detailed token accounting for this session and let us inspect it together publicly. Because from where I am standing, this is no longer a small pricing annoyance. It looks like something seriously wrong is happening, and users deserve a real explanation.

Hayrettin Tüzel

26,854 views • 6 months ago

I just compared Claude Code vs Codex vs Cursor CLI The task was to build a Next.js app with Tailwind 4 and shadcn components to collect customer feedback and showcase it with a widget. I gave all three the same prompt and let them go for 30 minutes to see what they came up with. Claude Code with Opus 4.1 Even though I told it to set up the app in the existing project folder, it tried to create a directory for it. After I interrupted and told it not to do that, it built a demo form and landing page with no errors. I had to ask it to make the demo interactive so users could submit a testimonial and preview it. The landing page looked like AI and was pretty basic, but it worked and it was done in a fraction of the time of the others. Total tokens used: 33k Codex with GPT-5 At the end of the 30 minutes I just could not get Codex to produce a working app. It got stuck in a loop of not being able to set up Tailwind 4 and despite many, MANY, attempts, I ended up with a "failed to compile" error. Total tokens used: 102k Cursor Agent with GPT-5 This was the slowest agent by far and a couple of times I actually thought it got stuck in a loop and was close to Ctrl+C'ing to cancel it. The TUI is really nice though, especially how it shows diffs and it did eventually build a working app (after one or two slight errors that needed fixing) The demo was interactive and it had a very minimal design that looked bare but also a lot less like an "AI generated" app than the Opus 4.1 design. It also wasn't too chatty and just did what it needed to do! Code quality was on a par with Opus 4.1, but it did use 5.5x as many tokens to get there. Still cheaper than Opus on a direct comparison but not when you factor in a Claude Code Max subscription. Total tokens: 188k I'll be able to do a proper comparison and record some videos when I'm back from holiday but for now, Opus is still the more capable model out of the box and Claude Code is the more complete CLI product. It will be interesting to see how Cursor evolve their CLI though with commands and subagents because I think with GPT-5 they have a real shot at providing competition for Claude Code if they can optimise output to get similar quality with less tokens. Jump to 0:40 in the video to see the two apps. Which do you think is which? ;)

Ian Nuttall

195,173 views • 1 year ago