Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🚨 codex users!! PLEASE DO NOT BUY CODEX!! I have $20 codex I used GPT-6 Astra of Max I asked it: "Here is the project, create a promo video. Give me your best" and disabled all the skills Now... the results are... BAD: - 2.4 Million tokens ONLY!!!! -...

56,183 Aufrufe • vor 4 Tagen •via X (Twitter)

75 Kommentare

Profilbild von noah helms
noah helmsvor 4 Tagen

I’ve found the same thing! Why use codex anymore when they’re cutting usage and releasing terrible models over and over? They’re really struggling

Profilbild von shownotover
shownotovervor 4 Tagen

same bro... Don't use it now... they were good 2 months ago

Profilbild von Arnold
Arnoldvor 4 Tagen

For Codex users from an EX Codex user... 75% in 11 minutes 👌🤦‍♂️🤦‍♂️🤦‍♂️

Profilbild von shownotover
shownotovervor 4 Tagen

That's sad 😭 Codex used to be so good

Profilbild von Morris Dweck
Morris Dweckvor 4 Tagen

Codex is still the best 20 dollar plan. not for value, but because Claude doesn't have anything that matches ma boi luna. 5x and 20x? go to claude. 20 dollar plan? Codex. For Luna.

Profilbild von shownotover
shownotovervor 4 Tagen

Yeah bro 💯 If you need unlimited coding, go with Codex because of Luna... I can't argue here... no sub does that

Profilbild von Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBackvor 4 Tagen

@shownotover careful with token limits! I hit a wall when I tried maxing out features. Sometimes less really is more.

Profilbild von shownotover
shownotovervor 4 Tagen

You are right but the problem is, right now, the task was entirely logical and had nothing to read about so it was a good decision to use it on Max. Otherwise it wouldn't make sense. 🥲

Profilbild von NiteLite
NiteLitevor 4 Tagen

You might get better results at a lower reasoning level. Asking it to do max reasoning for such a simple prompt is likely making the model over-reasoning. This is very far from the "PHD level research" that max reasoning is designed for.

Profilbild von shownotover
shownotovervor 4 Tagen

Okay sir... Will do a test for medium... Let's see the results

Profilbild von NiteLite
NiteLitevor 4 Tagen

Interesting, looking forward to seeing the results.

Profilbild von Dmytro Shevchenko 🇺🇦
Dmytro Shevchenko 🇺🇦vor 4 Tagen

on Pro 200$ you have approximately: 1b tokens on Astra 6b tokens on Sol (I have calculated before astra release) 14b Terra ~120b-140b Luna (not exact because it's between Astra and Sol 😅) Calculations based on usage

Profilbild von shownotover
shownotovervor 4 Tagen

If I assume that's weekly... Then you get around 400 Million tokens for Opus 5.5 (better than Astra) for only 1/10th of the price... I don't think Codex is worth it now 🥲

Profilbild von Dmytro Shevchenko 🇺🇦
Dmytro Shevchenko 🇺🇦vor 4 Tagen

yes, I have counted weekly limits :)

Profilbild von Teo Nex
Teo Nexvor 4 Tagen

I kinda solved this problem, just wrote a huge post about it

Profilbild von shownotover
shownotovervor 4 Tagen

Tag me in bro.

Profilbild von Teo Nex
Teo Nexvor 4 Tagen

Here you go, bro. I did a lot of research and built this harness to get more out of my subscription:

Profilbild von Watzon
Watzonvor 4 Tagen

Your mistake is using max. Don't get me wrong, Astra blows through limits, but you do not need max effort for just about anything.

Profilbild von shownotover
shownotovervor 4 Tagen

Bro, this was a logical task. That's why I used max effort. I did the same with Opus 5.5 on max (same prompt) and it didn't blow up with the usage limit like the one Astra did.

Profilbild von Watzon
Watzonvor 4 Tagen

Opus 5.5 also shouldn't be used on max. See Theo's Opus 5.5 video for the reasons why, they're valid. It's not even about usage. Max effort forces the models to think more, which not only leads to more token usage, but it also can lead to worse outcomes because they think themselves in circles.

Profilbild von Shahzeb Naveed
Shahzeb Naveedvor 4 Tagen

Claude code 👑

Profilbild von shownotover
shownotovervor 4 Tagen

Using Opus 5.5 as my main orchastrator

Profilbild von GamerzArtist
GamerzArtistvor 4 Tagen

I want to switch, But can't since I'm banned from Anthropic for being Under 18

Profilbild von shownotover
shownotovervor 4 Tagen

Try @cognition ... Good usage for all models

Profilbild von Sahil Panhotra | Indie Builder
Sahil Panhotra | Indie Buildervor 4 Tagen

yeah the quality isn't as good as opus 5.5 but comparing api costs and tokens counts is not right way to compare codex vs claude

Profilbild von shownotover
shownotovervor 4 Tagen

So... what should be the way? We are paying $20 for both... One gets better quality, more work done tbh... ain't that good?

Profilbild von Sahil Panhotra | Indie Builder
Sahil Panhotra | Indie Buildervor 4 Tagen

yeah bro i agree but API costs and token counts isn't right way here 😅 codex is heavily efficient (uses way less tokens and costs way less too compared to claude) so you can't compare lemons to apples you getting it right ?

Profilbild von shownotover
shownotovervor 4 Tagen

Okay I understand the basics: one is token-efficient and the other one uses a lot of tokens. At the end of the day we all care about quality and the amount of work done. Seeing Codex and Claude code in action, I can say to you and write it for you on paper that Claude does two times the work Codex does for the same amount of money.

Profilbild von Sahil Panhotra | Indie Builder
Sahil Panhotra | Indie Buildervor 4 Tagen

okay bro 😅 so compare it like that make 2 branches and see which one is doing how much work and not just try to burn limits and yeah i know currently Opus 5.5 is doing more work (coz its cheaper a bit now)

Profilbild von yorunoken
yorunokenvor 4 Tagen

that's really good no? on my sub Astra on low lasts like 5 minutes and dies right after

Profilbild von shownotover
shownotovervor 4 Tagen

try Claude... and if not, try Cursor

Profilbild von Aapakari
Aapakarivor 4 Tagen

Quality wins. $5.7 barely dents the $20, so the 29 minutes are the real bill.

Profilbild von shownotover
shownotovervor 4 Tagen

bro 😭 Why you using AI

Profilbild von ミ Α Ω 彡
ミ Α Ω 彡vor 4 Tagen

Yeah I have tried with Astra and man it was horrible

Profilbild von shownotover
shownotovervor 4 Tagen

$20 on Astra is horribly horrible haha still, I only had $20 so... what could have I done 🤣

Profilbild von Shayan Banerjee
Shayan Banerjeevor 4 Tagen

@thsottiaux @victornunez @OpenAI @sama

Profilbild von shownotover
shownotovervor 4 Tagen

@thsottiaux @victornunez @OpenAI @sama Nope... they are gonna say "DevDay"

Profilbild von anytan
anytanvor 4 Tagen

2.4M tokens burned for a promo that still looks mid. that weekly % hit different

Profilbild von shownotover
shownotovervor 4 Tagen

Yeah bro... But tell me one thing Am I not good enough for a human being that you have to use AI to respond? If I wanted to talk to AI ... Why would I make this post?

Profilbild von Lee Wyatt Corp
Lee Wyatt Corpvor 4 Tagen

Burning 2.4M tokens on a bad video hurts. We'd rather burn them on our own apps.

Profilbild von shownotover
shownotovervor 4 Tagen

XFastest is my own app 🥲

Profilbild von Lee Wyatt Corp
Lee Wyatt Corpvor 4 Tagen

didn't realize XFastest was yours. that 2.4M token run hits different when it's yours

Profilbild von Kekius Pisimus
Kekius Pisimusvor 4 Tagen

A.I is do good now

Profilbild von shownotover
shownotovervor 4 Tagen

Its gonna come after us in 10 years haha

Profilbild von RV
RVvor 4 Tagen

dude

Profilbild von shownotover
shownotovervor 4 Tagen

What is it sir? 😀

Profilbild von Garfy
Garfyvor 4 Tagen

why are u so cringy bro....

Profilbild von shownotover
shownotovervor 4 Tagen

what do you mean?

Profilbild von Jeveloper
Jevelopervor 4 Tagen

Now we are monthmaxxing depending on the model 😄

Profilbild von shownotover
shownotovervor 4 Tagen

Well, gotta save $ where we can 🙂

Profilbild von André Felipe
André Felipevor 4 Tagen

Kkkkkkkkkkk você tá de sacanagem com a minha cara certo? É uma piada??? Com o plano plus apenas vc usa o astra no Max, envia um promot literalmente dizendo "crie um vídeo, faça seu melhor" recebe um resultado ruim e vem bostejando aqui? Primeiro aprenda usar ferramentas de ia

Profilbild von shownotover
shownotovervor 4 Tagen

This is what Claude did with exactly the same prompt. That is why I tested Astra like this.

Profilbild von Thanh Nam
Thanh Namvor 4 Tagen

Use Astra to create then Opus fix the "taste"

Profilbild von shownotover
shownotovervor 4 Tagen

That'll simply mean I am using Astra as implementor and Opus as Orchastrator

Profilbild von Peter
Petervor 4 Tagen

Good stats, Claude indeed seems way better!!!

Profilbild von shownotover
shownotovervor 4 Tagen

yeah bro! I had switched to claude

Profilbild von Nazran F.
Nazran F.vor 4 Tagen

Opus 5.5 is 50x better than this for the same use case. I don’t know how OpenAI randomly dropped the bag

Profilbild von shownotover
shownotovervor 4 Tagen

I even created a video on Opus 5.5 and it was so much better. OpenAI used to be a lot better.

Profilbild von Nazran F.
Nazran F.vor 4 Tagen

It’ll flip again when the new models come out I’m sure. It’s just one beating the other in a loop.

Profilbild von Wytze H.
Wytze H.vor 4 Tagen

Ok, I’m usually not the guy to comment on this, but this pissed me off a little. Most of the time a model is as good as the prompt you give it. Of course sometimes a model comes around that does better with less, however, more often then not the more specific you explain what you want the better the result will be. I did the same with GPT-6 Sol on High effort.

Profilbild von shownotover
shownotovervor 4 Tagen

Same prompt and same $20 sub btw Now choice is yours if you get pissed or open your eyes

Profilbild von Adriel S.
Adriel S.vor 4 Tagen

perahps high or extra high could've been more than enough; htese models are pretty good and avoid it having it "overthink" itslef in a loop at times

Profilbild von shownotover
shownotovervor 4 Tagen

Hmm... When I use Opus 5.5 on max, that doesn't hurt... but when I use Astra on Max, that starts hurting?? why is that? I want maximum capability by the Model that I can get out for one task that I want to be the best. am I missing something?

Profilbild von Adam
Adamvor 4 Tagen

these really work well

Profilbild von shownotover
shownotovervor 4 Tagen

🥲 I guess if it works for you, I support you bro

Profilbild von Adam
Adamvor 4 Tagen

Your posting style 🤣 not codex

Profilbild von shownotover
shownotovervor 4 Tagen

😁 From Social media, I have learnt that the more you have emotions in your content, the more people will relate to it. ofc yapping doesn't count so have to do it like this ☺

Profilbild von Adam
Adamvor 4 Tagen

Yeah thats why I shout all the time 🤣. The more emotions you can stir the better

Profilbild von Mr Leefy
Mr Leefyvor 4 Tagen

"First time?"

Profilbild von Javier Jiménez
Javier Jiménezvor 4 Tagen

Tremendo lameculo de Darío y famboy de anthropic

Profilbild von shownotover
shownotovervor 4 Tagen

you had some serious brainwash to not see the clear difference between the quality to cost ratio can't argue with you

Profilbild von Javier Jiménez
Javier Jiménezvor 4 Tagen

A ver todo sabe que anthropic le importa sus cliente empresarial y nos a su usuario normales y lo que hace con 5.5 opus es para su IPO y cuando ya eso pase será la misma mierda que antes y openAI tiene modelo más potente intenamete ok

Profilbild von Conza Walker
Conza Walkervor 4 Tagen

So let’s see if Tuesday addresses this, I like the idea that the 6th dot is for multi agent mode, and ties into the M&Ms they’re going for

Profilbild von Expansion
Expansionvor 4 Tagen

no one cares

Profilbild von Crio Songo
Crio Songovor 4 Tagen

This is really unacceptable. Codex's quota is too tight compared to Claude, I'd better not buy it for now.

Ähnliche Videos

I tested Claude Code on a fresh account - 1,500 lines of HTML cost me 50% of my window. Full video and summary is here.. I just ran a recorded test on Claude Code with a fresh account (Pro, not Max - my main account was 20x Max) , and the result is honestly insane. The task was trivial: create 3 simple demo HTML pages, around 500 lines each. Roughly 1,500 lines of code total. Nothing massive. Nothing enterprise-grade. Nothing that should meaningfully stress a premium coding product. And yet Claude Code burned through 40% of my 5-hour window almost immediately. I ran the exact same test with Codex, and it consumed only 2%. Then it got even worse: after the session ended, I did absolutely nothing for 15 minutes, and Claude still ate another 10%. Total: 50% of the 5-hour window gone for a tiny HTML demo. My weekly usage had already started at 2% before I even really used it, and after this tiny test it jumped to 8%. Now let us be generous and assume this entire run used around 30k tokens total. If 30k tokens represents 10% of weekly usage, that implies around 300k tokens per week. That is roughly 1.2M-1.3M tokens per month, and even if you round up aggressively, you are still in the 1.5M token range. Using the Sonnet 4.6 pricing you list: $3 per 1M input tokens $15 per 1M output tokens How exactly is this supposed to make sense for a paid coding product? Because from the user side, this no longer looks like "premium usage protection." It looks like a quota system that is either wildly inefficient, badly broken, or being accounted in a way users are not being told about. And that is before I even get to my main account: my $200 Max plan now dies in a single day. Just a few months ago, similar or heavier usage would last me about a week. So no, I do not buy the "maybe you just used it more" excuse anymore. Something is clearly broken in Claude Code. Either token accounting is broken, context handling is broken, background consumption is broken, or all three. Alex Albert is this really the experience you want users to pay for? Just watch the video. I tried to be very transparent and clear for your team! I was fan of Claude but just disappointed! And if you want, send me the detailed token accounting for this session and let us inspect it together publicly. Because from where I am standing, this is no longer a small pricing annoyance. It looks like something seriously wrong is happening, and users deserve a real explanation.

Hayrettin Tüzel

26,854 Aufrufe • vor 6 Monaten

I just compared Claude Code vs Codex vs Cursor CLI The task was to build a Next.js app with Tailwind 4 and shadcn components to collect customer feedback and showcase it with a widget. I gave all three the same prompt and let them go for 30 minutes to see what they came up with. Claude Code with Opus 4.1 Even though I told it to set up the app in the existing project folder, it tried to create a directory for it. After I interrupted and told it not to do that, it built a demo form and landing page with no errors. I had to ask it to make the demo interactive so users could submit a testimonial and preview it. The landing page looked like AI and was pretty basic, but it worked and it was done in a fraction of the time of the others. Total tokens used: 33k Codex with GPT-5 At the end of the 30 minutes I just could not get Codex to produce a working app. It got stuck in a loop of not being able to set up Tailwind 4 and despite many, MANY, attempts, I ended up with a "failed to compile" error. Total tokens used: 102k Cursor Agent with GPT-5 This was the slowest agent by far and a couple of times I actually thought it got stuck in a loop and was close to Ctrl+C'ing to cancel it. The TUI is really nice though, especially how it shows diffs and it did eventually build a working app (after one or two slight errors that needed fixing) The demo was interactive and it had a very minimal design that looked bare but also a lot less like an "AI generated" app than the Opus 4.1 design. It also wasn't too chatty and just did what it needed to do! Code quality was on a par with Opus 4.1, but it did use 5.5x as many tokens to get there. Still cheaper than Opus on a direct comparison but not when you factor in a Claude Code Max subscription. Total tokens: 188k I'll be able to do a proper comparison and record some videos when I'm back from holiday but for now, Opus is still the more capable model out of the box and Claude Code is the more complete CLI product. It will be interesting to see how Cursor evolve their CLI though with commands and subagents because I think with GPT-5 they have a real shot at providing competition for Claude Code if they can optimise output to get similar quality with less tokens. Jump to 0:40 in the video to see the two apps. Which do you think is which? ;)

Ian Nuttall

195,173 Aufrufe • vor 1 Jahr