Загрузка видео...

Не удалось загрузить видео

На главную

🚨 codex users!! I tested GPT-6.1 Sol at max and.... it is BAD I gave it a prompt to create a motion graphics video it used: - 2% weekly limit (15% 5 hr limit) - 1.7 Million tokens - 58 minutes of work - $1.8 API usage So... the...

29,342 просмотров • 1 день назад •via X (Twitter)

Комментарии: 90

Фото профиля shownotover
shownotover1 день назад

If you wanna grow on X

Фото профиля Harshith
Harshith1 день назад

just $360 we get nearly $1k with Claude pro

Фото профиля shownotover
shownotover1 день назад

Yeah I know that... I made 4 posts about it...

Фото профиля John Rood
John Rood1 день назад

the trap is a task with no verifier: nothing can fail, so nothing ends the loop and it just keeps polishing. the runs in my fleet that burn budget always look like this.

Фото профиля shownotover
shownotover1 день назад

Bro, it has to verify itself. It was literally written inside the instructions to watch the video before shipping it. This is the final output after an hour of work. Opus did so much better than this.

Фото профиля :
:1 день назад

@johnroodepic where is the opus video for comparison? i smell bullshit

Фото профиля shownotover
shownotover1 день назад

@johnroodepic Is this bullshit by Opus 5.5??

Фото профиля Anupam Haldkar 
Anupam Haldkar 1 день назад

Sol is cheap compute for agents, not a director, $1.8 and 58 min for that clip is the actual product review.

Фото профиля Ravi Kushwaha
Ravi Kushwaha1 день назад

hahahahhaa

Фото профиля shownotover
shownotover1 день назад

🫠

Фото профиля Sahil Panhotra | Indie Builder
Sahil Panhotra | Indie Builder1 день назад

opus is better

Фото профиля shownotover
shownotover1 день назад

It is not just better, it's so much better 🥲

Фото профиля Yipeng | Daily AI
Yipeng | Daily AI1 день назад

This is a useful distinction: low quota burn doesn’t help much if the output needs a full redo. I’d love to see the same prompt and workflow run through Sol, Astra, and Opus side by side—quality per usable result is the number that matters.

Фото профиля shownotover
shownotover1 день назад

All of them are on my profile... You can have a look and then decide

Фото профиля Yipeng | Daily AI
Yipeng | Daily AI1 день назад

Thanks, I’ll take a look. The side-by-side outputs are exactly what I was hoping to see. Were Sol, Astra, and Opus given the same prompt and workflow?

Фото профиля shownotover
shownotover1 день назад

Yes. Exactly same prompts, same env and everything was done under keen observation by yours truely 🙂

Фото профиля Yipeng | Daily AI
Yipeng | Daily AI1 день назад

That makes the comparison much more useful. Thanks for controlling the prompt and environment 🙂 Did you measure how much human cleanup each final video needed? That would help put the cost and quality differences in context.

Фото профиля shownotover
shownotover1 день назад

Well I didn't control anything at all. I just took the output and posted on the post.

Фото профиля Prompt Heat
Prompt Heat1 день назад

claude 20 dolar > gpt 20 dolar acc

Фото профиля shownotover
shownotover1 день назад

100% agree with that bro.

Фото профиля Aden
Aden1 день назад

That ribbon looks great for 58 minutes and $1.8

Фото профиля shownotover
shownotover1 день назад

Nah... Check Claude $20 ... This usage is crap compared to Claude $20

Фото профиля Jack Dev
Jack Dev1 день назад

Try at xhigh, I've been running planner xhigh and reviewer xhigh with the implementation as medium/low

Фото профиля shownotover
shownotover1 день назад

I know the usage limit is very, very good. The only problem I have right now is the quality that you are getting for the price you are paying on Codex and on Claude.

Фото профиля Jack Dev
Jack Dev1 день назад

Yeah, OpenAI seems to be focusing more on cost now. which is great for users who doesn't have a lot of money like me, switching between the two works too but every week another model beat another model, back and forth so i stay with openai

Фото профиля Miro Jomaa
Miro Jomaa1 день назад

I prefer this over opus. Opus felt like adhd for the sake of something moving

Фото профиля shownotover
shownotover1 день назад

I mean, it's your choice but check the results opus got me... They are crazy

Фото профиля Miro Jomaa
Miro Jomaa1 день назад

yeah opus one is better, just a bit less adhd. I also noticed codex basicly does the same design. small text over the headline text and the end of the headline text is colored.

Фото профиля Silicon Rumors
Silicon Rumors1 день назад

I don't think this accurately shows performance of the model, it only shows that performance in making these videos.

Фото профиля shownotover
shownotover1 день назад

What will show accurate performance then? I'll test that out

Фото профиля Silicon Rumors
Silicon Rumors1 день назад

I think complex coding and tasks will show more accurate performance because i believe its more tuned for that anyways rather than making graphic videos like these

Фото профиля shownotover
shownotover1 день назад

Okay explain to me one: what should I test and what exactly should the output feel like? I have to post these results and if I don't post the results in quality (and/or if it's difficult for a normal person to see the difference), then I believe I cannot conduct the tests because that wouldn't actually make sense to anybody. I want you to come up with a test that you think would best fit every model that I test, in the same environment with the same prompt. I'll test these out for you and then paste the results right here along with the usages and everything.

Фото профиля Thanh Nam
Thanh Nam1 день назад

Opus

Фото профиля shownotover
shownotover1 день назад

Good choice

Фото профиля Anima
Anima1 день назад

Astra 6.1 got delayed unfortunately So no good motion design OAI users

Фото профиля shownotover
shownotover1 день назад

It's delayed until late October so... We are gonna have plenty of time to use Claude

Фото профиля Darhan
Darhan1 день назад

The shit test

Фото профиля shownotover
shownotover1 день назад

What do you mean 😭

Фото профиля CauseWhyNot
CauseWhyNot1 день назад

Better than anything I can do but pretty sad compared to opus 🥲

Фото профиля shownotover
shownotover1 день назад

Vibe coders can't do this 😔

Фото профиля CauseWhyNot
CauseWhyNot1 день назад

😭 ouch 😂

Фото профиля Victor Pimentel
Victor Pimentel1 день назад

Horrível.

Фото профиля shownotover
shownotover1 день назад

They really gotta improve ngl

Фото профиля Eric
Eric1 день назад

“I gave ai 1 task out of millions of tasks it can do and based on this 1 task you shouldn’t buy this “

Фото профиля shownotover
shownotover1 день назад

Well it's just not one task. This is the second task it's failing on... But yeah you can stay on Codex. Don't worry about that.

Фото профиля Gabriele Bolognese | SaaS founder
Gabriele Bolognese | SaaS founder1 день назад

claude is ALWAYS one step above

Фото профиля shownotover
shownotover1 день назад

💯 agreed

Фото профиля JP
JP1 день назад

Can’t compare it with Opus 5.5 or sonnet 5:5

Фото профиля Arnold
Arnold1 день назад

Thanks for doing thi Tests for us and being honest 🤝

Фото профиля shownotover
shownotover1 день назад

No problem... More tests coming soon

Фото профиля HASSAN
HASSAN1 день назад

Scamgpt 6.1 sol is not eve close to sonnet 5.

Фото профиля shownotover
shownotover1 день назад

I know and this is very sad 😭

Фото профиля Josh
Josh1 день назад

We would not call it bad if it wasn’t for opus.

Фото профиля shownotover
shownotover1 день назад

This is their latest 2nd gen model 🥲

Фото профиля AnimeVibeX
AnimeVibeX1 день назад

Which Plan?

Фото профиля shownotover
shownotover1 день назад

$20 codex plan

Фото профиля AnimeVibeX
AnimeVibeX1 день назад

I bought Claude );

Фото профиля shownotover
shownotover1 день назад

I have both and I test them every day or two.

Фото профиля AnimeVibeX
AnimeVibeX1 день назад

Opus 5.5 still Beats every ai Model

Фото профиля Shahzeb Naveed
Shahzeb Naveed1 день назад

$20 codex still ain't worth it 😢

Фото профиля shownotover
shownotover1 день назад

Better buy a Claude $20 account

Фото профиля hisoka
hisoka1 день назад

It’s amazing. So good and so efficient.

Фото профиля shownotover
shownotover1 день назад

And painfully slow

Фото профиля hisoka
hisoka1 день назад

Not for me

Фото профиля avhq
avhq1 день назад

Space bunny...

Фото профиля shownotover
shownotover1 день назад

I haven't tested it but it's free and very good... Free users should definitely use that 🙂

Фото профиля avhq
avhq1 день назад

No... It's better at ui and definitely motion that chatgpt. There's so many stunning posts on x

Фото профиля Peter
Peter1 день назад

Happy to see your initial feelings changed after you tested; seemingly, the big SELLING point is the cost-effectiveness. I am super happy that we finally have many options to choose from (and all are OK or awesome quality)

Фото профиля shownotover
shownotover1 день назад

Yeah ... Now it's about if you want more quality for more cost or wanna be cost effective 🥲

Фото профиля Peter
Peter1 день назад

yap, but AT least you have choices!

Фото профиля Linus
Linus1 день назад

A final visual check should be non-negotiable here.

Фото профиля shownotover
shownotover1 день назад

That's on you bro 🫵

Фото профиля cxan
cxan1 день назад

You remember long the other models took to make the video? Have been using sol 6.1 this morning and indeed it felt very slow. Usage seems decent but slowing down tokens is also a way to effectively reduce it....

Фото профиля shownotover
shownotover1 день назад

You know what? I feel like we are "feeling" that the usage is great because they have slowed down the models. If the model was fast enough, the usage would have drained faster. To be honest, this is what I feel.

Фото профиля cxan
cxan1 день назад

yes honestly I really really hope this isn't the case but I also thought about that. Honestly, the only thing thats keeping me from cancelling my $100 plan right now is my hopes in 6.1 sol being a decent daily-driver with good usage. If that now turns out false I'm back to claude

Фото профиля dhru
dhru1 день назад

i bought it few days before claude launched opus 5.5

Фото профиля shownotover
shownotover1 день назад

That hurts bro. 😭

Фото профиля dhru
dhru1 день назад

i know, will switch next month now 😭

Фото профиля Kekius Pisimus
Kekius Pisimus1 день назад

Why the hell bro like GPT model so much???

Фото профиля shownotover
shownotover1 день назад

Where did I say I like it? I'm just giving you guys information about whether or not you should buy a Codex $20 account.

Фото профиля Kekius Pisimus
Kekius Pisimus1 день назад

You said you will delete codex and now you're using it 😂

Фото профиля shownotover
shownotover1 день назад

I have a $20 sub and I don't wanna let my $20 go to waste just because I dont like it

Фото профиля Kekius Pisimus
Kekius Pisimus1 день назад

I think I said to you a long time ago... and you're still using codex. 😂

Фото профиля shownotover
shownotover1 день назад

it may have been a week ago bro.

Фото профиля Nabendu Biswas
Nabendu Biswas1 день назад

I am testing now and gave the same prompt as I gave to Sol 6, Astra, Opus 5.5 and Sonnet 5.5. It is to create a Subway surfer game without Blender. Let's see what it does and will post tomorrow

Фото профиля shownotover
shownotover1 день назад

Cool... Lemme know about that

Фото профиля James Malsawm
James Malsawm1 день назад

If Opus 5.5 don't come out, Sol 6.1 definitely is awesome model.

Фото профиля shownotover
shownotover1 день назад

What? Opus 5.5 is already out

Фото профиля Cuong Quach
Cuong Quach1 день назад

Why not build a page to track those numbers? It’d make them much easier to compare.

Фото профиля shownotover
shownotover1 день назад

I'd love to... I should definitely build one 🫠

Похожие видео

I tested Claude Code on a fresh account - 1,500 lines of HTML cost me 50% of my window. Full video and summary is here.. I just ran a recorded test on Claude Code with a fresh account (Pro, not Max - my main account was 20x Max) , and the result is honestly insane. The task was trivial: create 3 simple demo HTML pages, around 500 lines each. Roughly 1,500 lines of code total. Nothing massive. Nothing enterprise-grade. Nothing that should meaningfully stress a premium coding product. And yet Claude Code burned through 40% of my 5-hour window almost immediately. I ran the exact same test with Codex, and it consumed only 2%. Then it got even worse: after the session ended, I did absolutely nothing for 15 minutes, and Claude still ate another 10%. Total: 50% of the 5-hour window gone for a tiny HTML demo. My weekly usage had already started at 2% before I even really used it, and after this tiny test it jumped to 8%. Now let us be generous and assume this entire run used around 30k tokens total. If 30k tokens represents 10% of weekly usage, that implies around 300k tokens per week. That is roughly 1.2M-1.3M tokens per month, and even if you round up aggressively, you are still in the 1.5M token range. Using the Sonnet 4.6 pricing you list: $3 per 1M input tokens $15 per 1M output tokens How exactly is this supposed to make sense for a paid coding product? Because from the user side, this no longer looks like "premium usage protection." It looks like a quota system that is either wildly inefficient, badly broken, or being accounted in a way users are not being told about. And that is before I even get to my main account: my $200 Max plan now dies in a single day. Just a few months ago, similar or heavier usage would last me about a week. So no, I do not buy the "maybe you just used it more" excuse anymore. Something is clearly broken in Claude Code. Either token accounting is broken, context handling is broken, background consumption is broken, or all three. Alex Albert is this really the experience you want users to pay for? Just watch the video. I tried to be very transparent and clear for your team! I was fan of Claude but just disappointed! And if you want, send me the detailed token accounting for this session and let us inspect it together publicly. Because from where I am standing, this is no longer a small pricing annoyance. It looks like something seriously wrong is happening, and users deserve a real explanation.

Hayrettin Tüzel

26,854 просмотров • 6 месяцев назад