Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚨 codex users!! I tested GPT-6.1 Sol at max and.... it is BAD I gave it a prompt to create a motion graphics video it used: - 2% weekly limit (15% 5 hr limit) - 1.7 Million tokens - 58 minutes of work - $1.8 API usage So... the...

29,342 görüntüleme • 1 gün önce •via X (Twitter)

90 Yorum

shownotover profil fotoğrafı
shownotover1 gün önce

If you wanna grow on X

Harshith profil fotoğrafı
Harshith1 gün önce

just $360 we get nearly $1k with Claude pro

shownotover profil fotoğrafı
shownotover1 gün önce

Yeah I know that... I made 4 posts about it...

John Rood profil fotoğrafı
John Rood1 gün önce

the trap is a task with no verifier: nothing can fail, so nothing ends the loop and it just keeps polishing. the runs in my fleet that burn budget always look like this.

shownotover profil fotoğrafı
shownotover1 gün önce

Bro, it has to verify itself. It was literally written inside the instructions to watch the video before shipping it. This is the final output after an hour of work. Opus did so much better than this.

: profil fotoğrafı
:1 gün önce

@johnroodepic where is the opus video for comparison? i smell bullshit

shownotover profil fotoğrafı
shownotover1 gün önce

@johnroodepic Is this bullshit by Opus 5.5??

Anupam Haldkar  profil fotoğrafı
Anupam Haldkar 1 gün önce

Sol is cheap compute for agents, not a director, $1.8 and 58 min for that clip is the actual product review.

Ravi Kushwaha profil fotoğrafı
Ravi Kushwaha1 gün önce

hahahahhaa

shownotover profil fotoğrafı
shownotover1 gün önce

🫠

Sahil Panhotra | Indie Builder profil fotoğrafı
Sahil Panhotra | Indie Builder1 gün önce

opus is better

shownotover profil fotoğrafı
shownotover1 gün önce

It is not just better, it's so much better 🥲

Yipeng | Daily AI profil fotoğrafı
Yipeng | Daily AI1 gün önce

This is a useful distinction: low quota burn doesn’t help much if the output needs a full redo. I’d love to see the same prompt and workflow run through Sol, Astra, and Opus side by side—quality per usable result is the number that matters.

shownotover profil fotoğrafı
shownotover1 gün önce

All of them are on my profile... You can have a look and then decide

Yipeng | Daily AI profil fotoğrafı
Yipeng | Daily AI1 gün önce

Thanks, I’ll take a look. The side-by-side outputs are exactly what I was hoping to see. Were Sol, Astra, and Opus given the same prompt and workflow?

shownotover profil fotoğrafı
shownotover1 gün önce

Yes. Exactly same prompts, same env and everything was done under keen observation by yours truely 🙂

Yipeng | Daily AI profil fotoğrafı
Yipeng | Daily AI1 gün önce

That makes the comparison much more useful. Thanks for controlling the prompt and environment 🙂 Did you measure how much human cleanup each final video needed? That would help put the cost and quality differences in context.

shownotover profil fotoğrafı
shownotover1 gün önce

Well I didn't control anything at all. I just took the output and posted on the post.

Prompt Heat profil fotoğrafı
Prompt Heat1 gün önce

claude 20 dolar > gpt 20 dolar acc

shownotover profil fotoğrafı
shownotover1 gün önce

100% agree with that bro.

Aden profil fotoğrafı
Aden1 gün önce

That ribbon looks great for 58 minutes and $1.8

shownotover profil fotoğrafı
shownotover1 gün önce

Nah... Check Claude $20 ... This usage is crap compared to Claude $20

Jack Dev profil fotoğrafı
Jack Dev1 gün önce

Try at xhigh, I've been running planner xhigh and reviewer xhigh with the implementation as medium/low

shownotover profil fotoğrafı
shownotover1 gün önce

I know the usage limit is very, very good. The only problem I have right now is the quality that you are getting for the price you are paying on Codex and on Claude.

Jack Dev profil fotoğrafı
Jack Dev1 gün önce

Yeah, OpenAI seems to be focusing more on cost now. which is great for users who doesn't have a lot of money like me, switching between the two works too but every week another model beat another model, back and forth so i stay with openai

Miro Jomaa profil fotoğrafı
Miro Jomaa1 gün önce

I prefer this over opus. Opus felt like adhd for the sake of something moving

shownotover profil fotoğrafı
shownotover1 gün önce

I mean, it's your choice but check the results opus got me... They are crazy

Miro Jomaa profil fotoğrafı
Miro Jomaa1 gün önce

yeah opus one is better, just a bit less adhd. I also noticed codex basicly does the same design. small text over the headline text and the end of the headline text is colored.

Silicon Rumors profil fotoğrafı
Silicon Rumors1 gün önce

I don't think this accurately shows performance of the model, it only shows that performance in making these videos.

shownotover profil fotoğrafı
shownotover1 gün önce

What will show accurate performance then? I'll test that out

Silicon Rumors profil fotoğrafı
Silicon Rumors1 gün önce

I think complex coding and tasks will show more accurate performance because i believe its more tuned for that anyways rather than making graphic videos like these

shownotover profil fotoğrafı
shownotover1 gün önce

Okay explain to me one: what should I test and what exactly should the output feel like? I have to post these results and if I don't post the results in quality (and/or if it's difficult for a normal person to see the difference), then I believe I cannot conduct the tests because that wouldn't actually make sense to anybody. I want you to come up with a test that you think would best fit every model that I test, in the same environment with the same prompt. I'll test these out for you and then paste the results right here along with the usages and everything.

Thanh Nam profil fotoğrafı
Thanh Nam1 gün önce

Opus

shownotover profil fotoğrafı
shownotover1 gün önce

Good choice

Anima profil fotoğrafı
Anima1 gün önce

Astra 6.1 got delayed unfortunately So no good motion design OAI users

shownotover profil fotoğrafı
shownotover1 gün önce

It's delayed until late October so... We are gonna have plenty of time to use Claude

Darhan profil fotoğrafı
Darhan1 gün önce

The shit test

shownotover profil fotoğrafı
shownotover1 gün önce

What do you mean 😭

CauseWhyNot profil fotoğrafı
CauseWhyNot1 gün önce

Better than anything I can do but pretty sad compared to opus 🥲

shownotover profil fotoğrafı
shownotover1 gün önce

Vibe coders can't do this 😔

CauseWhyNot profil fotoğrafı
CauseWhyNot1 gün önce

😭 ouch 😂

Victor Pimentel profil fotoğrafı
Victor Pimentel1 gün önce

Horrível.

shownotover profil fotoğrafı
shownotover1 gün önce

They really gotta improve ngl

Eric profil fotoğrafı
Eric1 gün önce

“I gave ai 1 task out of millions of tasks it can do and based on this 1 task you shouldn’t buy this “

shownotover profil fotoğrafı
shownotover1 gün önce

Well it's just not one task. This is the second task it's failing on... But yeah you can stay on Codex. Don't worry about that.

Gabriele Bolognese | SaaS founder profil fotoğrafı
Gabriele Bolognese | SaaS founder1 gün önce

claude is ALWAYS one step above

shownotover profil fotoğrafı
shownotover1 gün önce

💯 agreed

JP profil fotoğrafı
JP1 gün önce

Can’t compare it with Opus 5.5 or sonnet 5:5

Arnold profil fotoğrafı
Arnold1 gün önce

Thanks for doing thi Tests for us and being honest 🤝

shownotover profil fotoğrafı
shownotover1 gün önce

No problem... More tests coming soon

HASSAN profil fotoğrafı
HASSAN1 gün önce

Scamgpt 6.1 sol is not eve close to sonnet 5.

shownotover profil fotoğrafı
shownotover1 gün önce

I know and this is very sad 😭

Josh profil fotoğrafı
Josh1 gün önce

We would not call it bad if it wasn’t for opus.

shownotover profil fotoğrafı
shownotover1 gün önce

This is their latest 2nd gen model 🥲

AnimeVibeX profil fotoğrafı
AnimeVibeX1 gün önce

Which Plan?

shownotover profil fotoğrafı
shownotover1 gün önce

$20 codex plan

AnimeVibeX profil fotoğrafı
AnimeVibeX1 gün önce

I bought Claude );

shownotover profil fotoğrafı
shownotover1 gün önce

I have both and I test them every day or two.

AnimeVibeX profil fotoğrafı
AnimeVibeX1 gün önce

Opus 5.5 still Beats every ai Model

Shahzeb Naveed profil fotoğrafı
Shahzeb Naveed1 gün önce

$20 codex still ain't worth it 😢

shownotover profil fotoğrafı
shownotover1 gün önce

Better buy a Claude $20 account

hisoka profil fotoğrafı
hisoka1 gün önce

It’s amazing. So good and so efficient.

shownotover profil fotoğrafı
shownotover1 gün önce

And painfully slow

hisoka profil fotoğrafı
hisoka1 gün önce

Not for me

avhq profil fotoğrafı
avhq1 gün önce

Space bunny...

shownotover profil fotoğrafı
shownotover1 gün önce

I haven't tested it but it's free and very good... Free users should definitely use that 🙂

avhq profil fotoğrafı
avhq1 gün önce

No... It's better at ui and definitely motion that chatgpt. There's so many stunning posts on x

Peter profil fotoğrafı
Peter1 gün önce

Happy to see your initial feelings changed after you tested; seemingly, the big SELLING point is the cost-effectiveness. I am super happy that we finally have many options to choose from (and all are OK or awesome quality)

shownotover profil fotoğrafı
shownotover1 gün önce

Yeah ... Now it's about if you want more quality for more cost or wanna be cost effective 🥲

Peter profil fotoğrafı
Peter1 gün önce

yap, but AT least you have choices!

Linus profil fotoğrafı
Linus1 gün önce

A final visual check should be non-negotiable here.

shownotover profil fotoğrafı
shownotover1 gün önce

That's on you bro 🫵

cxan profil fotoğrafı
cxan1 gün önce

You remember long the other models took to make the video? Have been using sol 6.1 this morning and indeed it felt very slow. Usage seems decent but slowing down tokens is also a way to effectively reduce it....

shownotover profil fotoğrafı
shownotover1 gün önce

You know what? I feel like we are "feeling" that the usage is great because they have slowed down the models. If the model was fast enough, the usage would have drained faster. To be honest, this is what I feel.

cxan profil fotoğrafı
cxan1 gün önce

yes honestly I really really hope this isn't the case but I also thought about that. Honestly, the only thing thats keeping me from cancelling my $100 plan right now is my hopes in 6.1 sol being a decent daily-driver with good usage. If that now turns out false I'm back to claude

dhru profil fotoğrafı
dhru1 gün önce

i bought it few days before claude launched opus 5.5

shownotover profil fotoğrafı
shownotover1 gün önce

That hurts bro. 😭

dhru profil fotoğrafı
dhru1 gün önce

i know, will switch next month now 😭

Kekius Pisimus profil fotoğrafı
Kekius Pisimus1 gün önce

Why the hell bro like GPT model so much???

shownotover profil fotoğrafı
shownotover1 gün önce

Where did I say I like it? I'm just giving you guys information about whether or not you should buy a Codex $20 account.

Kekius Pisimus profil fotoğrafı
Kekius Pisimus1 gün önce

You said you will delete codex and now you're using it 😂

shownotover profil fotoğrafı
shownotover1 gün önce

I have a $20 sub and I don't wanna let my $20 go to waste just because I dont like it

Kekius Pisimus profil fotoğrafı
Kekius Pisimus1 gün önce

I think I said to you a long time ago... and you're still using codex. 😂

shownotover profil fotoğrafı
shownotover1 gün önce

it may have been a week ago bro.

Nabendu Biswas profil fotoğrafı
Nabendu Biswas1 gün önce

I am testing now and gave the same prompt as I gave to Sol 6, Astra, Opus 5.5 and Sonnet 5.5. It is to create a Subway surfer game without Blender. Let's see what it does and will post tomorrow

shownotover profil fotoğrafı
shownotover1 gün önce

Cool... Lemme know about that

James Malsawm profil fotoğrafı
James Malsawm1 gün önce

If Opus 5.5 don't come out, Sol 6.1 definitely is awesome model.

shownotover profil fotoğrafı
shownotover1 gün önce

What? Opus 5.5 is already out

Cuong Quach profil fotoğrafı
Cuong Quach1 gün önce

Why not build a page to track those numbers? It’d make them much easier to compare.

shownotover profil fotoğrafı
shownotover1 gün önce

I'd love to... I should definitely build one 🫠

Benzer Videolar

I tested Claude Code on a fresh account - 1,500 lines of HTML cost me 50% of my window. Full video and summary is here.. I just ran a recorded test on Claude Code with a fresh account (Pro, not Max - my main account was 20x Max) , and the result is honestly insane. The task was trivial: create 3 simple demo HTML pages, around 500 lines each. Roughly 1,500 lines of code total. Nothing massive. Nothing enterprise-grade. Nothing that should meaningfully stress a premium coding product. And yet Claude Code burned through 40% of my 5-hour window almost immediately. I ran the exact same test with Codex, and it consumed only 2%. Then it got even worse: after the session ended, I did absolutely nothing for 15 minutes, and Claude still ate another 10%. Total: 50% of the 5-hour window gone for a tiny HTML demo. My weekly usage had already started at 2% before I even really used it, and after this tiny test it jumped to 8%. Now let us be generous and assume this entire run used around 30k tokens total. If 30k tokens represents 10% of weekly usage, that implies around 300k tokens per week. That is roughly 1.2M-1.3M tokens per month, and even if you round up aggressively, you are still in the 1.5M token range. Using the Sonnet 4.6 pricing you list: $3 per 1M input tokens $15 per 1M output tokens How exactly is this supposed to make sense for a paid coding product? Because from the user side, this no longer looks like "premium usage protection." It looks like a quota system that is either wildly inefficient, badly broken, or being accounted in a way users are not being told about. And that is before I even get to my main account: my $200 Max plan now dies in a single day. Just a few months ago, similar or heavier usage would last me about a week. So no, I do not buy the "maybe you just used it more" excuse anymore. Something is clearly broken in Claude Code. Either token accounting is broken, context handling is broken, background consumption is broken, or all three. Alex Albert is this really the experience you want users to pay for? Just watch the video. I tried to be very transparent and clear for your team! I was fan of Claude but just disappointed! And if you want, send me the detailed token accounting for this session and let us inspect it together publicly. Because from where I am standing, this is no longer a small pricing annoyance. It looks like something seriously wrong is happening, and users deserve a real explanation.

Hayrettin Tüzel

26,854 görüntüleme • 6 ay önce