Video wird geladen...
Video konnte nicht geladen werden
🚨 codex users!! I tested GPT-6.1 Sol at max and.... it is BAD I gave it a prompt to create a motion graphics video it used: - 2% weekly limit (15% 5 hr limit) - 1.7 Million tokens - 58 minutes of work - $1.8 API usage So... the... show more
28,875 Aufrufe • vor 1 Tag •via X (Twitter)
90 Kommentare

If you wanna grow on X

just $360 we get nearly $1k with Claude pro

Yeah I know that... I made 4 posts about it...

the trap is a task with no verifier: nothing can fail, so nothing ends the loop and it just keeps polishing. the runs in my fleet that burn budget always look like this.

Bro, it has to verify itself. It was literally written inside the instructions to watch the video before shipping it. This is the final output after an hour of work. Opus did so much better than this.

@johnroodepic where is the opus video for comparison? i smell bullshit

@johnroodepic Is this bullshit by Opus 5.5??

Sol is cheap compute for agents, not a director, $1.8 and 58 min for that clip is the actual product review.

hahahahhaa

🫠

opus is better

It is not just better, it's so much better 🥲

This is a useful distinction: low quota burn doesn’t help much if the output needs a full redo. I’d love to see the same prompt and workflow run through Sol, Astra, and Opus side by side—quality per usable result is the number that matters.

All of them are on my profile... You can have a look and then decide

Thanks, I’ll take a look. The side-by-side outputs are exactly what I was hoping to see. Were Sol, Astra, and Opus given the same prompt and workflow?

Yes. Exactly same prompts, same env and everything was done under keen observation by yours truely 🙂

That makes the comparison much more useful. Thanks for controlling the prompt and environment 🙂 Did you measure how much human cleanup each final video needed? That would help put the cost and quality differences in context.

Well I didn't control anything at all. I just took the output and posted on the post.

claude 20 dolar > gpt 20 dolar acc

100% agree with that bro.

That ribbon looks great for 58 minutes and $1.8

Nah... Check Claude $20 ... This usage is crap compared to Claude $20

Try at xhigh, I've been running planner xhigh and reviewer xhigh with the implementation as medium/low

I know the usage limit is very, very good. The only problem I have right now is the quality that you are getting for the price you are paying on Codex and on Claude.

Yeah, OpenAI seems to be focusing more on cost now. which is great for users who doesn't have a lot of money like me, switching between the two works too but every week another model beat another model, back and forth so i stay with openai

I prefer this over opus. Opus felt like adhd for the sake of something moving

I mean, it's your choice but check the results opus got me... They are crazy

yeah opus one is better, just a bit less adhd. I also noticed codex basicly does the same design. small text over the headline text and the end of the headline text is colored.

I don't think this accurately shows performance of the model, it only shows that performance in making these videos.

What will show accurate performance then? I'll test that out

I think complex coding and tasks will show more accurate performance because i believe its more tuned for that anyways rather than making graphic videos like these

Okay explain to me one: what should I test and what exactly should the output feel like? I have to post these results and if I don't post the results in quality (and/or if it's difficult for a normal person to see the difference), then I believe I cannot conduct the tests because that wouldn't actually make sense to anybody. I want you to come up with a test that you think would best fit every model that I test, in the same environment with the same prompt. I'll test these out for you and then paste the results right here along with the usages and everything.

Opus

Good choice

Astra 6.1 got delayed unfortunately So no good motion design OAI users

It's delayed until late October so... We are gonna have plenty of time to use Claude

The shit test

What do you mean 😭

Better than anything I can do but pretty sad compared to opus 🥲

Vibe coders can't do this 😔

😭 ouch 😂

Horrível.

They really gotta improve ngl

“I gave ai 1 task out of millions of tasks it can do and based on this 1 task you shouldn’t buy this “

Well it's just not one task. This is the second task it's failing on... But yeah you can stay on Codex. Don't worry about that.

claude is ALWAYS one step above

💯 agreed

Can’t compare it with Opus 5.5 or sonnet 5:5

Thanks for doing thi Tests for us and being honest 🤝

No problem... More tests coming soon

Scamgpt 6.1 sol is not eve close to sonnet 5.

I know and this is very sad 😭

We would not call it bad if it wasn’t for opus.

This is their latest 2nd gen model 🥲

Which Plan?

$20 codex plan

I bought Claude );

I have both and I test them every day or two.

Opus 5.5 still Beats every ai Model

$20 codex still ain't worth it 😢

Better buy a Claude $20 account

It’s amazing. So good and so efficient.

And painfully slow

Not for me

Space bunny...

I haven't tested it but it's free and very good... Free users should definitely use that 🙂

No... It's better at ui and definitely motion that chatgpt. There's so many stunning posts on x

Happy to see your initial feelings changed after you tested; seemingly, the big SELLING point is the cost-effectiveness. I am super happy that we finally have many options to choose from (and all are OK or awesome quality)

Yeah ... Now it's about if you want more quality for more cost or wanna be cost effective 🥲

yap, but AT least you have choices!

A final visual check should be non-negotiable here.

That's on you bro 🫵

You remember long the other models took to make the video? Have been using sol 6.1 this morning and indeed it felt very slow. Usage seems decent but slowing down tokens is also a way to effectively reduce it....

You know what? I feel like we are "feeling" that the usage is great because they have slowed down the models. If the model was fast enough, the usage would have drained faster. To be honest, this is what I feel.

yes honestly I really really hope this isn't the case but I also thought about that. Honestly, the only thing thats keeping me from cancelling my $100 plan right now is my hopes in 6.1 sol being a decent daily-driver with good usage. If that now turns out false I'm back to claude

i bought it few days before claude launched opus 5.5

That hurts bro. 😭

i know, will switch next month now 😭

Why the hell bro like GPT model so much???

Where did I say I like it? I'm just giving you guys information about whether or not you should buy a Codex $20 account.

You said you will delete codex and now you're using it 😂

I have a $20 sub and I don't wanna let my $20 go to waste just because I dont like it

I think I said to you a long time ago... and you're still using codex. 😂

it may have been a week ago bro.

I am testing now and gave the same prompt as I gave to Sol 6, Astra, Opus 5.5 and Sonnet 5.5. It is to create a Subway surfer game without Blender. Let's see what it does and will post tomorrow

Cool... Lemme know about that

If Opus 5.5 don't come out, Sol 6.1 definitely is awesome model.

What? Opus 5.5 is already out

Why not build a page to track those numbers? It’d make them much easier to compare.

I'd love to... I should definitely build one 🫠
