Загрузка видео...
Не удалось загрузить видео
Let's keep benchmarking with the Night train this time, with GPT-6 Sol and GPT-6 Luna, so here are all four! 🚀 I gave each model the exact same prompt: “Build a Night train.” •Opus 5.5 • GPT-6 Astra • GPT-6 Sol • GPT-6 Luna Something that really surprised me:... show more
112,633 просмотров • 8 дней назад •via X (Twitter)
Комментарии: 66

Sol copied Opus, Luna copied Astra. 😄

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

@SimonasLTU1 You can't do this with the prompt. Even putting references within the prompt distorts the result. You should use separate sanbox and worktree.

@SimonasLTU1 Not necessarily. I can explicitly tell each model to build from scratch and not use any existing work as a reference. But I take your point: separate worktrees would make it much easier to ensure they can’t see each other’s results. I’ll do that for future benchmarks.

@SimonasLTU1 You can do that, and it's generally effective, but as you can see from your results, you have no guarantee that agents won't peek elsewhere.

@SimonasLTU1 You’re right. An instruction helps, but it can’t guarantee they won’t look at files they can access. Separate environments are the way to make the next comparison properly independent. Thanks for pushing me on this!

Just to check, are you sure the models didn’t peek at each other’s work here?

Yes, I think Sol cheated by using Opus as an example. I need to make sure they don't look at any of the other ones and instead create everything from scratch. I've already adjusted the prompts.😂

@MongoMike ... Luna also peeked

@MongoMike Yeah, from now on, each model starts from scratch. No copycats!😂

There has to be more to this Opus 5.5 and Sol share the same mountain mesh, tree models and positioning, bridge and train models—identical Astra and Luna also very clearly using models and scenery

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

@bnj "Sol may have had access to the Opus version and reproduced it" — that defeats the entire point of your benchmark, doesn't it?

It does undermine the Sol comparison, yes. I won’t run future benchmarks that way. I think the main takeaway still comes through: Opus 5.5 produced the strongest result, and Astra’s original result was stronger than Sol’s and Luna’s. But I wouldn’t call the Sol and Luna runs a fair independent comparison. I’ll make sure every model starts fresh next time.

Sol just walked into the benchmark and stole the spotlight. 😂

Yeah, I think Sol may have had access to the Opus version and reproduced it. 😂 I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

there are only two models here Opus 5.5 and Astra

They’re different if you look closely, but Sol and Luna were influenced by the other models’ results. That was my mistake. I’ve already changed the prompts to make sure each model starts from scratch, so future benchmarks will be right

the camera framing details in Opus 5.5 and GPT Sol are nice, but the concrete pedestals that would block the train's path look illogical 🚞

Good catch! The framing looks great, but those structures don’t make sense if they block the train’s path. That’s exactly the kind of detail a good looking render can hide at first glance. And this was just one prompt, one shot—no follow-up fixes. 🚂👀

Due to a modeling defect that may not be concealed through rendering, it might be required to elaborate on the train tracks within the prompt. 🤝

No shot Luna made that

He did indeed. Yeah, I think Sol may have had access to the Gpt Astra version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

Yeah would love to see this test run again in my experience Luna was never able to do that

I might do it with Gpt Sol and Gpt Luna then!

Somehow, I'm surprised that Luna actually built a train

Yes not bad actually.

Sol found opus's project file and stole it and said it was theres

Yeah, from now on, each model starts from scratch. No copycats! 😂

Love this kind of hands-on comparison. The similar rocket result is a cool surprise, and the cost-per-task angle really hits home. Thanks for sharing!

Your welcome! It is interesting to benchmark all these models!

Same harness? Where’s tests? Numbers?

I gave them the same prompt: create the rocket launch in Three.js and render it as an MP4. This was a visual comparison, not a controlled benchmark with tests or performance numbers, so I don’t have those figures to share. Fair question, though!

Sol and Luna did really good!!

yes!

cheap subagents make the expensive model feel like a reviewer with a smaller bill.

Exactly! I like to call the main agent the orchestrator. 🎯

Luna is genuinely my goat

I think Sol is awesome for the price and what you get from it! 🔥

who is this test for?

Just benchmarking for whoever is interested.

Wtf that astra doing

That's why we're using Opus 5.5 now! 😂

Where's the rocket launch?

I know I mistyped ! 🤣 But they are on my feed

So who's distilling who? 😀

Haha, fair question! 😂 No actual model distillation here, but Sol and Luna could see the earlier work, which influenced their results. That was my setup mistake—I’ll run each model in a separate environment next time.

Fake

Not at all. True results.

Why is no one talking about how much longer it takes opus 5.5 to do the same thing than astra

True. Opus 5.5 take way longer than Astra, like 3 to 4 times longer. I guess is that speed gets less attention because the finished result is easier to show.

There is no way that’s astra. Do we remember how astra looked on release and now it’s putting up those results??

same benchmark that I had done. Just added Gpt Sol and Gpt Luna.

Half the price for comparable output is the real story here, not who 'wins' the benchmark. Are you routing by task type now, or just going with whichever's cheaper?

I’m starting to use this in my own workflow. Main agent on Astra. Sol sub-agents for cheap parallel work. Opus 5.5 for UI and heavy tasks. Sol is ridiculously cheap for what it can do. Opus is what I use when the output has to look right the first time. Knowing when to switch models based on the task, the quality you need, and the cost is becoming an essential skill.

why is GPT Sol so similar to Opus 5.5, and GPT Astra so similar to GPT Luna?

Yeah, I think Sol may have had access to the Opus 5.5 version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

Astra被随机路由到luna了

I selected Astra for that run. What makes you think it was routed to Luna?

因为OpenAI会偷偷降智把Astra路由到其他模型,减少算力消耗

super nice

Thanks

不是造夜班火车吗,怎么还整出火箭了

The night train wasn’t fast enough, so we upgraded to a rocket. 🚀😂 You’re right, though—I mixed this up with my other benchmark and typed the wrong thing. Good catch! 🤣

How good is the level of detail achieved with this prompt?

you already said sol copied opus then posted the chart anyway, thats a repost not a benchmark
