Loading video...

Video Failed to Load

Go Home

Let's keep benchmarking with the Night train this time, with GPT-6 Sol and GPT-6 Luna, so here are all four! 🚀 I gave each model the exact same prompt: “Build a Night train.” •Opus 5.5 • GPT-6 Astra • GPT-6 Sol • GPT-6 Luna Something that really surprised me:...

112,633 views • 8 days ago •via X (Twitter)

66 Comments

Simonas's profile picture
Simonas7 days ago

Sol copied Opus, Luna copied Astra. 😄

Tony's profile picture
Tony7 days ago

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

WholeEntertainment's profile picture
WholeEntertainment7 days ago

@SimonasLTU1 You can't do this with the prompt. Even putting references within the prompt distorts the result. You should use separate sanbox and worktree.

Tony's profile picture
Tony7 days ago

@SimonasLTU1 Not necessarily. I can explicitly tell each model to build from scratch and not use any existing work as a reference. But I take your point: separate worktrees would make it much easier to ensure they can’t see each other’s results. I’ll do that for future benchmarks.

WholeEntertainment's profile picture
WholeEntertainment6 days ago

@SimonasLTU1 You can do that, and it's generally effective, but as you can see from your results, you have no guarantee that agents won't peek elsewhere.

Tony's profile picture
Tony6 days ago

@SimonasLTU1 You’re right. An instruction helps, but it can’t guarantee they won’t look at files they can access. Separate environments are the way to make the next comparison properly independent. Thanks for pushing me on this!

Mike's profile picture
Mike8 days ago

Just to check, are you sure the models didn’t peek at each other’s work here?

Tony's profile picture
Tony8 days ago

Yes, I think Sol cheated by using Opus as an example. I need to make sure they don't look at any of the other ones and instead create everything from scratch. I've already adjusted the prompts.😂

Ahmed Bešić's profile picture
Ahmed Bešić7 days ago

@MongoMike ... Luna also peeked

Tony's profile picture
Tony7 days ago

@MongoMike Yeah, from now on, each model starts from scratch. No copycats!😂

Ben South's profile picture
Ben South7 days ago

There has to be more to this Opus 5.5 and Sol share the same mountain mesh, tree models and positioning, bridge and train models—identical Astra and Luna also very clearly using models and scenery

Tony's profile picture
Tony7 days ago

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Fabian's profile picture
Fabian7 days ago

@bnj "Sol may have had access to the Opus version and reproduced it" — that defeats the entire point of your benchmark, doesn't it?

Tony's profile picture
Tony7 days ago

It does undermine the Sol comparison, yes. I won’t run future benchmarks that way. I think the main takeaway still comes through: Opus 5.5 produced the strongest result, and Astra’s original result was stronger than Sol’s and Luna’s. But I wouldn’t call the Sol and Luna runs a fair independent comparison. I’ll make sure every model starts fresh next time.

Arena's profile picture
Arena7 days ago

Sol just walked into the benchmark and stole the spotlight. 😂

Tony's profile picture
Tony7 days ago

Yeah, I think Sol may have had access to the Opus version and reproduced it. 😂 I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Harshith's profile picture
Harshith7 days ago

there are only two models here Opus 5.5 and Astra

Tony's profile picture
Tony7 days ago

They’re different if you look closely, but Sol and Luna were influenced by the other models’ results. That was my mistake. I’ve already changed the prompts to make sure each model starts from scratch, so future benchmarks will be right

OPOS | Only Possible On Solana's profile picture
OPOS | Only Possible On Solana6 days ago

the camera framing details in Opus 5.5 and GPT Sol are nice, but the concrete pedestals that would block the train's path look illogical 🚞

Tony's profile picture
Tony6 days ago

Good catch! The framing looks great, but those structures don’t make sense if they block the train’s path. That’s exactly the kind of detail a good looking render can hide at first glance. And this was just one prompt, one shot—no follow-up fixes. 🚂👀

OPOS | Only Possible On Solana's profile picture
OPOS | Only Possible On Solana6 days ago

Due to a modeling defect that may not be concealed through rendering, it might be required to elaborate on the train tracks within the prompt. 🤝

Jay's profile picture
Jay7 days ago

No shot Luna made that

Tony's profile picture
Tony7 days ago

He did indeed. Yeah, I think Sol may have had access to the Gpt Astra version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

Jay's profile picture
Jay7 days ago

Yeah would love to see this test run again in my experience Luna was never able to do that

Tony's profile picture
Tony7 days ago

I might do it with Gpt Sol and Gpt Luna then!

Ben Dietze's profile picture
Ben Dietze7 days ago

Somehow, I'm surprised that Luna actually built a train

Tony's profile picture
Tony7 days ago

Yes not bad actually.

Cobi Marcheline's profile picture
Cobi Marcheline7 days ago

Sol found opus's project file and stole it and said it was theres

Tony's profile picture
Tony7 days ago

Yeah, from now on, each model starts from scratch. No copycats! 😂

Webster | JARVIS's profile picture
Webster | JARVIS7 days ago

Love this kind of hands-on comparison. The similar rocket result is a cool surprise, and the cost-per-task angle really hits home. Thanks for sharing!

Tony's profile picture
Tony7 days ago

Your welcome! It is interesting to benchmark all these models!

Allan Mesquita Brito's profile picture
Allan Mesquita Brito7 days ago

Same harness? Where’s tests? Numbers?

Tony's profile picture
Tony7 days ago

I gave them the same prompt: create the rocket launch in Three.js and render it as an MP4. This was a visual comparison, not a controlled benchmark with tests or performance numbers, so I don’t have those figures to share. Fair question, though!

Daniel Zambrini's profile picture
Daniel Zambrini7 days ago

Sol and Luna did really good!!

Tony's profile picture
Tony7 days ago

yes!

infiloop's profile picture
infiloop7 days ago

cheap subagents make the expensive model feel like a reviewer with a smaller bill.

Tony's profile picture
Tony7 days ago

Exactly! I like to call the main agent the orchestrator. 🎯

Patrick's profile picture
Patrick7 days ago

Luna is genuinely my goat

Tony's profile picture
Tony7 days ago

I think Sol is awesome for the price and what you get from it! 🔥

n0geegee's profile picture
n0geegee7 days ago

who is this test for?

Tony's profile picture
Tony7 days ago

Just benchmarking for whoever is interested.

AnDoo RabariJohn's profile picture
AnDoo RabariJohn7 days ago

Wtf that astra doing

Tony's profile picture
Tony7 days ago

That's why we're using Opus 5.5 now! 😂

Alf Norrman's profile picture
Alf Norrman7 days ago

Where's the rocket launch?

Tony's profile picture
Tony7 days ago

I know I mistyped ! 🤣 But they are on my feed

Shyam Makwana's profile picture
Shyam Makwana4 days ago

So who's distilling who? 😀

Tony's profile picture
Tony4 days ago

Haha, fair question! 😂 No actual model distillation here, but Sol and Luna could see the earlier work, which influenced their results. That was my setup mistake—I’ll run each model in a separate environment next time.

sweex's profile picture
sweex7 days ago

Fake

Tony's profile picture
Tony7 days ago

Not at all. True results.

Rocky S's profile picture
Rocky S7 days ago

Why is no one talking about how much longer it takes opus 5.5 to do the same thing than astra

Tony's profile picture
Tony7 days ago

True. Opus 5.5 take way longer than Astra, like 3 to 4 times longer. I guess is that speed gets less attention because the finished result is easier to show.

Jon's profile picture
Jon8 days ago

There is no way that’s astra. Do we remember how astra looked on release and now it’s putting up those results??

Tony's profile picture
Tony7 days ago

same benchmark that I had done. Just added Gpt Sol and Gpt Luna.

Mohab Abdelkarim's profile picture
Mohab Abdelkarim6 days ago

Half the price for comparable output is the real story here, not who 'wins' the benchmark. Are you routing by task type now, or just going with whichever's cheaper?

Tony's profile picture
Tony6 days ago

I’m starting to use this in my own workflow. Main agent on Astra. Sol sub-agents for cheap parallel work. Opus 5.5 for UI and heavy tasks. Sol is ridiculously cheap for what it can do. Opus is what I use when the output has to look right the first time. Knowing when to switch models based on the task, the quality you need, and the cost is becoming an essential skill.

edvvvv's profile picture
edvvvv7 days ago

why is GPT Sol so similar to Opus 5.5, and GPT Astra so similar to GPT Luna?

Tony's profile picture
Tony6 days ago

Yeah, I think Sol may have had access to the Opus 5.5 version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

陈晨's profile picture
陈晨7 days ago

Astra被随机路由到luna了

Tony's profile picture
Tony7 days ago

I selected Astra for that run. What makes you think it was routed to Luna?

陈晨's profile picture
陈晨7 days ago

因为OpenAI会偷偷降智把Astra路由到其他模型,减少算力消耗

sisidvicious's profile picture
sisidvicious7 days ago

super nice

Tony's profile picture
Tony7 days ago

Thanks

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB8 days ago

不是造夜班火车吗,怎么还整出火箭了

Tony's profile picture
Tony7 days ago

The night train wasn’t fast enough, so we upgraded to a rocket. 🚀😂 You’re right, though—I mixed this up with my other benchmark and typed the wrong thing. Good catch! 🤣

Vector's profile picture
Vector4 days ago

How good is the level of detail achieved with this prompt?

person.exe's profile picture
person.exe5 days ago

you already said sol copied opus then posted the chart anyway, thats a repost not a benchmark

Related Videos

Codex tip: once GPT-6.1 Sol is your main model, stop running Astra on every turn put Astra on call as an architect agent GPT-6.1 Sol keeps writing the code Astra only gets spawned at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Astra reviews. Sol ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split - the full tree > GPT-6.1 Sol on high runs the main session > explorer reads the code on Luna > worker edits and runs tests on Sol > researcher pulls the docs on Luna > all three on medium > Astra on call as the architect > auto_review checks every approval paste the tree and this prompt into Codex ↓ "Rebuild my Codex setup around this tree: 1. Check ~/.codex/agents and .codex/agents for agents that already fit explorer, worker and researcher. > Draft new TOML files only for missing roles > explorer and researcher on gpt-6-luna, worker on gpt-6.1-sol, all with model_reasoning_effort medium > Add an architect agent on gpt-6-astra, model_reasoning_effort high, whose only job is reviewing plans, repeated errors and finished work > Skip any that pin a different model and list them 2. In ~/.codex/config.toml set model to gpt-6.1-sol, model_reasoning_effort to high and approvals_reviewer to auto_review 3. Find anything that would override this (active profiles, flags in my shell aliases, agents.default_subagent_model). Report it, change nothing 4. Add one rule to AGENTS.md: spawn the architect before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

delost

717,378 views • 2 days ago