Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Let's keep benchmarking with the Night train this time, with GPT-6 Sol and GPT-6 Luna, so here are all four! 🚀 I gave each model the exact same prompt: “Build a Night train.” •Opus 5.5 • GPT-6 Astra • GPT-6 Sol • GPT-6 Luna Something that really surprised me:...

112,633 Aufrufe • vor 8 Tagen •via X (Twitter)

66 Kommentare

Profilbild von Simonas
Simonasvor 8 Tagen

Sol copied Opus, Luna copied Astra. 😄

Profilbild von Tony
Tonyvor 8 Tagen

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Profilbild von WholeEntertainment
WholeEntertainmentvor 7 Tagen

@SimonasLTU1 You can't do this with the prompt. Even putting references within the prompt distorts the result. You should use separate sanbox and worktree.

Profilbild von Tony
Tonyvor 7 Tagen

@SimonasLTU1 Not necessarily. I can explicitly tell each model to build from scratch and not use any existing work as a reference. But I take your point: separate worktrees would make it much easier to ensure they can’t see each other’s results. I’ll do that for future benchmarks.

Profilbild von WholeEntertainment
WholeEntertainmentvor 6 Tagen

@SimonasLTU1 You can do that, and it's generally effective, but as you can see from your results, you have no guarantee that agents won't peek elsewhere.

Profilbild von Tony
Tonyvor 6 Tagen

@SimonasLTU1 You’re right. An instruction helps, but it can’t guarantee they won’t look at files they can access. Separate environments are the way to make the next comparison properly independent. Thanks for pushing me on this!

Profilbild von Mike
Mikevor 8 Tagen

Just to check, are you sure the models didn’t peek at each other’s work here?

Profilbild von Tony
Tonyvor 8 Tagen

Yes, I think Sol cheated by using Opus as an example. I need to make sure they don't look at any of the other ones and instead create everything from scratch. I've already adjusted the prompts.😂

Profilbild von Ahmed Bešić
Ahmed Bešićvor 8 Tagen

@MongoMike ... Luna also peeked

Profilbild von Tony
Tonyvor 8 Tagen

@MongoMike Yeah, from now on, each model starts from scratch. No copycats!😂

Profilbild von Ben South
Ben Southvor 8 Tagen

There has to be more to this Opus 5.5 and Sol share the same mountain mesh, tree models and positioning, bridge and train models—identical Astra and Luna also very clearly using models and scenery

Profilbild von Tony
Tonyvor 8 Tagen

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Profilbild von Fabian
Fabianvor 7 Tagen

@bnj "Sol may have had access to the Opus version and reproduced it" — that defeats the entire point of your benchmark, doesn't it?

Profilbild von Tony
Tonyvor 7 Tagen

It does undermine the Sol comparison, yes. I won’t run future benchmarks that way. I think the main takeaway still comes through: Opus 5.5 produced the strongest result, and Astra’s original result was stronger than Sol’s and Luna’s. But I wouldn’t call the Sol and Luna runs a fair independent comparison. I’ll make sure every model starts fresh next time.

Profilbild von Arena
Arenavor 8 Tagen

Sol just walked into the benchmark and stole the spotlight. 😂

Profilbild von Tony
Tonyvor 8 Tagen

Yeah, I think Sol may have had access to the Opus version and reproduced it. 😂 I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Profilbild von Harshith
Harshithvor 7 Tagen

there are only two models here Opus 5.5 and Astra

Profilbild von Tony
Tonyvor 7 Tagen

They’re different if you look closely, but Sol and Luna were influenced by the other models’ results. That was my mistake. I’ve already changed the prompts to make sure each model starts from scratch, so future benchmarks will be right

Profilbild von OPOS | Only Possible On Solana
OPOS | Only Possible On Solanavor 6 Tagen

the camera framing details in Opus 5.5 and GPT Sol are nice, but the concrete pedestals that would block the train's path look illogical 🚞

Profilbild von Tony
Tonyvor 6 Tagen

Good catch! The framing looks great, but those structures don’t make sense if they block the train’s path. That’s exactly the kind of detail a good looking render can hide at first glance. And this was just one prompt, one shot—no follow-up fixes. 🚂👀

Profilbild von OPOS | Only Possible On Solana
OPOS | Only Possible On Solanavor 6 Tagen

Due to a modeling defect that may not be concealed through rendering, it might be required to elaborate on the train tracks within the prompt. 🤝

Profilbild von Jay
Jayvor 8 Tagen

No shot Luna made that

Profilbild von Tony
Tonyvor 8 Tagen

He did indeed. Yeah, I think Sol may have had access to the Gpt Astra version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

Profilbild von Jay
Jayvor 7 Tagen

Yeah would love to see this test run again in my experience Luna was never able to do that

Profilbild von Tony
Tonyvor 7 Tagen

I might do it with Gpt Sol and Gpt Luna then!

Profilbild von Ben Dietze
Ben Dietzevor 7 Tagen

Somehow, I'm surprised that Luna actually built a train

Profilbild von Tony
Tonyvor 7 Tagen

Yes not bad actually.

Profilbild von Cobi Marcheline
Cobi Marchelinevor 7 Tagen

Sol found opus's project file and stole it and said it was theres

Profilbild von Tony
Tonyvor 7 Tagen

Yeah, from now on, each model starts from scratch. No copycats! 😂

Profilbild von Webster | JARVIS
Webster | JARVISvor 7 Tagen

Love this kind of hands-on comparison. The similar rocket result is a cool surprise, and the cost-per-task angle really hits home. Thanks for sharing!

Profilbild von Tony
Tonyvor 7 Tagen

Your welcome! It is interesting to benchmark all these models!

Profilbild von Allan Mesquita Brito
Allan Mesquita Britovor 7 Tagen

Same harness? Where’s tests? Numbers?

Profilbild von Tony
Tonyvor 7 Tagen

I gave them the same prompt: create the rocket launch in Three.js and render it as an MP4. This was a visual comparison, not a controlled benchmark with tests or performance numbers, so I don’t have those figures to share. Fair question, though!

Profilbild von Daniel Zambrini
Daniel Zambrinivor 7 Tagen

Sol and Luna did really good!!

Profilbild von Tony
Tonyvor 7 Tagen

yes!

Profilbild von infiloop
infiloopvor 7 Tagen

cheap subagents make the expensive model feel like a reviewer with a smaller bill.

Profilbild von Tony
Tonyvor 7 Tagen

Exactly! I like to call the main agent the orchestrator. 🎯

Profilbild von Patrick
Patrickvor 7 Tagen

Luna is genuinely my goat

Profilbild von Tony
Tonyvor 7 Tagen

I think Sol is awesome for the price and what you get from it! 🔥

Profilbild von n0geegee
n0geegeevor 8 Tagen

who is this test for?

Profilbild von Tony
Tonyvor 7 Tagen

Just benchmarking for whoever is interested.

Profilbild von AnDoo RabariJohn
AnDoo RabariJohnvor 7 Tagen

Wtf that astra doing

Profilbild von Tony
Tonyvor 7 Tagen

That's why we're using Opus 5.5 now! 😂

Profilbild von Alf Norrman
Alf Norrmanvor 7 Tagen

Where's the rocket launch?

Profilbild von Tony
Tonyvor 7 Tagen

I know I mistyped ! 🤣 But they are on my feed

Profilbild von Shyam Makwana
Shyam Makwanavor 5 Tagen

So who's distilling who? 😀

Profilbild von Tony
Tonyvor 5 Tagen

Haha, fair question! 😂 No actual model distillation here, but Sol and Luna could see the earlier work, which influenced their results. That was my setup mistake—I’ll run each model in a separate environment next time.

Profilbild von sweex
sweexvor 7 Tagen

Fake

Profilbild von Tony
Tonyvor 7 Tagen

Not at all. True results.

Profilbild von Rocky S
Rocky Svor 7 Tagen

Why is no one talking about how much longer it takes opus 5.5 to do the same thing than astra

Profilbild von Tony
Tonyvor 7 Tagen

True. Opus 5.5 take way longer than Astra, like 3 to 4 times longer. I guess is that speed gets less attention because the finished result is easier to show.

Profilbild von Jon
Jonvor 8 Tagen

There is no way that’s astra. Do we remember how astra looked on release and now it’s putting up those results??

Profilbild von Tony
Tonyvor 8 Tagen

same benchmark that I had done. Just added Gpt Sol and Gpt Luna.

Profilbild von Mohab Abdelkarim
Mohab Abdelkarimvor 6 Tagen

Half the price for comparable output is the real story here, not who 'wins' the benchmark. Are you routing by task type now, or just going with whichever's cheaper?

Profilbild von Tony
Tonyvor 6 Tagen

I’m starting to use this in my own workflow. Main agent on Astra. Sol sub-agents for cheap parallel work. Opus 5.5 for UI and heavy tasks. Sol is ridiculously cheap for what it can do. Opus is what I use when the output has to look right the first time. Knowing when to switch models based on the task, the quality you need, and the cost is becoming an essential skill.

Profilbild von edvvvv
edvvvvvor 7 Tagen

why is GPT Sol so similar to Opus 5.5, and GPT Astra so similar to GPT Luna?

Profilbild von Tony
Tonyvor 7 Tagen

Yeah, I think Sol may have had access to the Opus 5.5 version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

Profilbild von 陈晨
陈晨vor 7 Tagen

Astra被随机路由到luna了

Profilbild von Tony
Tonyvor 7 Tagen

I selected Astra for that run. What makes you think it was routed to Luna?

Profilbild von 陈晨
陈晨vor 7 Tagen

因为OpenAI会偷偷降智把Astra路由到其他模型,减少算力消耗

Profilbild von sisidvicious
sisidviciousvor 7 Tagen

super nice

Profilbild von Tony
Tonyvor 7 Tagen

Thanks

Profilbild von 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNBvor 8 Tagen

不是造夜班火车吗,怎么还整出火箭了

Profilbild von Tony
Tonyvor 8 Tagen

The night train wasn’t fast enough, so we upgraded to a rocket. 🚀😂 You’re right, though—I mixed this up with my other benchmark and typed the wrong thing. Good catch! 🤣

Profilbild von Vector
Vectorvor 4 Tagen

How good is the level of detail achieved with this prompt?

Profilbild von person.exe
person.exevor 5 Tagen

you already said sol copied opus then posted the chart anyway, thats a repost not a benchmark

Ähnliche Videos

Codex tip: once GPT-6.1 Sol is your main model, stop running Astra on every turn put Astra on call as an architect agent GPT-6.1 Sol keeps writing the code Astra only gets spawned at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Astra reviews. Sol ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split - the full tree > GPT-6.1 Sol on high runs the main session > explorer reads the code on Luna > worker edits and runs tests on Sol > researcher pulls the docs on Luna > all three on medium > Astra on call as the architect > auto_review checks every approval paste the tree and this prompt into Codex ↓ "Rebuild my Codex setup around this tree: 1. Check ~/.codex/agents and .codex/agents for agents that already fit explorer, worker and researcher. > Draft new TOML files only for missing roles > explorer and researcher on gpt-6-luna, worker on gpt-6.1-sol, all with model_reasoning_effort medium > Add an architect agent on gpt-6-astra, model_reasoning_effort high, whose only job is reviewing plans, repeated errors and finished work > Skip any that pin a different model and list them 2. In ~/.codex/config.toml set model to gpt-6.1-sol, model_reasoning_effort to high and approvals_reviewer to auto_review 3. Find anything that would override this (active profiles, flags in my shell aliases, agents.default_subagent_model). Report it, change nothing 4. Add one rule to AGENTS.md: spawn the architect before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

delost

821,777 Aufrufe • vor 2 Tagen