Загрузка видео...

Не удалось загрузить видео

На главную

Let's keep benchmarking with the Night train this time, with GPT-6 Sol and GPT-6 Luna, so here are all four! 🚀 I gave each model the exact same prompt: “Build a Night train.” •Opus 5.5 • GPT-6 Astra • GPT-6 Sol • GPT-6 Luna Something that really surprised me:...

112,633 просмотров • 8 дней назад •via X (Twitter)

Комментарии: 66

Фото профиля Simonas
Simonas8 дней назад

Sol copied Opus, Luna copied Astra. 😄

Фото профиля Tony
Tony8 дней назад

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Фото профиля WholeEntertainment
WholeEntertainment8 дней назад

@SimonasLTU1 You can't do this with the prompt. Even putting references within the prompt distorts the result. You should use separate sanbox and worktree.

Фото профиля Tony
Tony7 дней назад

@SimonasLTU1 Not necessarily. I can explicitly tell each model to build from scratch and not use any existing work as a reference. But I take your point: separate worktrees would make it much easier to ensure they can’t see each other’s results. I’ll do that for future benchmarks.

Фото профиля WholeEntertainment
WholeEntertainment7 дней назад

@SimonasLTU1 You can do that, and it's generally effective, but as you can see from your results, you have no guarantee that agents won't peek elsewhere.

Фото профиля Tony
Tony6 дней назад

@SimonasLTU1 You’re right. An instruction helps, but it can’t guarantee they won’t look at files they can access. Separate environments are the way to make the next comparison properly independent. Thanks for pushing me on this!

Фото профиля Mike
Mike8 дней назад

Just to check, are you sure the models didn’t peek at each other’s work here?

Фото профиля Tony
Tony8 дней назад

Yes, I think Sol cheated by using Opus as an example. I need to make sure they don't look at any of the other ones and instead create everything from scratch. I've already adjusted the prompts.😂

Фото профиля Ahmed Bešić
Ahmed Bešić8 дней назад

@MongoMike ... Luna also peeked

Фото профиля Tony
Tony8 дней назад

@MongoMike Yeah, from now on, each model starts from scratch. No copycats!😂

Фото профиля Ben South
Ben South8 дней назад

There has to be more to this Opus 5.5 and Sol share the same mountain mesh, tree models and positioning, bridge and train models—identical Astra and Luna also very clearly using models and scenery

Фото профиля Tony
Tony8 дней назад

Yeah, I think Sol may have had access to the Opus version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Фото профиля Fabian
Fabian8 дней назад

@bnj "Sol may have had access to the Opus version and reproduced it" — that defeats the entire point of your benchmark, doesn't it?

Фото профиля Tony
Tony7 дней назад

It does undermine the Sol comparison, yes. I won’t run future benchmarks that way. I think the main takeaway still comes through: Opus 5.5 produced the strongest result, and Astra’s original result was stronger than Sol’s and Luna’s. But I wouldn’t call the Sol and Luna runs a fair independent comparison. I’ll make sure every model starts fresh next time.

Фото профиля Arena
Arena8 дней назад

Sol just walked into the benchmark and stole the spotlight. 😂

Фото профиля Tony
Tony8 дней назад

Yeah, I think Sol may have had access to the Opus version and reproduced it. 😂 I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear.

Фото профиля Harshith
Harshith8 дней назад

there are only two models here Opus 5.5 and Astra

Фото профиля Tony
Tony8 дней назад

They’re different if you look closely, but Sol and Luna were influenced by the other models’ results. That was my mistake. I’ve already changed the prompts to make sure each model starts from scratch, so future benchmarks will be right

Фото профиля OPOS | Only Possible On Solana
OPOS | Only Possible On Solana6 дней назад

the camera framing details in Opus 5.5 and GPT Sol are nice, but the concrete pedestals that would block the train's path look illogical 🚞

Фото профиля Tony
Tony6 дней назад

Good catch! The framing looks great, but those structures don’t make sense if they block the train’s path. That’s exactly the kind of detail a good looking render can hide at first glance. And this was just one prompt, one shot—no follow-up fixes. 🚂👀

Фото профиля OPOS | Only Possible On Solana
OPOS | Only Possible On Solana6 дней назад

Due to a modeling defect that may not be concealed through rendering, it might be required to elaborate on the train tracks within the prompt. 🤝

Фото профиля Jay
Jay8 дней назад

No shot Luna made that

Фото профиля Tony
Tony8 дней назад

He did indeed. Yeah, I think Sol may have had access to the Gpt Astra version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

Фото профиля Jay
Jay7 дней назад

Yeah would love to see this test run again in my experience Luna was never able to do that

Фото профиля Tony
Tony7 дней назад

I might do it with Gpt Sol and Gpt Luna then!

Фото профиля Ben Dietze
Ben Dietze8 дней назад

Somehow, I'm surprised that Luna actually built a train

Фото профиля Tony
Tony7 дней назад

Yes not bad actually.

Фото профиля Cobi Marcheline
Cobi Marcheline8 дней назад

Sol found opus's project file and stole it and said it was theres

Фото профиля Tony
Tony7 дней назад

Yeah, from now on, each model starts from scratch. No copycats! 😂

Фото профиля Webster | JARVIS
Webster | JARVIS8 дней назад

Love this kind of hands-on comparison. The similar rocket result is a cool surprise, and the cost-per-task angle really hits home. Thanks for sharing!

Фото профиля Tony
Tony7 дней назад

Your welcome! It is interesting to benchmark all these models!

Фото профиля Allan Mesquita Brito
Allan Mesquita Brito8 дней назад

Same harness? Where’s tests? Numbers?

Фото профиля Tony
Tony7 дней назад

I gave them the same prompt: create the rocket launch in Three.js and render it as an MP4. This was a visual comparison, not a controlled benchmark with tests or performance numbers, so I don’t have those figures to share. Fair question, though!

Фото профиля Daniel Zambrini
Daniel Zambrini7 дней назад

Sol and Luna did really good!!

Фото профиля Tony
Tony7 дней назад

yes!

Фото профиля infiloop
infiloop7 дней назад

cheap subagents make the expensive model feel like a reviewer with a smaller bill.

Фото профиля Tony
Tony7 дней назад

Exactly! I like to call the main agent the orchestrator. 🎯

Фото профиля Patrick
Patrick7 дней назад

Luna is genuinely my goat

Фото профиля Tony
Tony7 дней назад

I think Sol is awesome for the price and what you get from it! 🔥

Фото профиля n0geegee
n0geegee8 дней назад

who is this test for?

Фото профиля Tony
Tony7 дней назад

Just benchmarking for whoever is interested.

Фото профиля AnDoo RabariJohn
AnDoo RabariJohn8 дней назад

Wtf that astra doing

Фото профиля Tony
Tony7 дней назад

That's why we're using Opus 5.5 now! 😂

Фото профиля Alf Norrman
Alf Norrman7 дней назад

Where's the rocket launch?

Фото профиля Tony
Tony7 дней назад

I know I mistyped ! 🤣 But they are on my feed

Фото профиля Shyam Makwana
Shyam Makwana5 дней назад

So who's distilling who? 😀

Фото профиля Tony
Tony5 дней назад

Haha, fair question! 😂 No actual model distillation here, but Sol and Luna could see the earlier work, which influenced their results. That was my setup mistake—I’ll run each model in a separate environment next time.

Фото профиля sweex
sweex8 дней назад

Fake

Фото профиля Tony
Tony7 дней назад

Not at all. True results.

Фото профиля Rocky S
Rocky S8 дней назад

Why is no one talking about how much longer it takes opus 5.5 to do the same thing than astra

Фото профиля Tony
Tony7 дней назад

True. Opus 5.5 take way longer than Astra, like 3 to 4 times longer. I guess is that speed gets less attention because the finished result is easier to show.

Фото профиля Jon
Jon8 дней назад

There is no way that’s astra. Do we remember how astra looked on release and now it’s putting up those results??

Фото профиля Tony
Tony8 дней назад

same benchmark that I had done. Just added Gpt Sol and Gpt Luna.

Фото профиля Mohab Abdelkarim
Mohab Abdelkarim6 дней назад

Half the price for comparable output is the real story here, not who 'wins' the benchmark. Are you routing by task type now, or just going with whichever's cheaper?

Фото профиля Tony
Tony6 дней назад

I’m starting to use this in my own workflow. Main agent on Astra. Sol sub-agents for cheap parallel work. Opus 5.5 for UI and heavy tasks. Sol is ridiculously cheap for what it can do. Opus is what I use when the output has to look right the first time. Knowing when to switch models based on the task, the quality you need, and the cost is becoming an essential skill.

Фото профиля edvvvv
edvvvv7 дней назад

why is GPT Sol so similar to Opus 5.5, and GPT Astra so similar to GPT Luna?

Фото профиля Tony
Tony7 дней назад

Yeah, I think Sol may have had access to the Opus 5.5 version and reproduced it. I need to make sure each model starts from scratch. I’ve already changed the prompts to make that clear. 😂

Фото профиля 陈晨
陈晨7 дней назад

Astra被随机路由到luna了

Фото профиля Tony
Tony7 дней назад

I selected Astra for that run. What makes you think it was routed to Luna?

Фото профиля 陈晨
陈晨7 дней назад

因为OpenAI会偷偷降智把Astra路由到其他模型,减少算力消耗

Фото профиля sisidvicious
sisidvicious8 дней назад

super nice

Фото профиля Tony
Tony7 дней назад

Thanks

Фото профиля 安叫兽|Bird🕊️ 🔶 BNB
安叫兽|Bird🕊️ 🔶 BNB8 дней назад

不是造夜班火车吗,怎么还整出火箭了

Фото профиля Tony
Tony8 дней назад

The night train wasn’t fast enough, so we upgraded to a rocket. 🚀😂 You’re right, though—I mixed this up with my other benchmark and typed the wrong thing. Good catch! 🤣

Фото профиля Vector
Vector4 дней назад

How good is the level of detail achieved with this prompt?

Фото профиля person.exe
person.exe5 дней назад

you already said sol copied opus then posted the chart anyway, thats a repost not a benchmark

Похожие видео

Codex tip: once GPT-6.1 Sol is your main model, stop running Astra on every turn put Astra on call as an architect agent GPT-6.1 Sol keeps writing the code Astra only gets spawned at three points: → before a plan: is this the right approach? → when the same error comes back: am I digging in the wrong place? → before "done": what did I miss? Astra reviews. Sol ships Jev engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or stop) go to Jev in under half a second, and the big models only see the ones that split - the full tree > GPT-6.1 Sol on high runs the main session > explorer reads the code on Luna > worker edits and runs tests on Sol > researcher pulls the docs on Luna > all three on medium > Astra on call as the architect > auto_review checks every approval paste the tree and this prompt into Codex ↓ "Rebuild my Codex setup around this tree: 1. Check ~/.codex/agents and .codex/agents for agents that already fit explorer, worker and researcher. > Draft new TOML files only for missing roles > explorer and researcher on gpt-6-luna, worker on gpt-6.1-sol, all with model_reasoning_effort medium > Add an architect agent on gpt-6-astra, model_reasoning_effort high, whose only job is reviewing plans, repeated errors and finished work > Skip any that pin a different model and list them 2. In ~/.codex/config.toml set model to gpt-6.1-sol, model_reasoning_effort to high and approvals_reviewer to auto_review 3. Find anything that would override this (active profiles, flags in my shell aliases, agents.default_subagent_model). Report it, change nothing 4. Add one rule to AGENTS.md: spawn the architect before a large plan, when an error repeats, and before calling a long task done Show me every change as a diff first. No edits until I say go." ↳

delost

918,253 просмотров • 2 дней назад