Video yükleniyor...
Video Yüklenemedi
opus 5 is getting routed to opus 5.2 tested opus 5 on xhigh with the same prompt in devin and claude code but results came out really different > claude code one is clean and more detailed > while devin one is almost there but not the same opus... show more
140,113 görüntüleme • 10 gün önce •via X (Twitter)
44 Yorum

here's another voxel test. take a look at detailing looks like we gonna see opus new model this thursday

what happened to opus 5.1 ?

they skipped 5.1 no idea why but 5.2 has been spotted already

saw this guys post from the other day about it but he fakes some stuff so idk

was confirmed by leo and it was leaked in repo too

ah nice

I think this is just harness difference

I'd check the id before assuming a swap. Devin and Claude Code ship different system prompts and different tools, so the same prompt is not the same test. /status in Claude Code shows the exact model it's talking to, opus 5 or 5.2.

we are running some knowledge cut off question and for some users it seems opus 5.2 is available lemme do few more before i could say for confirm

same model name, different harness = different animal. if the ui says opus 5 and the logs say 5.2, you weren't comparing models. you were comparing wrappers. i only trust the session that dumps the exact model string before the first tool call.

Hey, i Just came here to remind those that are reading this. You CAN test all of these models in Devin but also, SWE 2.0 is free

bro whatever else than sonnet i use these days the cost is not reasonable at all

use swe 2.. opus 5 and kimi k3 level model for free in devin 20 bucks plan for month unlimited

Ooooo

what's the cutoff question

ask it if it knows tibo the reset guy without search it should know

@be_arsh using hermes - op5 high "do you know tibo the reset guy? tell me without search": "Partially... Thibault "Tibo" Louis-Lucas (@tibo_maker).. The "reset guy" tag is where my confidence drops... So: ~80% on who Tibo is, ~30% that the reset guy is him.

@be_arsh @tibo_maker won't work in hermes just claude sub not api

Yes it is. Opus Max uses much more thinking Tokens than Opus 5 since yesterday. Can confirm the routing

why not Opus 5.1 v Opus 5.2?

no idea why but they skipped opus 5.1 or maybe this is opus 5.1 or 5.2 can't say for sure

the routing to 5.2 tracks, but i'd bet cc's agentic scaffolding explains most of that quality gap, not the model itself. devin handles context chunking totally differently. did you try a raw api call to isolate the model from the tool loop?

Well Opus 5 devin made it noticeably better

post the screenshots of the terminal and api calls, other wise bs

same prompt across two coding tools is still a different test since each tool shapes the result. routing may be happening, but output alone can't prove it.

the same prompt producing cleaner output in claude code is a pretty strong clue. did you compare the raw model/version metadata too, or just the generations?

Fable 5.1 flat out felt better to talk to so an improvement but for Opus’ issues would be huge

This would explain why it completely smoked sol in the exact same task on isolated worktree's, was just letting fable5.1 look for related issues in the codex adapter but it found none...

Bro finally lock in

Silent model upgrades are the best kind — one day the code's just cleaner and you can't even explain why. Secret routing stays winning. 🔍

You believe that this is a routing block and not a harness difference ?

The output shift is a useful clue, but it does not prove a backend change. JAZII, could you repeat both runs with fixed settings and compare latency, token use, and any model metadata?

Something changed in the last 2 days" — yeah, quiet model updates are back. Gotta love when they push better weights under the hood without telling anyone

Could simply be the harness

The only thing that’s clear is that Claude Code turned out much neater

ngl same model, different harness, different personality. The routing story is just the wrapper leaking into the answer.

这两天确实像偷偷换了个脑子,输出突然勤快了

Clean split: freeze a small probe battery, run it daily, and compare devin vs Claude Code traces across runs - a routing change shows up as a shift, not one sample. And keep the boring hypothesis: if only Claude Code got less lazy, that is a harness update, not a swap.

CC 里 `/status` 先看一眼,确认是 opus 5 还是 5.2,再对比。两边系统提示不一样,光看输出差异很难分。

can you provide the prompt for this, or a live link to view this?

And Opus 5.1? They are just gonna skip it :c

yeah

感受不到,opus实在太强了,完全摸不到上限,astra完全无法处理的难题交给opus5.0都可以一次完成

See2

