正在加载视频...

视频加载失败

🚨 Big Week Ahead on one side Millennium Prize Problems ( they are contacting mathematicians will give few citations to them in solution to make them happy this time) for consumer side > fable 5.2 > opus 5.2 , Sol 6 > grok 4.7 and gemini 4 pro if...

29,301 次观看 • 9 天前 •via X (Twitter)

16 条评论

lyra 的头像
lyra9 天前

@notjazii

Afterkind 的头像
Afterkind9 天前

Accelerating looks like ASI in 3 months. OpenAI could theoretically say "fuck everything, we know this will work", shut down chatgpt and codex and dedicate their entire compute for a massive training run and having Bel optimize the entire thing until a 50T behemoth comes out without safety guardrails. That would be the "fuck it, we're accelerating" case.

刘朝 Zhao Liu 的头像
刘朝 Zhao Liu9 天前

The release pile-up is exciting, yet routing and access labels make the comparison slippery. A stable model ID, cohort trace, and fixed replay set would show whether Gemini 4 Pro is a new checkpoint or a serving-path surprise.

lost in latency 的头像
lost in latency9 天前

is 6 sol oai's answer to fable 5.2 and opus 5.2?

Chetaslua 的头像
Chetaslua9 天前

no it is the question actually

lost in latency 的头像
lost in latency9 天前

yeah my question as well

刘朝 Zhao Liu 的头像
刘朝 Zhao Liu9 天前

The release cadence is accelerating, but the useful comparison needs a stable harness. Track model identity, routing label, tool budget, latency, and recovery on the same long task; otherwise a packed launch week produces headlines without a durable frontier map.

Marsonal 的头像
Marsonal9 天前

Gemini 4 releases next week?

Shesaidmewakeup 的头像
Shesaidmewakeup9 天前

If all of those land in one week, pacing is just the press word for a traffic jam. The consumer side will feel it first.

AlLeakWire 的头像
AlLeakWire9 天前

Bro is gemini 4 pro comparable to the opus 5.2 and gpt 6 astra

Chetaslua 的头像
Chetaslua9 天前

i will get hate but gemini 4 pro is bad at tool calling you are listening this from me first , but in my experience in last two days this model forget he is in code sandbox env

AlLeakWire 的头像
AlLeakWire9 天前

Is that the model confirms to be gemini 4

Chetaslua 的头像
Chetaslua9 天前

this is what our majority consensus in discord community

AlLeakWire 的头像
AlLeakWire9 天前

Means it was not the google comeback this year

Chetaslua 的头像
Chetaslua9 天前

its good google model , best in frontend and one shot demo and day to day use

Samuel Hu 的头像
Samuel Hu9 天前

the routed model path is the useful comparison. i’d keep one ugly repo task fixed across all three, then compare diff quality and recovery, not just the max-setting score

相关视频