Загрузка видео...
Не удалось загрузить видео
🚨 Big Week Ahead on one side Millennium Prize Problems ( they are contacting mathematicians will give few citations to them in solution to make them happy this time) for consumer side > fable 5.2 > opus 5.2 , Sol 6 > grok 4.7 and gemini 4 pro if... show more
26,767 просмотров • 1 день назад •via X (Twitter)
Комментарии: 16

@notjazii

Accelerating looks like ASI in 3 months. OpenAI could theoretically say "fuck everything, we know this will work", shut down chatgpt and codex and dedicate their entire compute for a massive training run and having Bel optimize the entire thing until a 50T behemoth comes out without safety guardrails. That would be the "fuck it, we're accelerating" case.

The release pile-up is exciting, yet routing and access labels make the comparison slippery. A stable model ID, cohort trace, and fixed replay set would show whether Gemini 4 Pro is a new checkpoint or a serving-path surprise.

is 6 sol oai's answer to fable 5.2 and opus 5.2?

no it is the question actually

yeah my question as well

The release cadence is accelerating, but the useful comparison needs a stable harness. Track model identity, routing label, tool budget, latency, and recovery on the same long task; otherwise a packed launch week produces headlines without a durable frontier map.

Gemini 4 releases next week?

If all of those land in one week, pacing is just the press word for a traffic jam. The consumer side will feel it first.

Bro is gemini 4 pro comparable to the opus 5.2 and gpt 6 astra

i will get hate but gemini 4 pro is bad at tool calling you are listening this from me first , but in my experience in last two days this model forget he is in code sandbox env

Is that the model confirms to be gemini 4

this is what our majority consensus in discord community

Means it was not the google comeback this year

its good google model , best in frontend and one shot demo and day to day use

the routed model path is the useful comparison. i’d keep one ugly repo task fixed across all three, then compare diff quality and recovery, not just the max-setting score
