Loading video...
Video Failed to Load
Sam Altman just said an internal post-Astra model can solve things the world's best mathematicians cannot. in his conversation with Mark Benioff (co-founder and CEO of Salesforce) "GPT 5.5 was maybe as good as like an average math professor. 5.6 was as good as like a maybe top one... show more
344,654 views • 7 days ago •via X (Twitter)
30 Comments

Wow, I wonder what on Earth this could be

Full video

It cannot solve things out side of its data base training it takes a data base to train it on operations transforms coordinates vectors etc with out that it can't perform the operations needed it is not the Ai that does its a human and Ai collaboration this post is mis conception

5.5 is an average math professor, 5.6 beats the world's top mathematicians, and the one on your credit card is a quota wall making you beg for a reset.

He's keening for his IPO. That's all.

Average math professor and top one or two percentile are ranking claims, so some instrument has to produce that ranking. Which one is it? And for the internal model, is there a named problem plus a reviewer outside the lab, or only the description?

I so love it when these technocrats demonstrate their wisdom. My whole day revolves around it!

AI models teasing out solutions with humans using it to brut force problems is not a math professor.

They must put more weight on the value of alignment in post-training than pre if there are still large gaps between release and pre-train.

he said Dyson cages lol ..

Called it. After coding saturated, math was always going to be the next capability claim. I'd still want to see the eval set before believing the post-Astra model beats the best human mathematicians

Ever done sports trivia with AI? Maybe 3rd grade level. Constantly confidently wrong.

yeah, it’s called Mathstra

We went from AI struggling with basic math to AI potentially solving problems the best mathematicians in the world can't. And we're still using it to write “just following up” emails.

I wish people like Sam would have a bit more humility about their claims and not just be all about hype.

5.5 average prof, 5.6 top 1-2%, astra a bit better, internal beats all mathematicians. curve isn't flattening

Los políticos son los que piensan en matar las máquinas resuelves algoritmos

idk, solve things mathematicians cannot is doing a lot of work here

Average math professor → top percentile → better than the best. The jump that matters for education is the opposite direction: can it explain the mistake a Year 12 just made without skipping the why?

Great find, thanks!

一个小版本就从普通数学教授跳到顶尖,这台阶迈得也太快了吧

Hey look over here. Astra can also fly. You are in the right place "DREAMFARCE"

Đối với các AI hiện nay thì có 2 tình huống khác nhau cần phải phân biệt : 1) Có những vấn đề mà các nhà toán học rất giỏi cũng không giải quyết được nhưng AI giải quyết được . Ví dụ : NHỮNG BÀI TOÁN THẾ KỶ . 2) Nhưng cũng có những ( Còn tiếp )

I'll believe when I see that all these math results have a positive effect on the world we live in

That's disrespectful

A model nobody can use solving problems nobody can verify. Convenient.

A lot depends on what domains lend themselves to takeoff. Math seems to me to be one of the most obvious candidates for fast takeoff because you can relentlessly RL train it on problems with either known or provable answers and those answers don't need a wetware lab or something to show they are right. It can all happen within compute. I don't think there are many domains that applies to. In fact math is the only one I can think of in a fully abstracted space where the answers are fully abstracted and can also be proven in that same abstracted space. One could say tailor made for RL where you only need to build out a fairly minimal training harness. Other things will require much more elaborate test setups. ExploitGym is certainly an example but obviously had a few downsides and it's clearly harder to get right. But assuming the labs *can* "perfect" test setups then yeah you aim RL at it and get something superhuman relatively quickly, clearly.

The claim needs a task boundary before it becomes a useful benchmark. Publish problem statements, tool access, proof verification, and independent reruns. “Beyond the best mathematicians” is testable only when the search space and failure rate are visible.

The math benchmark just got a lot more interesting.

'Beats the best mathematicians' and 'does math the way they do' are different claims. GPT-5.6 can hit the score; the open question is whether it produces a checkable proof a human can extend. If not, it's an oracle, not a colleague.
Related Videos
Sensitive content
Lana Rhoades talks about agents in the advlt industry. She said some agents are more pimp-like than others. She revealed that she had 2 agents during her time in the industry and she preferred one to the other because he was calmer and groomed her pretty well. While the other which was considered as one of the best agents in the industry as at that time was not so good to her, he yelled and forced her to do a lot of things she wouldn't have wanted to.
vision
96,048 views • 19 days ago
