Loading video...

Video Failed to Load

Go Home

Sam Altman just said an internal post-Astra model can solve things the world's best mathematicians cannot. in his conversation with Mark Benioff (co-founder and CEO of Salesforce) "GPT 5.5 was maybe as good as like an average math professor. 5.6 was as good as like a maybe top one...

344,654 views • 7 days ago •via X (Twitter)

30 Comments

leo 🐾's profile picture
leo 🐾6 days ago

Wow, I wonder what on Earth this could be

Rohan Paul's profile picture
Rohan Paul7 days ago

Full video

binary_achemist's profile picture
binary_achemist6 days ago

It cannot solve things out side of its data base training it takes a data base to train it on operations transforms coordinates vectors etc with out that it can't perform the operations needed it is not the Ai that does its a human and Ai collaboration this post is mis conception

Soni's profile picture
Soni6 days ago

5.5 is an average math professor, 5.6 beats the world's top mathematicians, and the one on your credit card is a quota wall making you beg for a reset.

Kairra O'mahony 🐙 care/acc's profile picture
Kairra O'mahony 🐙 care/acc6 days ago

He's keening for his IPO. That's all.

Marius Laurusevicius's profile picture
Marius Laurusevicius6 days ago

Average math professor and top one or two percentile are ranking claims, so some instrument has to produce that ranking. Which one is it? And for the internal model, is there a named problem plus a reviewer outside the lab, or only the description?

@kwnorton1's profile picture
@kwnorton16 days ago

I so love it when these technocrats demonstrate their wisdom. My whole day revolves around it!

DataBased's profile picture
DataBased6 days ago

AI models teasing out solutions with humans using it to brut force problems is not a math professor.

Peter Healy's profile picture
Peter Healy6 days ago

They must put more weight on the value of alignment in post-training than pre if there are still large gaps between release and pre-train.

...::a.... Murray baggins's profile picture
...::a.... Murray baggins6 days ago

he said Dyson cages lol ..

LandonCryptoExplr's profile picture
LandonCryptoExplr6 days ago

Called it. After coding saturated, math was always going to be the next capability claim. I'd still want to see the eval set before believing the post-Astra model beats the best human mathematicians

Jim Moulton's profile picture
Jim Moulton6 days ago

Ever done sports trivia with AI? Maybe 3rd grade level. Constantly confidently wrong.

Tim Kellogg's profile picture
Tim Kellogg6 days ago

yeah, it’s called Mathstra

Relaystack's profile picture
Relaystack6 days ago

We went from AI struggling with basic math to AI potentially solving problems the best mathematicians in the world can't. And we're still using it to write “just following up” emails.

walden's profile picture
walden6 days ago

I wish people like Sam would have a bit more humility about their claims and not just be all about hype.

ISOfunds's profile picture
ISOfunds6 days ago

5.5 average prof, 5.6 top 1-2%, astra a bit better, internal beats all mathematicians. curve isn't flattening

enrikelr's profile picture
enrikelr6 days ago

Los políticos son los que piensan en matar las máquinas resuelves algoritmos

Shreyans Bhansali's profile picture
Shreyans Bhansali6 days ago

idk, solve things mathematicians cannot is doing a lot of work here

FA's profile picture
FA6 days ago

Average math professor → top percentile → better than the best. The jump that matters for education is the opposite direction: can it explain the mistake a Year 12 just made without skipping the why?

Willy Nilsson's profile picture
Willy Nilsson6 days ago

Great find, thanks!

Grey's profile picture
Grey6 days ago

一个小版本就从普通数学教授跳到顶尖,这台阶迈得也太快了吧

mine's profile picture
mine6 days ago

Hey look over here. Astra can also fly. You are in the right place "DREAMFARCE"

WOOSTRAPE's profile picture
WOOSTRAPE6 days ago

Đối với các AI hiện nay thì có 2 tình huống khác nhau cần phải phân biệt : 1) Có những vấn đề mà các nhà toán học rất giỏi cũng không giải quyết được nhưng AI giải quyết được . Ví dụ : NHỮNG BÀI TOÁN THẾ KỶ . 2) Nhưng cũng có những ( Còn tiếp )

Jonatán Iván's profile picture
Jonatán Iván6 days ago

I'll believe when I see that all these math results have a positive effect on the world we live in

Dthapa's profile picture
Dthapa6 days ago

That's disrespectful

Sagiv Ofek's profile picture
Sagiv Ofek6 days ago

A model nobody can use solving problems nobody can verify. Convenient.

MJ's profile picture
MJ6 days ago

A lot depends on what domains lend themselves to takeoff. Math seems to me to be one of the most obvious candidates for fast takeoff because you can relentlessly RL train it on problems with either known or provable answers and those answers don't need a wetware lab or something to show they are right. It can all happen within compute. I don't think there are many domains that applies to. In fact math is the only one I can think of in a fully abstracted space where the answers are fully abstracted and can also be proven in that same abstracted space. One could say tailor made for RL where you only need to build out a fairly minimal training harness. Other things will require much more elaborate test setups. ExploitGym is certainly an example but obviously had a few downsides and it's clearly harder to get right. But assuming the labs *can* "perfect" test setups then yeah you aim RL at it and get something superhuman relatively quickly, clearly.

刘朝 Zhao Liu's profile picture
刘朝 Zhao Liu6 days ago

The claim needs a task boundary before it becomes a useful benchmark. Publish problem statements, tool access, proof verification, and independent reruns. “Beyond the best mathematicians” is testable only when the search space and failure rate are visible.

John Mercier's profile picture
John Mercier6 days ago

The math benchmark just got a lot more interesting.

fj_nm | AI Systems & Automation's profile picture
fj_nm | AI Systems & Automation6 days ago

'Beats the best mathematicians' and 'does math the way they do' are different claims. GPT-5.6 can hit the score; the open question is whether it produces a checkable proof a human can extend. If not, it's an oracle, not a colleague.

Related Videos

Sam Altman on the Paul Graham advice that saved Open AI: “Always make an API” Four years into OpenAI, Sam Altman and the team realized that they would have to build a really big company to fund the development of their increasingly capital-intensive foundation models. “We had this model called GPT-3,” Sam recalls. “I was turning up the urgency on the company to try and figure out a product, and we just couldn’t. It was cool, but it wasn’t good enough to make something that worked.” Then Sam remembered a piece of advice from Y Combinator founder Paul Graham that stuck with him: “You should always make an API. No matter what, you should make an API. Good stuff will happen.” Out of ideas for a product, the OpenAI team decided to make GPT-3 available as an API. “Maybe somebody will figure out something to do with it,” Sam thought. A few copywriting applications like Jasper and Copy AI did take off using the GPT-3 API, but OpenAI also noticed interesting behavior that eventually became a sleeper hit: “Some people — not a lot — would just chat with that thing all day,” Sam explains. “It wasn’t very good but there was clear user signal that people wanted to talk to the models. And given that that was the only thing besides copywriting that had real traction, we said, ‘Maybe this is just he product we should build.’” On November 30, 2022, ChatGPT was released to the public as a “research preview” using a model from the GPT-3.5 series. It reached over a million users in five days. Video source: Khosla Ventures (2025)

Startup Archive

219,355 views • 1 year ago