Загрузка видео...

Не удалось загрузить видео

На главную

Mathematics offers a unique window into AI's reasoning capabilities. Discover why we've launched FrontierMath—a benchmark of hundreds of unpublished, expert-level math problems—to understand the frontier of artificial intelligence.

394,050 просмотров • 1 год назад •via X (Twitter)

Комментарии: 10

Фото профиля Epoch AI
Epoch AI1 год назад

Learn more about FrontierMath, explore sample problems with detailed solutions, and see how current AI systems perform on FrontierMath:

Фото профиля Daniel Paleka
Daniel Paleka1 год назад

How many problems are there in the dataset? I guess about 290?

Фото профиля Tilman Bayer
Tilman Bayer1 год назад

SOTA among math benchmarks... I see what you did there 😉

Фото профиля Alfred Wahlforss
Alfred Wahlforss1 год назад

Go Elliot!

Фото профиля Fred Zhang
Fred Zhang1 год назад

are problem statements formalized?

Фото профиля Charlie Snell
Charlie Snell1 год назад

Banger

Фото профиля Luis Garicano 🇪🇺🇺🇦
Luis Garicano 🇪🇺🇺🇦1 год назад

You did not try o1-preview?

Фото профиля João Augusto
João Augusto1 год назад

That is awesome guys!!

Фото профиля YJxAI – e/acc
YJxAI – e/acc1 год назад

IF Humans can take weeks to achieve the result. Should not o1 be given the same amount of time to think . Given its increase in performance by increasing Test time Compute

Фото профиля Jason Rute @ JMM 2025
Jason Rute @ JMM 20251 год назад

What is the process by which researchers can submit models to be evaluated on this benchmark? Or are you only interested in leading foundation models? (But even then, you have to consider prompting strategies, no?)

Похожие видео