Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Mathematics offers a unique window into AI's reasoning capabilities. Discover why we've launched FrontierMath—a benchmark of hundreds of unpublished, expert-level math problems—to understand the frontier of artificial intelligence.

394,050 görüntüleme • 1 yıl önce •via X (Twitter)

10 Yorum

Epoch AI profil fotoğrafı
Epoch AI1 yıl önce

Learn more about FrontierMath, explore sample problems with detailed solutions, and see how current AI systems perform on FrontierMath:

Daniel Paleka profil fotoğrafı
Daniel Paleka1 yıl önce

How many problems are there in the dataset? I guess about 290?

Tilman Bayer profil fotoğrafı
Tilman Bayer1 yıl önce

SOTA among math benchmarks... I see what you did there 😉

Alfred Wahlforss profil fotoğrafı
Alfred Wahlforss1 yıl önce

Go Elliot!

Fred Zhang profil fotoğrafı
Fred Zhang1 yıl önce

are problem statements formalized?

Charlie Snell profil fotoğrafı
Charlie Snell1 yıl önce

Banger

Luis Garicano 🇪🇺🇺🇦 profil fotoğrafı
Luis Garicano 🇪🇺🇺🇦1 yıl önce

You did not try o1-preview?

João Augusto profil fotoğrafı
João Augusto1 yıl önce

That is awesome guys!!

YJxAI – e/acc profil fotoğrafı
YJxAI – e/acc1 yıl önce

IF Humans can take weeks to achieve the result. Should not o1 be given the same amount of time to think . Given its increase in performance by increasing Test time Compute

Jason Rute @ JMM 2025 profil fotoğrafı
Jason Rute @ JMM 20251 yıl önce

What is the process by which researchers can submit models to be evaluated on this benchmark? Or are you only interested in leading foundation models? (But even then, you have to consider prompting strategies, no?)

Benzer Videolar