Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Mathematics offers a unique window into AI's reasoning capabilities. Discover why we've launched FrontierMath—a benchmark of hundreds of unpublished, expert-level math problems—to understand the frontier of artificial intelligence.

394,050 Aufrufe • vor 1 Jahr •via X (Twitter)

10 Kommentare

Profilbild von Epoch AI
Epoch AIvor 1 Jahr

Learn more about FrontierMath, explore sample problems with detailed solutions, and see how current AI systems perform on FrontierMath:

Profilbild von Daniel Paleka
Daniel Palekavor 1 Jahr

How many problems are there in the dataset? I guess about 290?

Profilbild von Tilman Bayer
Tilman Bayervor 1 Jahr

SOTA among math benchmarks... I see what you did there 😉

Profilbild von Alfred Wahlforss
Alfred Wahlforssvor 1 Jahr

Go Elliot!

Profilbild von Fred Zhang
Fred Zhangvor 1 Jahr

are problem statements formalized?

Profilbild von Charlie Snell
Charlie Snellvor 1 Jahr

Banger

Profilbild von Luis Garicano 🇪🇺🇺🇦
Luis Garicano 🇪🇺🇺🇦vor 1 Jahr

You did not try o1-preview?

Profilbild von João Augusto
João Augustovor 1 Jahr

That is awesome guys!!

Profilbild von YJxAI – e/acc
YJxAI – e/accvor 1 Jahr

IF Humans can take weeks to achieve the result. Should not o1 be given the same amount of time to think . Given its increase in performance by increasing Test time Compute

Profilbild von Jason Rute @ JMM 2025
Jason Rute @ JMM 2025vor 1 Jahr

What is the process by which researchers can submit models to be evaluated on this benchmark? Or are you only interested in leading foundation models? (But even then, you have to consider prompting strategies, no?)

Ähnliche Videos