Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

can AI do research-level mathematics? make conjectures? prove theorems? there’s a moving frontier between what can and cannot be done with LLMs. that boundary just shifted a little. this is my experience with AI proving a new theorem. 1/

355,922 görüntüleme • 1 yıl önce •via X (Twitter)

23 Yorum

prof-g profil fotoğrafı
prof-g1 yıl önce

this is joint work with julian gould and miguel lopez, phd students in my lab @Penn . this is also joint work with claude-3.5-sonnet, gemini-1.5-pro, gpt-4o, and gpt-o1-mini . it’s a 25 page paper on network information flows and lattice theory. (link to preprint at end) 2/

prof-g profil fotoğrafı
prof-g1 yıl önce

quick summary: aug ’24 : claude-3.5/gpt-4o conjectured a new theorem. sept ’24 : generated many wrong proofs w/claude+gpt+gemini, mapping out the subspace of latent proof-space. sept 13 ’24 : gpt-o1-mini dropped & nailed a correct + elegant proof. oct 1 ’24 => arxiv preprint. 3/

prof-g profil fotoğrafı
prof-g1 yıl önce

these theorems were born in a pub on aug 10. i was on my phone w/gpt+claude, having a wide-ranging chat about lattice theory and financial models. on a whim, i asked if the results could give a new lattice-theoretic proof of the classical max-flow-min-cut theorem. 4/

prof-g profil fotoğrafı
prof-g1 yıl önce

recall, MFMC says that in a source-target directed network, if you have a conservative flow which respects edge capacity constraints, the maximal flow value equals the minimal value of a cut separating source from target. classic. 5/

prof-g profil fotoğrafı
prof-g1 yıl önce

gpt/claude gave a faulty lattice-theoretic proof of the classic MFMC theorem. alas. but it led to the question, could one have capacity constraints that were valued in arbitrary lattices? that is when claude-3.5 and gpt-4o both surprised me with the same conjecture. 6/

prof-g profil fotoğrafı
prof-g1 yıl önce

recall, lattices are partially-ordered sets with join (V) and meet (^) operations that act like union/max/or & intersection/min/and respectively. 7/

prof-g profil fotoğrafı
prof-g1 yıl önce

lattices include everything from Boolean algebras to power sets; they play an important role in logic, information theory, computer science, and much more. (these have no relation to crystal lattices appearing in physics) 8/

prof-g profil fotoğrafı
prof-g1 yıl önce

both gpt-4o and claude-3.5 independently guessed that a version of MFMC with lattice-valued edge capacities and flow-conservation using the join V should be true. that would be an interesting new result. but is it true? 9/

prof-g profil fotoğrafı
prof-g1 yıl önce

well, no, it’s not true in general. claude helped find a counterexample in the case of a non-distributive lattice. we [humans] found a counterexample in the case of a non-modular lattice. so, if true, the lattice would definitely need to be distributive, maybe more. 10/

prof-g profil fotoğrafı
prof-g1 yıl önce

we branched investigations into three lattice-valued duality conjectures: * flow-cut duality (MFMC) * path-cut duality (bottleneck) * chain-antichain duality (Dilworth’s theorem) all novel; the most primal form seemed to be bottleneck duality. 11/

prof-g profil fotoğrafı
prof-g1 yıl önce

i spent late aug + early sept using claude-3.5/gemini-1.5-pro/gpt-4o generating proofs. all wrong, but increasingly subtle & hard to tell if right or not. meanwhile, my team came up with human-generated independent proofs. it was a contest. 12/

prof-g profil fotoğrafı
prof-g1 yıl önce

humans in the lead! the language model proofs were sometimes convincing but wrong. this was frustrating. at one point claude got frustrated too and suggested i talk to an expert. it was such a rollercoaster seeing what looked plausible, but then, no. 13/

prof-g profil fotoğrafı
prof-g1 yıl önce

we soon noticed patterns in the types of proofs AI tried to write. it felt like we were mapping out a subspace of wrong proofs in latent space. looking for a new basis vector… 14/

prof-g profil fotoğrafı
prof-g1 yıl önce

sept 13 ’24 : gpt-o1-preview and gpt-o1-mini drops. i had a stack of prompts ready to roll. i fed the bottleneck conjecture into gpt-o1-preview. it failed miserably, going down the most typical failure mode of proof. >sigh 15/

prof-g profil fotoğrafı
prof-g1 yıl önce

i fed a failed proof from claude-3.5 into gpt-o1-mini and asked for an analysis. it found the mistake and suggested an improvement. the `improved proof’ had an obvious reversed inequality inference error. >sigh 16/

prof-g profil fotoğrafı
prof-g1 yıl önce

i pointed out the reversed inequality. gpt-o1-mini said “To address this, let's revisit and refine the proof...” >> 43 seconds of thought later << an entirely new, clever, correct proof. more elegant than the human proof. 17/

prof-g profil fotoğrafı
prof-g1 yıl önce

this result appears to lie right on the boundary of what is & is not provable via LLMs. i could almost but not quite get it with gpt-4o/claude-3.5/gemini1.5-pro what pushed the frontier was the new gpt-o1-mini model mapping out the failure modes was crucial 18/

prof-g profil fotoğrafı
prof-g1 yıl önce

the rest of sept was spent writing up the preprint… you can find it here: this took more time than I thought it would & indeed, it would have been faster to do it all w/o AI 19/

prof-g profil fotoğrafı
prof-g1 yıl önce

however, the paper is *much* better with the use of AI AI assisted in the initial conjectures, some of the proofs, and most of the applications it was truly a collaborative effort 20/

prof-g profil fotoğrafı
prof-g1 yıl önce

i spoke about the result at the CRM three weeks ago in Barcelona it’s funny… lots of people asked me about how i made the artwork for my slides [AI+c4d] nobody asked me about the AI-generated theorem/proof… 21/

prof-g profil fotoğrafı
prof-g1 yıl önce

you can read for yourself the details of the theorems & applications. most importantly, there is an appendix that documents the process. it’s just the facts – no commentary on what the means for the future. but of course… the future awaits. 22/

prof-g profil fotoğrafı
prof-g1 yıl önce

i went back and forth between outrageous optimism and frustration through this process. i believe that the current models can reason – however you want to interpret that. i also believe that there is a long way to go before we get to true depth of mathematical results. 23/

prof-g profil fotoğrafı
prof-g1 yıl önce

the best way to prepare to see what future models can do [imo] is to explore those corners of math latent space where proofs almost work. such corners exist and are far from the clay prize problems. but there are corners… 24/

Benzer Videolar