Sharing AI progress in mathematics
- zone411 - 4577 sekunder sedanA quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).
The highest ranked would be:
| 22 | Hilbert’s tenth problem over ℚ |
| 29 | Unique Games |
| 31 | Anderson-model extended states |
| 37 | Spacetime Penrose inequality |
| 48 | Nonexistence of Landau–Siegel zeros |
| 52 | Baum–Connes |
| 78 | Abundance |
| 80 | Hadwiger |
| 87 | Bose–Einstein condensation |
| 92 | Two-dimensional entanglement area law |
- prideout - 5870 sekunder sedanThis includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.
https://github.com/openai/math/blob/main/preprints/Paired-st...
- schleck8 - 749 sekunder sedanLevent Alpöge (Anthropic mathematician) comment on the significance:
> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.
- NotOscarWilde - 4056 sekunder sedanAs a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...
- xanderlewis - 3344 sekunder sedanAs Kevin Buzzard recently said:
> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.
- enoether - 7054 sekunder sedanUnique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!
[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...
- sebmellen - 7110 sekunder sedanIt’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces
Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...
- gizmodo59 - 7079 sekunder sedanThis is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
- againstapples - 3642 sekunder sedanAs an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
- foota - 4695 sekunder sedanFrom their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
- kingstnap - 5868 sekunder sedanSome of these are interesting ngl.
109. Integer multiplication below n log n
Surprising that this is possible.
158. The Euclidean plane cannot be colored with five colors.
Only 6 and 7 remain!
376. Universal computation in forced Navier–Stokes flows.
Morning coffee proven turing complete
- dekhn - 6190 sekunder sedanI'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
- karahime - 7522 sekunder sedanExtremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
- ravenical - 7491 sekunder sedan
- binlog - 7283 sekunder sedanSo happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
- philipfweiss - 432 sekunder sedanClaude, how important are these results?
> If it holds up, it's the biggest single event in the history of mathematics by a wide margin.
- ed - 7095 sekunder sedanActual results: https://github.com/openai/math/blob/main/overview.pdf
- open592 - 6685 sekunder sedanLet's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
- pavitheran - 6360 sekunder sedanFrom the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
- ks2048 - 6485 sekunder sedanI think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
- lf88 - 1205 sekunder sedanIn some ways, this feels more like an ominous warning about the times to come than something to celebrate.
- sigbottle - 3024 sekunder sedanUnique games conjecture and matmul <= 2.25. What the hell.
- rinconrex - 1950 sekunder sedanThe math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.
- avd201 - 2102 sekunder sedanWow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
- closetheloopdev - 2283 sekunder sedanHopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!
It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!
- TheMrZZ - 2813 sekunder sedanThese results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.
But having so many of them at once? Damn. We really live in the future.
- NegativeLatency - 2483 sekunder sedanWhy should I care?
- xydac - 2131 sekunder sedani wonder what it means for maths researchers, and how it aligns with how they approach math problems.
- curtis-jm - 3902 sekunder sedanYou can read the papers here: https://hub.valency.io/collections/openai-math
- - 6019 sekunder sedan
- dgacmu - 3184 sekunder sedanI find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.
- matapassiones - 1902 sekunder sedanValency has the papers up on Valency Hub
- yewenjie - 5717 sekunder sedanA lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.
That copium didn't last for what, three months?
- aaraujo002 - 6475 sekunder sedanThe Advisory Group states in its recommendations [1]:
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
- lokl - 2055 sekunder sedanDo applied math next.
- digitaltrees - 1771 sekunder sedanGross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
- Catloafdev - 6846 sekunder sedanThis is a pretty hilarious thing to read juxtaposed with AGMAI's requests.
Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"
- jrflo - 4003 sekunder sedanGlad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
- hi__dang - 1352 sekunder sedanMathematics is solved.
- applicative - 3630 sekunder sedanI wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
- - 7671 sekunder sedan
- mi_lk - 5826 sekunder sedanCurious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
- kevinwang - 5888 sekunder sedanwow
- connor11528 - 5083 sekunder sedanwill this make the math for building data centers work?
- tootie - 3324 sekunder sedanSeemingly none are vetted and reviewed yet
- nautilus12 - 4267 sekunder sedanHave any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?
The ones with lean proofs could still be formulated incorrectly
- sandworm101 - 2087 sekunder sedanSo the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?
- k2xl - 7116 sekunder sedanCan someone knowledgeable about the subject outline the most significant portions of the results?
- redox99 - 5305 sekunder sedanThe stochastic parrots have predicted the next token once again.
- applicative - 4249 sekunder sedanWhy didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
- pugfugly - 2388 sekunder sedanholy fucking shit
- mathisfun123 - 6332 sekunder sedanWith so many results in so many different areas no way they even remotely spot checked well enough.
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
- dpweb - 5630 sekunder sedan[dead]
- rafterydj - 7266 sekunder sedanI don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
- senderista - 7654 sekunder sedanGood to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
Nördnytt! 🤓