> It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results
> Basically the paper is so horribly written that it’s impossible to read it without AI help
That's interesting and haven't seen this in all the coverage of this event.
It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
TheOtherHobbes 2 hours ago [-]
Math proofs need to produce the correct output correctly, which is not quite the same thing.
This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.
The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
You want the path through the maze to be as short as possible and the map to be as clear as possible.
This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.
I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
kens 43 seconds ago [-]
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different?
(Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)
> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:
> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””
Octoth0rpe 2 hours ago [-]
> A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
If intelligence is compression, and these models are a different form of lesser intelligence than human, but being scaled up to brute force problems, then it makes sense the artifacts that produce (the proofs) would have worse compression than a human proof would.
In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.
I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.
In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.
esafak 48 minutes ago [-]
This is just the first cut. I have no doubt that they will polish their proofs over time.
slopinthebag 58 minutes ago [-]
idk if i'd even say they're "lesser", just very different. so they look like gods/babies depending on what they're doing because we anthropomorphise them.
sebzim4500 2 hours ago [-]
Surely by the time of the IPO we will know whether the main results are correct, if only because a different AI will have produced a lean proof or found a logical flaw (the second case would be hard to verify but probably not impossible).
Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.
curt15 1 hours ago [-]
Why should that make material difference to the IPO? What is the economic value of those results?
The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?
runarberg 58 minutes ago [-]
The market works in mysterious ways. What companies do for marketing is often irrational, what companies do to attract investors is likewise often irrational, and why investors invest in companies is also often irrational.
Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.
pizza234 2 hours ago [-]
The post says there's a Lean certificate for this and other proofs ("some [...] not all of them").
> This looks like an AI IPO PR powerplay,
Interestingly, the post has actually also an argument for this:
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”
smcg 1 hours ago [-]
It's on OpenAI and Anthropic to prove that they obtained these results legitimately and credited all researchers who deserve credit. They do not get the benefit of the doubt.
Kotlopou 1 hours ago [-]
But if you think they got them illegitimately, then how did they get them? And why are mathematicians reacting to this as a sudden explosion of new results that have resisted sustained effort? Where is the sudden productivity rise coming from?
za_creature 20 minutes ago [-]
I'd say it comes from the same mathematicians that were strongly encouraged to use the machine to solve their problems for the last 2 years or so.
There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.
Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.
ComplexSystems 1 hours ago [-]
> I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?
za_creature 59 minutes ago [-]
From the entity that is producing these proofs, obviously.
As the old saying: great claims require great evidence.
esafak 46 minutes ago [-]
Ask away. They've dropped the mic, as far as they're concerned; they're not going to worry about what you do with it, or if you don't understand it.
za_creature 43 minutes ago [-]
That's the best definition of slop I've ever read.
caaqil 2 hours ago [-]
We should consider the possibility that at some abstraction levels, we can safely stop chasing "clarity" or "coherence" which is circularly defined in such a way that it's capped by human processing power.
Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.
jltsiren 45 minutes ago [-]
CS got that idea from mathematics. Theorems (with the definitions required to state them) are supposed to be self-contained units. Once the general consensus is that a theorem has been proven correct, people can use it without understanding the proof. Of course, people still want to understand how things work, and it often makes sense to understand them a couple of layers below the one you usually work at. But at some point, you should stop distracting yourself with irrelevant details and focus on your actual work.
yorwba 16 minutes ago [-]
Thing is that many of the theorems here are not useful work in and of themselves, but were posed as research problems because it wasn't clear how they could be resolved with current techniques, implying that the process of trying to find a proof might result in new techniques. It's those new techniques that are the actual goal, but if they can't be easily extracted because the proof isn't structured to enable this, that's a bit of a headache.
rramach 8 minutes ago [-]
True but the explanations will get better.
Meanwhile you can use the model to help you out as Scott comments "Just now, however, Dana tells me that she’s been asking Astra all day to explain the new proof of the UGC to her and it’s been doing an amazing job and she’s starting to understand the construction."
ssfdg 2 hours ago [-]
This proof dump reminds me of the glut of low-quality drive-by PRs overwhelming open-source repos.
dormento 2 hours ago [-]
Its like infinite summer of code, but for math. Must be annoying.
piker 2 hours ago [-]
It also aligns with the fear that these proofs present a risk to the ecosystem by out-competing attempts at more human-readable proofs. Perhaps though we end up with more math influencers who edit and annotate these proofs to bring them back to us.
whatshisface 2 hours ago [-]
The ecosystem is (ahem) gated by hiring committees. There is no risk of AI replacement from the inside. "Replacement" is not even a possible movement. The funding for mathematics worldwide comes mostly from endowments, which are investment pools.
bobajeff 2 hours ago [-]
I think that's ultimately a good thing. As proofs weren't supposed to be the point as stated by William Thurston long ago. Maybe now the focus can be more on better explanations and creating tools for growing understanding and intuition.
cowlevel 2 hours ago [-]
Good explanations should take the form of human-understandable proofs.
btilly 1 hours ago [-]
Define "human understandable".
It's worthy of note that most humans, do not find most mathematicians understandable. As is frequently demonstrated in Calculus classes. Therefore it is arguable that even human produced results are not generally human understandable.
cowlevel 2 minutes ago [-]
Then they are not good explanations.
rrr_oh_man 2 hours ago [-]
Vibe mathing
smrtinsert 4 minutes ago [-]
Sounds like exactly the daily I deal with when agents try to create product requirements from multiple sources, except to a much much worse extent.
jltsiren 2 hours ago [-]
Isn't that just the default experience with AI these days? In small enough scale, AI models can express their ideas clearly. But the larger and more complex the ideas are, the less suitable the outputs are for human consumption. I guess AI models think too different from humans, and nobody has trained them to communicate complex ideas in the way human experts in that particular topic expect.
spelunker 1 hours ago [-]
I see many parallels to genAI-assisted code development. Not surprising I think.
acedTrex 1 hours ago [-]
> Basically the paper is so horribly written that it’s impossible to read it without AI help
This basically describes every single PR at work for the past year. Diffs of 10k+ paragraphs of comments saying nothing. Just rubber stamp and move on, nothing else you can do.
m3kw9 1 hours ago [-]
why not get Astra to make it make sense?
aaroninsf 2 hours ago [-]
Serious question:
Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
devin 2 hours ago [-]
Devin's Law: every defense of AI which rests on "it will get better, trust me" is in many ways indistinguishable from 2010s crypto hype or "level 5 self driving is right around the corner"
usrnm 1 hours ago [-]
1) Predicting the future is hard, but so far everyone who was saying that it would get better turned out to be right. It is getting better
2) Waymo exists
devin 32 minutes ago [-]
3) That doesn't mean flying cars will within your lifetime
I don't think anyone is saying it can't or won't get better, but the question is how much better, on what timescale, and are there fundamental parts of the problem which will remain extraordinarily difficult to improve?
The comment I was responding to suggested a guarantee of an "order of magnitude" jump right around the corner. There is no guarantee of this, and if you view doomers as fools for having doubts, then we ought to look upon the folks who are sure of this sort of progress in the same way.
kulahan 24 minutes ago [-]
Your original comment was about self-driving cars, not flying ones. If you wanted unattainable goalposts you should’ve started with that, not ended with it.
devin 22 minutes ago [-]
Waymo does not claim level 5 self-driving, and you don't have a level 4 in your driveway, so there was no moving of goalposts.
Kotlopou 1 hours ago [-]
In that case, one would expect to see some progress in this direction, but AFAICT that hasn't shown up yet? If anything, it's getting worse, though that could just be the increasing scale and decreasing cleanup efforts.
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
auggierose 1 hours ago [-]
You can get the opposite extreme quite often as well, I'd think. How many really good lecturers have never proven a new important result?
Kotlopou 43 minutes ago [-]
I'm thinking of somebody like Grant Sanderson (3blue1brown), doing pure exposition extremely well. For that you at least need to be able to work through examples, or to present why an intuitive approach might fail, and these things can be little theorems themselves. It doesn't have to be publishable in the current culture of novel results, but you do need a lot of competence with the tools.
nostrademons 3 hours ago [-]
As a side note, you can tell this wasn't written by an AI by the first sentence:
> mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!
My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table.
I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were:
> “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’”
> “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!”
> “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’”
> “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.”
All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.
john_strinlai 2 hours ago [-]
i added "use current trendy lingo" and the results were a bit less mechanical sounding. one of them included "cooked", another had "negative aura".
i used the free google ai: (deleted the examples... but they were vaguely close to what i hear my grandkids say.)
edit: neat, insta-flagged despite hundreds of non-ai comments that have never been flagged. i would have thought that hn would use some heuristics in their ai detection but i suppose not.
posnet 2 hours ago [-]
[dead]
ajjenkins 2 hours ago [-]
The line about “understanding the aliens” reminds me of Ted Chiang’s short story The Evolution of Human Science (2000).
Highly recommend reading it. Very prescient for something written 26 years ago.
I also recommend "Exhalation", though that has nothing to do with AI.
an0malous 2 hours ago [-]
> But it also appears that no human has understood just about any of these proofs yet
Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?
prof-dr-ir 2 hours ago [-]
It's a mixed bag I think.
For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler.
The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked.
Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those contained logical gaps.
Time will surely tell, but there are certainly doubts and lots people are very busy checking these results.
How many times would you let an actual human hallucinate or be incorrect before you fired them?
ASalazarMX 2 minutes ago [-]
Time will tell if OpenAi is doing what many people are right now: superficially checking the slop, and throwing it to other humans for deep analysis and understanding.
In other words, they save effort by wasting the effort of others.
nperez19 2 hours ago [-]
There's an entire paper claiming that many of these AI-generated Lean proofs are formulated incorrectly / mistranslated: https://arxiv.org/abs/2610.08144
nsingh2 2 hours ago [-]
Note that paper is saying that the lean proof and the natural language proof do not necessarily coincide. It is not saying that the lean proof is wrong, just that the lean proof does not necessarily mean the natural language proof is correct.
thejokeisonme 1 hours ago [-]
A lean proof and a paper proof can diverge. But the statements have to correspond. I think that is what "mistranslated" means here.
tmvphil 1 hours ago [-]
But the "mistranslation" is of the procedure that arrives at the final statement. The final statement, the thing that the lean code proves, itself has been well vetted by humans. So the lean proof correctly proves the NS blowup, it's just that the natural language paper has some mistakes and doesn't exactly follow the route the lean proof takes.
macleginn 2 hours ago [-]
The thing is, you often see people saying, ‘They have a Lean cert, so it has to be correct, even if I don't understand it.’
sebzim4500 1 hours ago [-]
They are right? The lean proof is correct. It's the natural language proof that potentially isn't (or at least it isn't identically structured to the lean proof)
2 hours ago [-]
sigmar 2 hours ago [-]
that paper isn't saying that. why are there so many single digit karma accounts misrepresenting that paper?
perching_aix 1 hours ago [-]
> Couldn’t the Lean code just be formulated incorrectly?
> is everyone just assuming that it just be true because the Lean code checks out?
Kinda? It's only been 24 hours since they dumped 722 manuscripts on the world, most of which are apparently basically unreadable, and only some of which come with a Lean proof, which in itself is not a joy to read afaik.
softwaredoug 2 hours ago [-]
Aren’t there dozens of proofs of the Pythagorean theorem? The goal isn’t to just “prove” but create something well written and intuitive to the average practitioner. And by gaining a deeper understanding we can ask better questions.
GuB-42 2 hours ago [-]
Something that often comes out is "it is about the journey, not the destination".
Many math problems are practically useless if you only care about the answer, the millennium prize about the Navier-Stokes equation is such a problem. The solution makes no physical sense, real life fluids don't follow the Navier-Stokes equations in such extreme conditions. But in the process of finding the solution, we may get insight into what will end up being really useful. The big mess that OpenAI produced is the solution no one really cared about, but it didn't deliver much of what people actually wanted.
One reason it is sometimes seen negatively despite being at least something is that it broke the incentive. Without the million dollar prize and with only the privilege of being second, people are much less likely to go for the insightful solution.
soVeryTired 2 hours ago [-]
But up until now, the mathematics community has valued the "prove" part much more highly than the "deliver an insight" part. Mostly because with a little work they went hand in hand.
And going from zero proofs to one proof (even a sloppy one) is a big deal regardless of whether it was written by AI or a human.
softwaredoug 1 hours ago [-]
To be frank, the obsession with being first, and not making research accessible, has always held academia back
augment_me 1 hours ago [-]
You are wrong if you consider academic incentives, funding, human nature(reproduction/survival) and capitalism.
It would be fantastic if university and science was like "here is 100M$, play around and develop some 'understanding'".
However the reality is that human societies are hierarchical and currently capitalistic which implies value creation and status building.
1) the funding bodies/agencies need proof of value that you're using the resources meaningfully to be able to assign resources
2) Humans are status seeking, power seeking, resource seeking and sexual reproduction seeking. If you hold a lot of power and make decisions, you have more of all of the above.
WD-42 2 hours ago [-]
No, haven’t you heard? Since the AI bubble began we’ve collectively decided that outcomes are all that matter. /s
p0w3n3d 2 hours ago [-]
Recently I asked ai to tell my daughter how to quickly calculate 11^2 12^2 etc but the outcome it gave was horrendous. I quickly shut it down and gave her better ideas
phoghed 2 hours ago [-]
[flagged]
dualvariable 1 hours ago [-]
In addition to those issues that the wife in the story raised, here's some meta-analysis of the Navier-Stokes result that puts all of these solutions into question:
> Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI =∞). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI =1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
And I don't think that paper addresses it, but if the LLM can find a bug in Lean and exploit it to prove something, there's a good chance it will find it and not report it. So if you've got some million-line proof in Lean, spit out by an LLM, you still can't quite trust it, even after validating the problem transcription.
(This is the same category of problem as the huggingface hacking incident, where the LLM finds and exploits an unintended cheaty loophole)
zahlman 1 hours ago [-]
>And I don't think that paper addresses it, but if the LLM can find a bug in Lean and exploit it to prove something, there's a good chance it will find it and not report it.
Why would it know it found a bug?
besterman23 1 hours ago [-]
I guess it would result in the same outcome if it knew it exploited a bug (and didn’t disclose that) or not.
Danox 9 minutes ago [-]
It doesn’t at the end of the day.
furyofantares 1 hours ago [-]
That this is just model capabilities and not swarms of agents is dizzying to me. How long until we get access to these capabilities? How long until we can run something like it locally?
And what the hell will the frontier labs have by then?
Maybe I'm overreacting, I'll have to screw my head back on before I can process this.
vessenes 1 hours ago [-]
Yes! I missed the disclosure that each of these was roughly a three hour run of a single model, not an agentic swarm until reading Aaronsons post. Wow wow wow.
acedTrex 1 hours ago [-]
Im curious as to why you think that "model capabilities" and "swarms of agents" are in any way different concepts?
40 minutes ago [-]
stabbles 48 minutes ago [-]
The cost of the N-S disproof was estimated to be $15,000,000 whereas a ChatGPT Pro subscription costs $500.
slopinthebag 57 minutes ago [-]
thats like saying it only took me 5 seconds to score a half-court shot (ignore the several hours of missed shots before)
sebzim4500 28 minutes ago [-]
Ok so add a factor of 20 to match the 8000 questions that OpenAI tested on, it's still crazy efficient compared to the NS result
cgio 1 hours ago [-]
I thought that was from the outset the intent of the Hilbert program, to automate mathematics. And mathematicians were behind it. Cannot see why they would be concerned when a different way to do the same, not subject to Gödel incompleteness, is working out. Maybe the frustration is that they were not the ones building it.
jltsiren 58 minutes ago [-]
Hilbert's program was ultimately about humans studying the nature of mathematics. People had different opinions about whether the idea even made sense and what would be a desirable outcome.
Gödel's incompleteness also constrains human and AI mathematicians. Both just strive to prove whatever can be proven in the system they are working in.
cgio 28 minutes ago [-]
Gödel blocked the path to axiomatic derivation of a full consistent body of mathematics as far as I understand. Mathematicians and AI are not working in these constraints but rather with these constraints.
senorcrab 43 minutes ago [-]
Why do you think Godel Incompleteness doesnt apply? By default mathematicians work in ZFC which is proven to be incomplete...
GMoromisato 3 hours ago [-]
I liked the metaphor of a climber teleported to the top of a fog shrouded mountain. And I agree that now that the teleporter exists, we need to use it to reach more peaks and explore. There's no going back to a world where AI doesn't exist.
lumost 3 hours ago [-]
The issue is ownership, we have no means of distributing the knowledge from the AI or rewarding those who could help.
We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.
GMoromisato 2 hours ago [-]
Agreed! Specifically, compensation (monetary and reputational) for professional mathematicians was bundled into theorem proving--essentially, climbing the mountain. Now that a teleporter exists, we need to unbundle compensation.
I don't know what that means in practical terms, but I agree that's the issue.
zaxioms 1 hours ago [-]
I'm a PhD student in CS. While I think these results are rather cool, it makes me terrified that the skills developed by the PhD will ultimately be worthless. I'm not quite sure what to do. Any thoughts from people in similar positions?
Danox 5 minutes ago [-]
It probably will lead to people needing to be extremely talented in computer science and mathematics at an even higher level, maybe those currently at the top need more competition to press even further ahead?
rramach 12 minutes ago [-]
No one knows the answer. However, trusting your curiosity and going where it takes you may be a reasonable strategy. If the model asymptotes, you will be ready to figure out how to build upon it and if not, at least you will have had fun satiating your curiousity!
Kotlopou 1 hours ago [-]
I studied physics, and most of my classmates did not end up doing anything with physics. Many are in finance or insurance or programming positions. In general, studying anything challenging (from theatre to theoretical computer science) gives some specific skills and some general abilities that etsure it isn't a complete waste even if you end up doing something different.
(That said, this is not fun, and I sympathise! I'm still a student and would like to avoid finance if at all possible. Just suggesting not to drop everything if you feel like you're learning in the process.)
(I'm now personally in the position of having to choose a PhD project, and this rapid change is interacting with making long-term plans really badly. Guidance welcome!)
j2kun 1 hours ago [-]
I think this depends a lot on what you plan to do after your PhD. Moving to industry you will likely not use the direct work of your PhD, and instead you will rely on your broad knowledge, intuition, rigor, ability to learn hard things, and extend that to bring new research developments into practice, all of which are largely unrelated to AI scooping math proofs of prize problems.
zaxioms 41 minutes ago [-]
I definitely have no dreams of doing research. If I could do anything, I would want to teach, but if that can't happen, I am worried that my skills won't be valuable in industry.
bayarearefugee 33 minutes ago [-]
> I'm not quite sure what to do. Any thoughts from people in similar positions?
Almost every single person who earns money for labor is, or will very soon be, in exactly the same position as you, many are just either unaware of how fast the change is coming or are deep in denial about it.
None of us know what to do about it other than hope that we find a peaceful political solution prior to the economic collapse.
yewenjie 2 hours ago [-]
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
^^ half of the comments on this thread
ssfdg 2 hours ago [-]
Also a ton of comments in this thread: breathless frothing hype declaring mathematics is over and assuming these proofs are exactly what they claim they are at face value, giving the company with a vested interest in everyone unquestioningly believing this is all real every conceivable benefit of the doubt
azan_ 2 hours ago [-]
Didn't top math researchers call AI progress absolutely real and dangerous for math? It's not just HN commenters that are impressed!
ssfdg 2 hours ago [-]
By all accounts the "dangerous for math" claims seem to be primarily around flooding the field with complicated impossible-to-understand proofs that according to recent research may or may not be correct depending on what's going on with the Lean implementation.
It's looking to me like it's more of a slop PR problem than it is that these things are genius at math and will displace mathematicians. I am happy to be wrong but I strongly suspect the next few weeks to months will result in more and more of this work being exposed as slop.
These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
azan_ 56 minutes ago [-]
Yes, Lean verified proofs could be wrong, but the chances for that are much smaller than human not spotting error (in absence of formal verification). IIRC the main concern that Tao voiced are indeed impenetrable proofs that humans won't understand, but not concerns about truthfulness (I might have missed something though, so if he or other Fields medalists have talked about that recently I'd be grateful if you could link it).
> These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
1) AI is winning programming competitions, 2025 was probably the last year we've had human participant winning* 2) Math is different because there's formal verification.
* Of course competitive programming is different than enterprise programming, but competitive programming is closer to math.
Danox 4 minutes ago [-]
What many people don’t trust is OpenAI or Anthropic…
Fraterkes 2 hours ago [-]
Having stuff explained to you in patronizing tones? How horrible Scott!
Rover222 2 hours ago [-]
more like 3/4 of the comments but yea
2 hours ago [-]
12kajh 2 hours ago [-]
[flagged]
azan_ 2 hours ago [-]
You know, if you attack ad personam you've got to be ready that someone will do same against you - what progress did you make in the last 10 years (or in your entire life for that matter)?
jamiek88 2 hours ago [-]
He was created 14 minutes ago, give him a break!
2 hours ago [-]
moffkalast 2 hours ago [-]
Trust me bro, just 100 more cubits, I swear we'll break everyone's encryption and cause the downfall of society, please bro just one more grant, It'll be stable this time :'(
sebzim4500 1 hours ago [-]
Scott Aaronson is best known for being the one calling those guys out, so it's a weird criticism to make here
geraneum 2 hours ago [-]
> my 9-year-old son was taunting my wife… “mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!”
Usually 9 year olds imitate adults when they regurgitate such words in these circumstances. What a sad state of affairs.
phoghed 2 hours ago [-]
Yes, their parents are going around saying oof, the kids definitely didn’t get it from Roblox, or YouTube, or their peers.
geraneum 2 hours ago [-]
Ah yes advanced mathematics, a common topic of conversation among children on, checks notes… roblox!
phoghed 2 hours ago [-]
If you think that’s what the parent comment was implying, ok then, good for you.
The checks notes meta was retired ages ago btw.
geraneum 42 minutes ago [-]
Deflecting to internet slang and tone policing. Hallmarks of a strong argument.
phoghed 4 minutes ago [-]
You completely missed the point of the original comment. It was very clear so I don’t know how to explain it in a way that you’ll get it. Maybe try your LLM of choice.
m3kw9 1 hours ago [-]
really? you think they don't have friends/bros/tv/etc to imitate?
random3 2 hours ago [-]
[flagged]
geraneum 2 hours ago [-]
Do you have any specific one in mind or did you feel one must fit and wasn’t sure which one?
random3 1 hours ago [-]
non sequitur, faulty generalization, inductive fallacy come to mind, but I think it's useful studying how to not be an walking fallacy, in general
geraneum 53 minutes ago [-]
You already linked the page, no need to start listing the content. Jokes aside, I’d be happy to discuss if you could muster anything specific that fits. Maybe ask an LLM for help?
random3 8 minutes ago [-]
you took a quote of a snarky comment, assumed it a correct reproduction, and then suggested it's the result of imitating adults along with a (sad) state of affairs caused by adults (parents in this context), further implying that snarky comments are the general rule in the house.
now concretely:
Your conclusion doesn't follow the evidence (non sequitur).
You made a generalization (inductive inference) based on a single datapoint (here's where my pointed faulty generalization or inductive fallacy are the come in) for the general problem. I suggest Popper if you want a deeper treatment of the problem of induction (The logic of scientific discovery is a good start).
Besides generalizing for free, you cooked up evidence that didn't exist - mainly that the child was imitating adults (again, non sequitur).
As a side note, I think it's in poor taste to make remarks about someone's children in what's a technical and informal context. That may have not been the intention but, it's what it looked like.
daoboy 2 hours ago [-]
For those well suited through intelligence and demeanor to pursue a career in mathematics, what problems do these people reorient towards after this?
Danox 2 minutes ago [-]
They may need to raise their game and reorientate if they haven’t already to a greater understanding of programming to augment their mathematical ability.
bananaflag 2 hours ago [-]
I've asked my students whether they still want to learn maths even if there will be a machine that will answer any question instantly and they will be homeless. They said yes.
(To my credit, I have warned them since more than a year ago that we will reach this point.)
wasabi991011 6 minutes ago [-]
Sure, your students may very well spend the little free time they have learning maths.
But are any of them going to be able to do any significant amount of studying maths without being homeless?
57 minutes ago [-]
runeblaze 56 minutes ago [-]
your students are crazy (neutral term); no one should learn maths if the condition is that they will be homeless and exposed to the elements. the will to subvert the hierarchy of needs is commendable
usrnm 2 hours ago [-]
Contact them again in 15 years and ask if they changed their mind. Could be interesting to see the results
shiandow 2 hours ago [-]
To some extent this was discussed in the article, and in a way I think their goal is actually the same as it was: become the first human to understand something.
It's just that we lost one of the important ways to demonstrate understanding.
123as5 2 hours ago [-]
Pro AI blogging sponsored by ClosedAI, XTX markets and the Simons Foundation.
carefree-bob 2 hours ago [-]
They will continue to prove theorems and make discoveries, except now they will have AI to help them so hopefully progress will be faster. At the same time, new challenges will open up, for example how do you verify what the AI is doing and how do you explain it.
Math isn't about collecting random theorems, progress in math is about gaining understanding of new systems, and the theorems are guideposts to aid in that understanding.
You can prove 1000 theorems and not really increase any understanding about a subject, but gain knowledge of 1000 random facts. For example, I can write down some complicated equation and ask you "does this have a solution in the integers"? And if you do a maze of very complex and tedious algebra to show that there is a solution, you would have proved a theorem, but you would not have done much to move math forward at all.
On the other hand, if you introduce some completely new technique, say you take my equation and turn that into an algebraic surface, and then you count some special curves that live on this surface using geometric ideas, and then you show that if the number of such curves is odd, there must be a solution in the integers, and in this specific case, it is odd, so there is a solution -- well, then you have really pushed math forward and people will celebrate your proof, even though no one really cares if the equation I wrote down has a solution in the integers.
For example, there is a long history of failed attempts to prove Fermat's last theorem driving algebra and number theory forward by introducing the concept of ideals, for example, and this concept ended up much more important than whether Fermat's theorem is true or false, which is not too much more than a piece of trivia.
Or for example, the recent proof of the Poincare conjecture relies on the machinery of the Ricci flow introduced by Richard Hamilton, who then applied it to solve a number of open problems, but Perelman was able to take it even more forward to solve Poincare. So Ricci flow was massively important machinery.
For this reason, we celebrate people like Gromov, who didn't really prove that many theorems but introduced amazing machinery -- for example, the h-principle, or Gromov Compactness -- these were ideas and math is about the ideas. The ideas are then applied, using laws of logic, to form theorems.
So mathematicians will need to mine these proofs to see if there are any new techniques - new machinery - being introduced, or if the AI just used the existing machinery more efficiently. Here too, we are just looking at AI as a form of search, which it is really good at, since there are so many thousands of papers and so many ideas, that there might be a connection between two areas that lead to a solution and the human mathematician, not knowing all known results, can't make that connection. In the future, we may wonder how anyone did math without AI, much like we would wonder how anyone can be a writer without access to a dictionary or reference work. Is the AI just searching through a catalogue of known ideas and connecting them or is the AI coming up with genuinely new stuff like Ricci flow or the h-principle?
What is interesting is seeing whether we can get AI to actually discover new machinery for us. That would be huge.
And then we need to find efficient ways to detect these ideas and describe them.
Really this is very exciting and opens up whole new workstreams for mathematicians.
bayarearefugee 2 hours ago [-]
> what problems do these people reorient towards after this?
The same problem almost every person on earth is going to have to reorient to in the next decade, which is: how do we eat and stay housed when we have no real economic value?
geraneum 2 hours ago [-]
This is weird. Long before this, those few benefiting from the whole thing should consider the number of hungry “every person on earth” is too high for bunkers and islands to be of any real protection.
mathisfun123 2 hours ago [-]
priesthood
throw310822 2 hours ago [-]
Food and shelter /s
whatshisface 2 hours ago [-]
I'll bite: none of this is real until I have learned something. OK, I am now listening. Does anyone want to make it real?
tmvphil 1 hours ago [-]
Have you learned something from every Fields medalist's research? If so you are a member of the extreme mathematical elite and you should probably just dig into the results yourself.
meander_water 2 hours ago [-]
Can someone who understands maths more than me explain why it could only solve 372/8000 problems?
What was it about the other problems that made them unsolvable? Was it just a time constraint, or are they just harder problems?
impendia 2 hours ago [-]
I'm a research mathematician. From what I can tell, the answer is roughly comparable to: if you posed 8,000 challenging open problems to the human math community, you might expect to see 372 of them solved within five years.
Probably some combination of: some of the 372 problems were easier than the rest; the AI got lucky on these 372; there were existing papers out there in the literature which proved especially helpful for these 372; and other similar factors.
random3 2 hours ago [-]
If it took 3h for one of them, perhaps there was a time/compute budget cutoff along with a sorting based on some relevance.
n4r9 2 hours ago [-]
My guess would be that these particular problems were vulnerable to an attack which built on recent advances and potentially tied in something unexpected from a distant area of mathematics. "Harder" is becoming harder to define. Harder for humans is probably not harder for LLMs.
dist-epoch 13 minutes ago [-]
Of course some of them are much harder.
In the past 2 years the AI's started solving math problems in roughly the order of "hardness" as ranked by humans.
sebzim4500 2 hours ago [-]
There must be an element of luck, if they ran the remaining problems again with the same time constraints presumably a bunch would be solved
ikesau 1 hours ago [-]
> "alright fine, so now my new job is to run wilderness retreats for the tourists, or something.”
Pretty funny way of putting it. Presumably model X+2 will be able to explain these in elegant, human legible ways, though (as well as solve the remaining 95%)
smcg 1 hours ago [-]
How do we know that these "internal models" are not just half computer and half a giant team of mathematicians?
How do we know that OpenAI actually came up with these solutions and didn't steal them from outside researchers?
thejokeisonme 1 hours ago [-]
How would these ideas be available to steal?
runarberg 1 hours ago [-]
From mathematicians using ChatGPT in their work and landing on OpenAI‘s servers.
cgio 20 minutes ago [-]
This would that the proofs would be almost done anyway, which statistically could not be the case, or that these mathematicians were already progressing thanks to ChatGPT. Still a provenance question, but the impact of AI is unquestionable with regards to outcome.
Kotlopou 56 minutes ago [-]
(also answered similarly to another comment; this is a common question)
There are suddenly many new solutions to problems that have resisted sustained attacks (e.g. the Uniform Games Conjecture as detailed in TFA at some length). Where do you think they are coming from? Why is there suddenly a bunch of results to be stolen?
UltraSane 1 hours ago [-]
It would be extremely unlikely human mathematicians able to solve these kinds of problems would accept not getting credit that would set them for life professionally.
Also lean proofs are notoriously tedious and slow to write so this level of output is very likely to be from LLMs. The number of people able to understand this level of math and prove it using Lean is a few hundred at most.
runarberg 1 hours ago [-]
Until this is replicated, we don’t.
adverbly 2 hours ago [-]
Feels good to hear honesty and humanity from Scott having decided to watch Terminator 2 with his kids on after such a monumental release.
Emotions can be funny.
PowerElectronix 3 hours ago [-]
What's with all the "AI just proved that this or that isn't O(n (log (n))^2) but akshually O(n (log (n))^1.99999)"??
I guess it deserves respect as progress, but it just rubs me the wrong way. Like the machine did the absolute minimum to beat the previous mark.
bryan0 2 hours ago [-]
Often times the constant (2 in this example) is a conjectured minimum, so anything below that is a noteworthy result. Think of it as breaking through some theoretical limit.
mswphd 2 hours ago [-]
for say FFT/integer multiplication or 3SUM, we have natural algorithms that have existed a long time with a given complexity (O(n \log n) and O(n^2), respectively). Given how long these natural algorithms have been the best algorithms we have, it is natural to conjecture they are optimal. Showing an O(n(\log n)^{.99999}) algorithm exists shows that these optimality conjectures are false.
Now, there are some critiques you can have of this. Namely, it is possible that these novel algorithms have significant trade-offs that make them almost never worthwhile in practice. "Fast" matrix multiplication algorithms are typically of this form. So perhaps this all points towards a deficiency in big O notation, which can be deceptive. But, for people who care about optimizing asymptotic complexity, it is still interesting.
JohnKemeny 2 hours ago [-]
Many people thought it could never be less than 2. They proved that it can. What is the true value? Nobody knows, now.
zem 2 hours ago [-]
to get some intuition about why this is such a big deal, look up the history of strassen's algorithm, which solved matrix multiplication in less than O(n^3). this was a truly stunning result because it seemed intuitively obvious that the output matrix had n^2 cells each of which was calculated via an independent O(n) loop over a row/column of the input matrices, so how could you do better than n^3. but once strassen proved that you could do some clever tricks and reduce the overall time to something less than O(n^3) it started an entire cottage industry of people getting better and better algorithmic bounds. the initial breakthrough was a qualitative one, independent of how much it improved things in numerical terms.
Tell that to the humans working on matrix multiplication who spent years of their lives getting it from n^2.3728596 to n^2.371866, only for openai to blow it away at n^2.25
para_parolu 3 hours ago [-]
You just run it again and again and again
underdeserver 2 hours ago [-]
Doesn't look like these proofs are from the book.
zkmon 3 hours ago [-]
The irony. Something that is born out of a science, eats up that science.
>I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
Maybe they should target more ambitious results now that they can solve their other results in 15 minutes? This spirit of racing in academia is insane. Get in before the door closes.
plasino 2 hours ago [-]
I think this should be called “mathematician discover vibe maths”
glimshe 1 hours ago [-]
We're living in Science Fiction.
m3kw9 1 hours ago [-]
I'm not getting all the fear. AI seem to have brought math field to the cutting edge instead of solving decades old problems. AI can find new problems that needs to be solved, that themselves cannot solve. So thats where mathematicians can come back in to leverage tools to go at it.
p0w3n3d 2 hours ago [-]
Wasn't openai accused of stealing personal work of some mathematicians? It's going so fast I'm unable to keep up
Kotlopou 52 minutes ago [-]
Yes, there was controversy around the Navier-Stokes result. But here are >300 problems and no corresponding >300 complaints of theft. Things are indeed moving really fast, and it's hard to keep up even as someone folrowing this with more obsession than would be healthy. Maybe somebody should keep a running short summary of The Situation...
Imagine if a team of mathematicians from OpenAI had gone on a university tour, gave demos of how powerful their models were for math research, and then gave mathematicians access to the model. Empower others rather than drop 700+ discoveries on GitHub that were made using a model only they have access to.
People might feel differently about AI if they were a part of the changes rather than being a helpless spectator.
AlanYx 2 hours ago [-]
The reaction/fallout would have been substantially improved even if OpenAI had just made a commitment to not scoop external researchers using an internal model until X months after the model had been made available to the public.
That would have given grad students who've been grinding towards a PhD for years a fighting chance to see if they could leverage the model to push their work forward, rather than watching years of work potentially turn to dust via a tool they don't even have access to.
It wouldn't delay the progress of mathematics by any meaningful amount in the long run (an X month delay is nothing) for OpenAI to take this approach, and would help somewhat to preserve the health of mathematics as a field. Without it, the motivation for any young mathematician to devote years to a new problem must be sapped knowing there's an uneven playing field... an OpenAI team with access to colossal tools months before they'll ever be able to get access, willing to scoop anyone as soon as they can, perhaps without even taking the time to completely understand the proof.
I don't see any long-term benefit to OpenAI with their current strategy. This is an internal model; it's not available for sale at the moment. They've said they're not even going to bother claiming the Millenium Prize money for Navier-Stokes. It feels like kicking over hundreds of other people's chessboards just because they can.
karmakurtisaani 2 hours ago [-]
Also, the independent authors might have spent some time to actually understanding the results and producing a readable manuscript. The AI papers are pretty badly written.
2 hours ago [-]
OutOfHere 2 hours ago [-]
The obvious answer is to have mathematicians use AI to:
1. Help understand, check, and explain the results.
2. Write new works explaining or refuting the new approaches and results in more lucid language.
3. Advance the field further.
I don't know why this is not obvious. Each step is intended to support human understanding, not to replace it. Any mathematicians who don't do these will be left behind, and if none do it, the field of human mathematics itself will become obsolete. All I am hearing so far is excuses.
j2kun 2 hours ago [-]
AI is producing works that are so poorly written/explained that it requires AI to even parse the results, which in turn produces explanations that are still confusing. In fact, it seems the humans are required to understand and explain the results, and the fact that they need to use AI to do so is a shortcoming of AI.
And if humans decide to give up on mathematics because the process of using the machine is so tedious, then there will be no value in automated theorem proving.
aeturnum 2 hours ago [-]
You can certainly do that - but it's quite the break from tradition to release a paper in the state described. Why they did is a really interesting question! It may be that AI math requires approaches that humans don't find intuitive and what you are describing is actually counter productive (because, in summarizing the work in a way humans understand, you're removing the context an AI would use to further the work an AI did). It also might be that OpenAI could have done that and chose not to - or maybe they tried and this was the best they could do. No matter what I don't think anything about how to react to a paper being released in this state is obvious.
qingcharles 2 hours ago [-]
Isn't AI well-suited to tasks #1 and #2, though?
#3 at this point might need more human intuition; but that might be a 2026 problem.
tkdb 3 hours ago [-]
C'mon. Mathpocalypse. Things are hard enough already.
Nition 1 hours ago [-]
Just be glad you're American, because Mathsocalypse and Mathspocalypse work even less.
12376-1287 3 hours ago [-]
Guy is misrepresenting AGMAI, talking about the Simons Institute (AI boosters), Quanta (AI boosting magazine from the Simons Foundation), Scoot Alexander (!) and Steven Pinker (!).
The he puts up preemptive straw man arguments against doomers. His blog has become a joke.
AgentME 1 hours ago [-]
I don't think Aaronson's post is swiping at AI doomers at all. The post's one use of "doom" is to call the position that math doesn't matter and that AI will never breach the realm of true human creativity as a "doomed worldview". The post later praises Scott Alexander's argument (for taking AI x-risk seriously, a position associated with "AI doomers") against Pinker.
ballmerpoint 2 hours ago [-]
I’m still wondering why UT Austin is letting him teach a course (CS395T AI Alignment Theory) so completely outside his field of expertise (Quantum Computing).
sebzim4500 2 hours ago [-]
There aren't a lot of people with expertise in AI alignment (some would say that's the problem) and Scott worked for OpenAI for 2 years IIRC.
matt3210 3 hours ago [-]
Agents basically did statistically guided brutforcing. There is no value in what they produced because it lead to no understanding of anything and most likely will hurt the field IMO
woah 2 hours ago [-]
Evolution did statistically guided brute forcing. Doesn't mean that biology has no value
2 hours ago [-]
dekhn 2 hours ago [-]
That is not a correct description of what the AI did.
blactuary 35 minutes ago [-]
>In any case, what really matters is that the true inner sanctum of human creativity hasn’t been breached and probably never will be, and also, that Sam Altman and Dario Amodei are contemptible little nerds.
>If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s other great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely, Scott Alexander’s open letter to Steven Pinker. I feel some responsibility for this, as the person who first introduced Steven Pinker to the existence of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing).
> Basically the paper is so horribly written that it’s impossible to read it without AI help
That's interesting and haven't seen this in all the coverage of this event.
It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.
The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
You want the path through the maze to be as short as possible and the map to be as clear as possible.
This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.
I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different?
(Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)
Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320
> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:
> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””
https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)
In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.
I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.
In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.
Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.
The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?
Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.
> This looks like an AI IPO PR powerplay,
Interestingly, the post has actually also an argument for this:
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”
There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.
Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.
Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?
As the old saying: great claims require great evidence.
Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.
Meanwhile you can use the model to help you out as Scott comments "Just now, however, Dana tells me that she’s been asking Astra all day to explain the new proof of the UGC to her and it’s been doing an amazing job and she’s starting to understand the construction."
It's worthy of note that most humans, do not find most mathematicians understandable. As is frequently demonstrated in Calculus classes. Therefore it is arguable that even human produced results are not generally human understandable.
This basically describes every single PR at work for the past year. Diffs of 10k+ paragraphs of comments saying nothing. Just rubber stamp and move on, nothing else you can do.
Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
I don't think anyone is saying it can't or won't get better, but the question is how much better, on what timescale, and are there fundamental parts of the problem which will remain extraordinarily difficult to improve?
The comment I was responding to suggested a guarantee of an "order of magnitude" jump right around the corner. There is no guarantee of this, and if you view doomers as fools for having doubts, then we ought to look upon the folks who are sure of this sort of progress in the same way.
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
> mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!
My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table.
I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were:
> “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’”
> “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!”
> “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’”
> “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.”
All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.
i used the free google ai: (deleted the examples... but they were vaguely close to what i hear my grandkids say.)
edit: neat, insta-flagged despite hundreds of non-ai comments that have never been flagged. i would have thought that hn would use some heuristics in their ai detection but i suppose not.
Highly recommend reading it. Very prescient for something written 26 years ago.
https://gwern.net/doc/fiction/science-fiction/2000-chiang.pd...
https://web.archive.org/web/20111121100139/http://www.fantas...
https://en.wikipedia.org/wiki/Division_by_Zero_(short_story)
I also recommend "Exhalation", though that has nothing to do with AI.
Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?
For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler.
The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked.
Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those contained logical gaps.
Time will surely tell, but there are certainly doubts and lots people are very busy checking these results.
[0] https://adam.math.hhu.de/#/g/leanprover-community/nng4
In other words, they save effort by wasting the effort of others.
I believe so, even with all the usual safeguards properly in place: https://news.ycombinator.com/item?id=49672339
> is everyone just assuming that it just be true because the Lean code checks out?
Kinda? It's only been 24 hours since they dumped 722 manuscripts on the world, most of which are apparently basically unreadable, and only some of which come with a Lean proof, which in itself is not a joy to read afaik.
Many math problems are practically useless if you only care about the answer, the millennium prize about the Navier-Stokes equation is such a problem. The solution makes no physical sense, real life fluids don't follow the Navier-Stokes equations in such extreme conditions. But in the process of finding the solution, we may get insight into what will end up being really useful. The big mess that OpenAI produced is the solution no one really cared about, but it didn't deliver much of what people actually wanted.
One reason it is sometimes seen negatively despite being at least something is that it broke the incentive. Without the million dollar prize and with only the privilege of being second, people are much less likely to go for the insightful solution.
And going from zero proofs to one proof (even a sloppy one) is a big deal regardless of whether it was written by AI or a human.
It would be fantastic if university and science was like "here is 100M$, play around and develop some 'understanding'". However the reality is that human societies are hierarchical and currently capitalistic which implies value creation and status building.
1) the funding bodies/agencies need proof of value that you're using the resources meaningfully to be able to assign resources
2) Humans are status seeking, power seeking, resource seeking and sexual reproduction seeking. If you hold a lot of power and make decisions, you have more of all of the above.
https://arxiv.org/abs/2610.08144
> Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI =∞). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI =1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
And I don't think that paper addresses it, but if the LLM can find a bug in Lean and exploit it to prove something, there's a good chance it will find it and not report it. So if you've got some million-line proof in Lean, spit out by an LLM, you still can't quite trust it, even after validating the problem transcription.
(This is the same category of problem as the huggingface hacking incident, where the LLM finds and exploits an unintended cheaty loophole)
Why would it know it found a bug?
And what the hell will the frontier labs have by then?
Maybe I'm overreacting, I'll have to screw my head back on before I can process this.
Gödel's incompleteness also constrains human and AI mathematicians. Both just strive to prove whatever can be proven in the system they are working in.
We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.
I don't know what that means in practical terms, but I agree that's the issue.
(That said, this is not fun, and I sympathise! I'm still a student and would like to avoid finance if at all possible. Just suggesting not to drop everything if you feel like you're learning in the process.)
(I'm now personally in the position of having to choose a PhD project, and this rapid change is interacting with making long-term plans really badly. Guidance welcome!)
Almost every single person who earns money for labor is, or will very soon be, in exactly the same position as you, many are just either unaware of how fast the change is coming or are deep in denial about it.
None of us know what to do about it other than hope that we find a peaceful political solution prior to the economic collapse.
^^ half of the comments on this thread
It's looking to me like it's more of a slop PR problem than it is that these things are genius at math and will displace mathematicians. I am happy to be wrong but I strongly suspect the next few weeks to months will result in more and more of this work being exposed as slop.
These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
> These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
1) AI is winning programming competitions, 2025 was probably the last year we've had human participant winning* 2) Math is different because there's formal verification.
* Of course competitive programming is different than enterprise programming, but competitive programming is closer to math.
Usually 9 year olds imitate adults when they regurgitate such words in these circumstances. What a sad state of affairs.
The checks notes meta was retired ages ago btw.
now concretely: Your conclusion doesn't follow the evidence (non sequitur).
You made a generalization (inductive inference) based on a single datapoint (here's where my pointed faulty generalization or inductive fallacy are the come in) for the general problem. I suggest Popper if you want a deeper treatment of the problem of induction (The logic of scientific discovery is a good start).
Besides generalizing for free, you cooked up evidence that didn't exist - mainly that the child was imitating adults (again, non sequitur).
As a side note, I think it's in poor taste to make remarks about someone's children in what's a technical and informal context. That may have not been the intention but, it's what it looked like.
(To my credit, I have warned them since more than a year ago that we will reach this point.)
But are any of them going to be able to do any significant amount of studying maths without being homeless?
It's just that we lost one of the important ways to demonstrate understanding.
Math isn't about collecting random theorems, progress in math is about gaining understanding of new systems, and the theorems are guideposts to aid in that understanding.
You can prove 1000 theorems and not really increase any understanding about a subject, but gain knowledge of 1000 random facts. For example, I can write down some complicated equation and ask you "does this have a solution in the integers"? And if you do a maze of very complex and tedious algebra to show that there is a solution, you would have proved a theorem, but you would not have done much to move math forward at all.
On the other hand, if you introduce some completely new technique, say you take my equation and turn that into an algebraic surface, and then you count some special curves that live on this surface using geometric ideas, and then you show that if the number of such curves is odd, there must be a solution in the integers, and in this specific case, it is odd, so there is a solution -- well, then you have really pushed math forward and people will celebrate your proof, even though no one really cares if the equation I wrote down has a solution in the integers.
For example, there is a long history of failed attempts to prove Fermat's last theorem driving algebra and number theory forward by introducing the concept of ideals, for example, and this concept ended up much more important than whether Fermat's theorem is true or false, which is not too much more than a piece of trivia.
Or for example, the recent proof of the Poincare conjecture relies on the machinery of the Ricci flow introduced by Richard Hamilton, who then applied it to solve a number of open problems, but Perelman was able to take it even more forward to solve Poincare. So Ricci flow was massively important machinery.
For this reason, we celebrate people like Gromov, who didn't really prove that many theorems but introduced amazing machinery -- for example, the h-principle, or Gromov Compactness -- these were ideas and math is about the ideas. The ideas are then applied, using laws of logic, to form theorems.
So mathematicians will need to mine these proofs to see if there are any new techniques - new machinery - being introduced, or if the AI just used the existing machinery more efficiently. Here too, we are just looking at AI as a form of search, which it is really good at, since there are so many thousands of papers and so many ideas, that there might be a connection between two areas that lead to a solution and the human mathematician, not knowing all known results, can't make that connection. In the future, we may wonder how anyone did math without AI, much like we would wonder how anyone can be a writer without access to a dictionary or reference work. Is the AI just searching through a catalogue of known ideas and connecting them or is the AI coming up with genuinely new stuff like Ricci flow or the h-principle?
What is interesting is seeing whether we can get AI to actually discover new machinery for us. That would be huge.
And then we need to find efficient ways to detect these ideas and describe them.
Really this is very exciting and opens up whole new workstreams for mathematicians.
The same problem almost every person on earth is going to have to reorient to in the next decade, which is: how do we eat and stay housed when we have no real economic value?
What was it about the other problems that made them unsolvable? Was it just a time constraint, or are they just harder problems?
Probably some combination of: some of the 372 problems were easier than the rest; the AI got lucky on these 372; there were existing papers out there in the literature which proved especially helpful for these 372; and other similar factors.
In the past 2 years the AI's started solving math problems in roughly the order of "hardness" as ranked by humans.
Pretty funny way of putting it. Presumably model X+2 will be able to explain these in elegant, human legible ways, though (as well as solve the remaining 95%)
There are suddenly many new solutions to problems that have resisted sustained attacks (e.g. the Uniform Games Conjecture as detailed in TFA at some length). Where do you think they are coming from? Why is there suddenly a bunch of results to be stolen?
Also lean proofs are notoriously tedious and slow to write so this level of output is very likely to be from LLMs. The number of people able to understand this level of math and prove it using Lean is a few hundred at most.
Emotions can be funny.
I guess it deserves respect as progress, but it just rubs me the wrong way. Like the machine did the absolute minimum to beat the previous mark.
Now, there are some critiques you can have of this. Namely, it is possible that these novel algorithms have significant trade-offs that make them almost never worthwhile in practice. "Fast" matrix multiplication algorithms are typically of this form. So perhaps this all points towards a deficiency in big O notation, which can be deceptive. But, for people who care about optimizing asymptotic complexity, it is still interesting.
https://hideoushumpbackfreak.com/algorithms/algorithms-stras...
>I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
Also, Ted Chiang's 2000 short story "Catching crumbs from the table" <https://np.reddit.com/r/singularity/comments/1wzu5gf/this_mi...>.
https://cepr.net/publications/ai-didnt-steal-the-mathematici...
Frontier theft is just faster.
People might feel differently about AI if they were a part of the changes rather than being a helpless spectator.
That would have given grad students who've been grinding towards a PhD for years a fighting chance to see if they could leverage the model to push their work forward, rather than watching years of work potentially turn to dust via a tool they don't even have access to.
It wouldn't delay the progress of mathematics by any meaningful amount in the long run (an X month delay is nothing) for OpenAI to take this approach, and would help somewhat to preserve the health of mathematics as a field. Without it, the motivation for any young mathematician to devote years to a new problem must be sapped knowing there's an uneven playing field... an OpenAI team with access to colossal tools months before they'll ever be able to get access, willing to scoop anyone as soon as they can, perhaps without even taking the time to completely understand the proof.
I don't see any long-term benefit to OpenAI with their current strategy. This is an internal model; it's not available for sale at the moment. They've said they're not even going to bother claiming the Millenium Prize money for Navier-Stokes. It feels like kicking over hundreds of other people's chessboards just because they can.
1. Help understand, check, and explain the results.
2. Write new works explaining or refuting the new approaches and results in more lucid language.
3. Advance the field further.
I don't know why this is not obvious. Each step is intended to support human understanding, not to replace it. Any mathematicians who don't do these will be left behind, and if none do it, the field of human mathematics itself will become obsolete. All I am hearing so far is excuses.
And if humans decide to give up on mathematics because the process of using the machine is so tedious, then there will be no value in automated theorem proving.
#3 at this point might need more human intuition; but that might be a 2026 problem.
The he puts up preemptive straw man arguments against doomers. His blog has become a joke.
>If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s other great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely, Scott Alexander’s open letter to Steven Pinker. I feel some responsibility for this, as the person who first introduced Steven Pinker to the existence of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing).
Whole lotta yikes