AI and Math Research

  • Topic: AI 
  • Thread starter Thread starter Hornbein
  • Start date Start date
  • Featured
Join the discussion
Registration is free. Ask a follow-up in this thread, or start your own.
41 replies · 4K views
Hornbein
Gold Member
Messages
4,248
Reaction score
3,377
----
@hhh4804
(As a physicist) i don't know, it feels like a chess player solving a position using a chess engine... i think we are completely missing the meaning of scientific research, there is no point in solving math problems using AI
----

Me
Well it depends. Do you want to solve problems or do you enjoy the thrill of the chase?

As for me I worked in a tiny field that very few were seriously interested in. Over a couple of decades I solved everything I wanted to to my satisfaction. I wrote up the results in friendly form. That was all very nice but now I no longer have anything to ponder in my spare time so I'm bored. Did I win or did I lose?

In my case it was strictly amateur. Suppose I were funding research. Would I want results as quickly and cheaply as possible? Or is my goal to keep gentlemen in endless character-building struggle?
 
Physics news on Phys.org
I thought this news fitted this thread:

Did AI Just Solve Navier-Stokes? What OpenAI's Claim Actually Proves

To work the problem, OpenAI used 10,000 parallel AI agents, 88 hours of compute, and roughly $22.5M in accumulated costs.
Here's where the fine print matters. The official Clay formulation for Navier-Stokes includes four options, labeled A, B, C, and D. Options A and B ask for global smooth solutions with no external forcing, in ordinary three-dimensional space or a periodic domain. Options C and D allow blow-up examples with a smooth external forcing term. OpenAI's proof targets options C and D, meaning the forced variant is genuinely part of the official Clay problem.

That said, many mathematicians consider options A and B to be the deeper question, because forcing is externally imposed: you're choosing the force to cause the blow-up, rather than asking whether the equations can break down on their own. Whether the approach can be extended to remove the forcing and address A and B is not yet clear, and that question is very much still open. The gap between what's been shown and what many experts consider the heart of the problem is the actual takeaway here.
[...] an independent researcher's year of unpublished work became the resource a well-funded lab ran a multi-million-dollar sprint around, [...]
If you aren't familiar with Lean, know it's a formal proof verification system that checks whether each logical step follows from the previous one. A Lean-verified proof can't be "talked into" looking correct: if it passes, the logical chain is sound. What Lean doesn't do is tell you whether the approach is conceptually meaningful, whether it addresses the problem as experts understand it, or whether a human mathematician would recognize it as a genuine solution.

Terence Tao, arguably the most prominent living mathematician, offered a warning about this dynamic on September 5th, before OpenAI's announcement. Writing on Mathstodon in response to rumors circulating at the time, he noted it was a hypothetical concern: "there is a substantial opportunity cost in converting a historically productive and motivating problem such as Navier-Stokes regularity into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them." After the announcement, Tao praised the Buckmaster-Alpöge work as "a remarkable achievement" and noted their arguments had been formalized in Lean. He hasn't publicly endorsed OpenAI's specific claimed proof.
Axios put the uncomfortable question plainly: what happens when the company providing scientists with AI research tools can also mobilize vastly more resources to compete with them? Buckmaster and Alpöge were using OpenAI's own Codex throughout their year of work. The tools a researcher uses to build toward a result can be owned by the same organization that can outpace them with those same tools at 10,000x scale. That's not a conspiracy. It's a structural feature of the current moment, and it's the scenario people have been worried about, independent of whether OpenAI did anything wrong here.
The sprint model produced a result. Whether it produced a result, in the sense mathematicians mean, is a different question, and one that'll take considerably more time to answer.

OpenAi's article: On the Navier–Stokes Millennium Prize Problem
 
  • Agree
  • Like
Likes   Reactions: FactChecker and jack action
javisot said:
(I have been trying since yesterday to apply OpenAI's work to gain a deeper understanding of black hole jets)
IMO, this is an important point. It is hoped that a solution (AI or otherwise) of a math/physics problem gives some insight to a new principle that advances our understanding of the subject.
 
Last edited:
javisot said:
I am amazed at how long it has taken for this matter to be mentioned on PF. A Millennium problem, C-D, has been solved (or so it seems). We should have been talking about this since yesterday.
A discussion of this topic and some controversy around it was started two days ago in this thread:
 
renormalize said:
A discussion of this topic and some controversy around it was started two days ago in this thread:
Oops, thanks, I didn't see it.
 
  • Like
Likes   Reactions: jack action and renormalize
maybe not exactly research related, but I just want to suggest one always double check results announced by AI for accuracy.
Here is an example: In some of my (joint) papers there are new methods introduced, not for themselves, but to solve a specific problem. Hence the paper and its title and discussion focus only on the solved problem and not the new method, which is buried in the proofs of the theorems.
Then later other people sometimes introduced the same method and called attention to it as such in their title, i.e. their paper was not about solving a new problem but just about introducing a generalized version of a known method. Subsequently people tend to cite that later work as the origin of the generalized method.
So I asked an AI agent if a certain generalized method was introduced in one of our papers, naming it and dating it. The answer was no, that the method was introduced by someone else several years later. Then I asked again, but made the ask declarative, claiming affirmatively that indeed we had introduced this method in that same paper, this time giving a page number. the answer this time was yes, this was correct, the method was introduced in that paper. This just seems to me like useless sycophantic response, and pretty worthless. ...
I wonder if the AI agent could actually solve the problem we solved, having awareness only of the later introductions to the method? I'll ask.
Ok ,AI cited a later paper, 8 years after ours, for the proof, which does give a more general but different proof. Now I'll ask it to explain how we first did it...

Ok I asked and got back a hallucinatory account, which is totally inaccurate, and a mishmash of other peoples results that do not even imply the desired result and are misunderstood completely by AI. Our new method introduced in that paper is not even mentioned. So it obviously did not read our paper and relied on summaries of related papers that it also did not understand.
It cites a paper by an expert written a year before ours, misunderstanding his notation, a paper which does not settle (or even consider) the matter.

I conclude that if AI gives you something useful, wonderful, but absolutely do not rely on the truth of any claim it makes, unverified. I try to remember, AI does not feel bad when it lies and hallucinates, as there is no one there, even if it uses the pronoun "I".

Added later: the AI agents I used also seem to have no memory. This time I asked the leading version of my question first, i.e. whether an early paper proving something actually also trivially implied a famous corollary, and it correctly said yes indeed, (I think it used the word "demolished") and even laid out the logic impeccably. Then I asked again, without the early reference, who had first proved the famous corollary, and this time it incorrectly gave as answer a reference to a much later paper written by someone who had ignored the implications of the earlier work.
Its memory also seemed to vary as I asked again and again. The next answer mentioned the earlier paper as part of the historical context, but seemed to have forgotten that it had actually settled the whole matter. Later answers seemed to have forgotten the earlier paper entirely.
This AI behavior mirrors exactly what one finds asserted in the literature, since various writers have a varying degree of familiarity with the actual state of affairs historically. So AI apparently just quoted as true whatever statement it found in the source it happened to consult. All answers, no matter how contradictory, were always asserted in a completely dogmatic tone, which I fear may be misleading to the human user.
 
Last edited:
  • Like
Likes   Reactions: jack action and symbolipoint
From Hornbein,
----
@hhh4804
(As a physicist) i don't know, it feels like a chess player solving a position using a chess engine... i think we are completely missing the meaning of scientific research, there is no point in solving math problems using AI
----
Yes! The right way to think. Investigate, study, practice, learn. Work directly with the material and use only real, not artificial* tools.

*Some expected controversy on that.


Me
Well it depends. Do you want to solve problems or do you enjoy the thrill of the chase?

As for me I worked in a tiny field that very few were seriously interested in. Over a couple of decades I solved everything I wanted to to my satisfaction. I wrote up the results in friendly form. That was all very nice but now I no longer have anything to ponder in my spare time so I'm bored. Did I win or did I lose?

In my case it was strictly amateur. Suppose I were funding research. Would I want results as quickly and cheaply as possible? Or is my goal to keep gentlemen in endless character-building struggle?
You would find more to explore if you look for it, even if later and not sooner. You so far won; at least some. If you are later planning to do funded research, then who gets credit when you produce what you find? You and any colleagues, or the artificial intelligence system?
 
symbolipoint said:
who gets credit when you produce what you find? You and any colleagues, or the artificial intelligence system?
You have a point regarding funding, but, IMO, there is value in knowing what is right and what is wrong, regardless of who gets the credit.
 
FactChecker (2026). Post #9 in thread: AI and Math Research. Physics Forums. https://www.physicsforums.com/posts/7322472/

The value is also in the personal in-the-mind knowledge and skill of knowing what is right and what is wrong. My focus is, what does the person understand and know how to do. Not what does a system or artificial intelligence understand and know how to do.
 
symbolipoint said:
FactChecker (2026). Post #9 in thread: AI and Math Research. Physics Forums. https://www.physicsforums.com/posts/7322472/

The value is also in the personal in-the-mind knowledge and skill of knowing what is right and what is wrong. My focus is, what does the person understand and know how to do. Not what does a system or artificial intelligence understand and know how to do.
Good point. The concerning issue is that it is often so hard to understand why AI comes to the conclusion it does. It is not intellectually satisfying to know the answer without understanding why.
 
  • Agree
Likes   Reactions: symbolipoint
Hornbein said:
What did he say?
He begins to say something or seems to say something between two and three quarter minutes and three minutes into the video...., I am still watching.

By the ten minutes part, I had no more patience. I do not know what he says. Also the speaking speed is too fast and I wonder if somebody along the line manipulated the video to make for high speed speaking.
 
Last edited:
Hornbein said:
What did he say?

He talks about AI progress and impacts on mathematics research and humanity. But, a clip from the end is getting highlighted in social media, where he says

It's amazing how willing we are to change everything, without having any idea what's going to happen afterwards. It's extremely non-linear dynamics, ..., so we have to slow down, it's insane, the pace, and there is no reason to be this fast.

 
symbolipoint said:
He begins to say something or seems to say something between two and three quarter minutes and three minutes into the video...., I am still watching.

By the ten minutes part, I had no more patience. I do not know what he says. Also the speaking speed is too fast and I wonder if somebody along the line manipulated the video to make for high speed speaking.
He always talks like that. You may use that gear symbol to reduce the speed.
 
  • Informative
Likes   Reactions: symbolipoint
  • Like
  • Informative
Likes   Reactions: javisot, jack action and symbolipoint
Hornbein said:
----
@hhh4804
(As a physicist) i don't know, it feels like a chess player solving a position using a chess engine... i think we are completely missing the meaning of scientific research, there is no point in solving math problems using AI
----

Me
Well it depends. Do you want to solve problems or do you enjoy the thrill of the chase?

As for me I worked in a tiny field that very few were seriously interested in. Over a couple of decades I solved everything I wanted to to my satisfaction. I wrote up the results in friendly form. That was all very nice but now I no longer have anything to ponder in my spare time so I'm bored. Did I win or did I lose?

In my case it was strictly amateur. Suppose I were funding research. Would I want results as quickly and cheaply as possible? Or is my goal to keep gentlemen in endless character-building struggle?
According to leading mathematicians the value in higher math isn't the results, which are generally of no practical value. The worth of those famous questions is largely in what may be discovered while trying to answer them. Or as I heard Gerhard Ringel say, it gives mathematics teachers something to do. So getting a proof isn't the true goal. I'm told only a hundred people watched the computer Go championship. It was no fun watching machines make moves you don't understand.

I read that learning to play music causes the brain to grow differently. This can be seen during autopsies. So there is a benefit even if no one else likes what you play. Japan places a huge emphasis on music : maybe that's why it's such a nice place. In my case the friends I made through music have been the most lasting.

Will reliance on AI cause degeneration of society? I think television has, so I'd guess yes. Of course nothing is going to stop anything that gives the user a big competitive advantage.
 
  • Like
  • Agree
Likes   Reactions: javisot, Greg Bernhardt and OmCheeto
Hornbein said:
Will reliance on AI cause degeneration of society? I think television has, so I'd guess yes. Of course nothing is going to stop anything that gives the user a big competitive advantage.
Television may have caused the degeneration of society as we knew it, but the societies of now and the future are different. It is always happening.
Our society of "Come in! Bonanza's on!" was not that bad.
 
  • Wow
Likes   Reactions: symbolipoint
Hornbein said:
According to leading mathematicians the value in higher math isn't the results, which are generally of no practical value. The worth of those famous questions is largely in what may be discovered while trying to answer them.
To reinforce the point you are making, we could simply assume right now that the answer to the Riemann hypothesis is "yes" or "no," accept that answer, and carry on doing mathematics. But doing so would be absurd and trivial; what matters is what we had to do to prove it and what mathematical tools we had to create.

If we create a tool capable of solving the Millennium Problems without using new tools... something isn't working.
 
@Greg Bernhardt AI not able write up its reasoning well will only be temporary. it is a minor engineering problem the big tech companies can choose to fix whenever they feel like getting around to it. It will be si.ilat when MS copilot was found not able to do simple arithmetic correctly. Look how far LLMs and Agentic AI, with tens of millions of dollars in tokens, have come. It went from a toddler to being able compete with the best of human mathematicans. i still remember when someone on this forum asked me whether i can trust AI to teach me about homological algebra given it makes mistakes in doing simple arithmetic calculations. I trust the AI since it doesn't give me attitude, don't do any kind of gate keeping and has an infinite amount of patience, available 24/7 to answer my math/science questions. I can't wait when the whole autoformulizations and interactive prove assistants for lean are made to be user friendly. Think about it, students has an AI that can check the correctness of their proofs to exercises questions. That will really threaten math graduate students everywhere.
 
Last edited:
elias001 said:
@Greg Bernhardt AI not able write up its reasoning well will only be temporary. it is a minor engineering problem the big tech companies can choose to fix whenever they feel like getting around to it.
That depends on what AI tools are used. The use of logic manipulation can be listed out, although it might be millions of lines. On the other hand, trained neural networks are notorious for being unable to clarify the reason for results.
 
@FactChecker the models that are used to solve the problems are not accessible to the public. Whatever reasons you list, any of those AI companies has access to more computational resources than anyone can afford in a life time. Also, just because it is not good now at whatever it is you think that needs improvement. Wait a few months to a year and see how much things will improve. Also, the AI companies wants to solve scientific problems, they have lean prove checkers. If the math community keep complaining about this or that, they don't really need the math community for how to improve math writing skils. There are enough math journal articles online and offline that can be used as training data.
 
elias001 said:
@FactChecker the models that are used to solve the problems are not accessible to the public. Whatever reasons you list, any of those AI companies has access to more computational resources than anyone can afford in a life time. Also, just because it is not good now at whatever it is you think that needs improvement. Wait a few months to a year and see how much things will improve.
If they use deep neural networks in any part, you can forget about explaining the results in any satisfying way. This has been a problem for decades.
 
@FactChecker just because they can't do a good job of explaining, it doesn't mean it will remain that way forever. Deep neural network can be represented mathematically and that means it is a math problem and is not an impossible solution. Anyways, whatever it is that can't be done or not being done well, tech companies or even you yourself might one day find a better solutions to the problem.
 
elias001 said:
@FactChecker just because they can't do a good job of explaining, it doesn't mean it will remain that way forever. Deep neural network can be represented mathematically and that means it is a math problem and is not an impossible solution. Anyways, whatever it is that can't be done or not being done well, tech companies or even you yourself might one day find a better solutions to the problem.
First of all, I should admit that I don't know if neural networks or anything like it is relevant to the AI that does mathematical proofs. But your confidence that AI results are explainable seems to be more general than that particular AI application.
In the context of deep neural networks, if you accept "AI found obscure patterns and tendencies in a massive amount of training data." as an acceptable explanation, then you are right. Otherwise, I have my doubts.
 
  • Like
Likes   Reactions: PeroK and symbolipoint
FactChecker said:
If they use deep neural networks in any part, you can forget about explaining the results in any satisfying way. This has been a problem for decades.
In general, a computer system (AI or otherwise) would need additional functionality to explain it to humans.

For example, a chess engine finds the best moves. But, it's very different to explain them to a human.

A mathematical proof is similar. There is the line by line proof, but an explanation of the ideas and strategy behind the proof is very different.
 
  • Informative
  • Like
Likes   Reactions: javisot and symbolipoint
@FactChecker i am not saying i am confident about AI being able to explain complicated results all by itself in a future time. i am saying thayt researchers in the AI community will innovate new methods for AI to explain complicated results in a way that is understable to humans. Open AI or any of the big research AI labs has s lot of computational resources to get their own AI systems with guidance from human researchers can find solution to this problems. Put it this way, they have to find it eventually if AI is going to be a scientist that works along side human assisting them in conducting scientific research. It is just a matter of time. As for mathematics, there are the entire catalogue of past issues of American mathematical monthly, Annals of Mathematics and other journals serving as data sets which a capable AI system can train on for writing mathematics well. It is only a matter of priority if large AI research labs choose to spend time on it.