AI and Math Research

  • Thread starter Thread starter Hornbein
  • Start date Start date
  • Featured
Join the discussion
Registration is free. Ask a follow-up in this thread, or start your own.
10 replies · 829 views
Hornbein
Gold Member
Messages
4,205
Reaction score
3,351
----
@hhh4804
(As a physicist) i don't know, it feels like a chess player solving a position using a chess engine... i think we are completely missing the meaning of scientific research, there is no point in solving math problems using AI
----

Me
Well it depends. Do you want to solve problems or do you enjoy the thrill of the chase?

As for me I worked in a tiny field that very few were seriously interested in. Over a couple of decades I solved everything I wanted to to my satisfaction. I wrote up the results in friendly form. That was all very nice but now I no longer have anything to ponder in my spare time so I'm bored. Did I win or did I lose?

In my case it was strictly amateur. Suppose I were funding research. Would I want results as quickly and cheaply as possible? Or is my goal to keep gentlemen in endless character-building struggle?
 
Physics news on Phys.org
I thought this news fitted this thread:

Did AI Just Solve Navier-Stokes? What OpenAI's Claim Actually Proves

To work the problem, OpenAI used 10,000 parallel AI agents, 88 hours of compute, and roughly $22.5M in accumulated costs.
Here's where the fine print matters. The official Clay formulation for Navier-Stokes includes four options, labeled A, B, C, and D. Options A and B ask for global smooth solutions with no external forcing, in ordinary three-dimensional space or a periodic domain. Options C and D allow blow-up examples with a smooth external forcing term. OpenAI's proof targets options C and D, meaning the forced variant is genuinely part of the official Clay problem.

That said, many mathematicians consider options A and B to be the deeper question, because forcing is externally imposed: you're choosing the force to cause the blow-up, rather than asking whether the equations can break down on their own. Whether the approach can be extended to remove the forcing and address A and B is not yet clear, and that question is very much still open. The gap between what's been shown and what many experts consider the heart of the problem is the actual takeaway here.
[...] an independent researcher's year of unpublished work became the resource a well-funded lab ran a multi-million-dollar sprint around, [...]
If you aren't familiar with Lean, know it's a formal proof verification system that checks whether each logical step follows from the previous one. A Lean-verified proof can't be "talked into" looking correct: if it passes, the logical chain is sound. What Lean doesn't do is tell you whether the approach is conceptually meaningful, whether it addresses the problem as experts understand it, or whether a human mathematician would recognize it as a genuine solution.

Terence Tao, arguably the most prominent living mathematician, offered a warning about this dynamic on September 5th, before OpenAI's announcement. Writing on Mathstodon in response to rumors circulating at the time, he noted it was a hypothetical concern: "there is a substantial opportunity cost in converting a historically productive and motivating problem such as Navier-Stokes regularity into a mere viral social media post advertising some benchmark progress, rather than actually advancing the field and developing the next generation of both problems to ask, and people to work on them." After the announcement, Tao praised the Buckmaster-Alpöge work as "a remarkable achievement" and noted their arguments had been formalized in Lean. He hasn't publicly endorsed OpenAI's specific claimed proof.
Axios put the uncomfortable question plainly: what happens when the company providing scientists with AI research tools can also mobilize vastly more resources to compete with them? Buckmaster and Alpöge were using OpenAI's own Codex throughout their year of work. The tools a researcher uses to build toward a result can be owned by the same organization that can outpace them with those same tools at 10,000x scale. That's not a conspiracy. It's a structural feature of the current moment, and it's the scenario people have been worried about, independent of whether OpenAI did anything wrong here.
The sprint model produced a result. Whether it produced a result, in the sense mathematicians mean, is a different question, and one that'll take considerably more time to answer.

OpenAi's article: On the Navier–Stokes Millennium Prize Problem
 
Reply
  • Like
Likes   Reactions: javisot
Reply
  • Agree
  • Like
Likes   Reactions: FactChecker and jack action
javisot said:
(I have been trying since yesterday to apply OpenAI's work to gain a deeper understanding of black hole jets)
IMO, this is an important point. It is hoped that a solution (AI or otherwise) of a math/physics problem gives some insight to a new principle that advances our understanding of the subject.
 
Last edited:
Reply
  • Like
Likes   Reactions: javisot
javisot said:
I am amazed at how long it has taken for this matter to be mentioned on PF. A Millennium problem, C-D, has been solved (or so it seems). We should have been talking about this since yesterday.
A discussion of this topic and some controversy around it was started two days ago in this thread:
 
Reply
  • Like
Likes   Reactions: javisot
renormalize said:
A discussion of this topic and some controversy around it was started two days ago in this thread:
Oops, thanks, I didn't see it.
 
Reply
  • Like
Likes   Reactions: jack action and renormalize
maybe not exactly research related, but I just want to suggest one always double check results announced by AI for accuracy.
Here is an example: In some of my (joint) papers there are new methods introduced, not for themselves, but to solve a specific problem. Hence the paper and its title and discussion focus only on the solved problem and not the new method, which is buried in the proofs of the theorems.
Then later other people sometimes introduced the same method and called attention to it as such in their title, i.e. their paper was not about solving a new problem but just about introducing a generalized version of a known method. Subsequently people tend to cite that later work as the origin of the generalized method.
So I asked an AI agent if a certain generalized method was introduced in one of our papers, naming it and dating it. The answer was no, that the method was introduced by someone else several years later. Then I asked again, but made the ask declarative, claiming affirmatively that indeed we had introduced this method in that same paper, this time giving a page number. the answer this time was yes, this was correct, the method was introduced in that paper. This just seems to me like useless sycophantic response, and pretty worthless. ...
I wonder if the AI agent could actually solve the problem we solved, having awareness only of the later introductions to the method? I'll ask.
Ok ,AI cited a later paper, 8 years after ours, for the proof, which does give a more general but different proof. Now I'll ask it to explain how we first did it...

Ok I asked and got back a hallucinatory account, which is totally inaccurate, and a mishmash of other peoples results that do not even imply the desired result and are misunderstood completely by AI. Our new method introduced in that paper is not even mentioned. So it obviously did not read our paper and relied on summaries of related papers that it also did not understand.
It cites a paper by an expert written a year before ours, misunderstanding his notation, a paper which does not settle (or even consider) the matter.

I conclude that if AI gives you something useful, wonderful, but absolutely do not rely on the truth of any claim it makes, unverified. I try to remember, AI does not feel bad when it lies and hallucinates, as there is no one there, even if it uses the pronoun "I".

Added later: the AI agents I used also seem to have no memory. This time I asked the leading version of my question first, i.e. whether an early paper proving something actually also trivially implied a famous corollary, and it correctly said yes indeed, (I think it used the word "demolished") and even laid out the logic impeccably. Then I asked again, without the early reference, who had first proved the famous corollary, and this time it incorrectly gave as answer a reference to a much later paper written by someone who had ignored the implications of the earlier work.
Its memory also seemed to vary as I asked again and again. The next answer mentioned the earlier paper as part of the historical context, but seemed to have forgotten that it had actually settled the whole matter. Later answers seemed to have forgotten the earlier paper entirely.
This AI behavior mirrors exactly what one finds asserted in the literature, since various writers have a varying degree of familiarity with the actual state of affairs historically. So AI apparently just quoted as true whatever statement it found in the source it happened to consult. All answers, no matter how contradictory, were always asserted in a completely dogmatic tone, which I fear may be misleading to the human user.
 
Last edited:
Reply
  • Like
Likes   Reactions: jack action and symbolipoint
From Hornbein,
----
@hhh4804
(As a physicist) i don't know, it feels like a chess player solving a position using a chess engine... i think we are completely missing the meaning of scientific research, there is no point in solving math problems using AI
----
Yes! The right way to think. Investigate, study, practice, learn. Work directly with the material and use only real, not artificial* tools.

*Some expected controversy on that.


Me
Well it depends. Do you want to solve problems or do you enjoy the thrill of the chase?

As for me I worked in a tiny field that very few were seriously interested in. Over a couple of decades I solved everything I wanted to to my satisfaction. I wrote up the results in friendly form. That was all very nice but now I no longer have anything to ponder in my spare time so I'm bored. Did I win or did I lose?

In my case it was strictly amateur. Suppose I were funding research. Would I want results as quickly and cheaply as possible? Or is my goal to keep gentlemen in endless character-building struggle?
You would find more to explore if you look for it, even if later and not sooner. You so far won; at least some. If you are later planning to do funded research, then who gets credit when you produce what you find? You and any colleagues, or the artificial intelligence system?
 
symbolipoint said:
who gets credit when you produce what you find? You and any colleagues, or the artificial intelligence system?
You have a point regarding funding, but, IMO, there is value in knowing what is right and what is wrong, regardless of who gets the credit.
 
FactChecker (2026). Post #9 in thread: AI and Math Research. Physics Forums. https://www.physicsforums.com/posts/7322472/

The value is also in the personal in-the-mind knowledge and skill of knowing what is right and what is wrong. My focus is, what does the person understand and know how to do. Not what does a system or artificial intelligence understand and know how to do.
 
Reply
  • Like
Likes   Reactions: FactChecker
symbolipoint said:
FactChecker (2026). Post #9 in thread: AI and Math Research. Physics Forums. https://www.physicsforums.com/posts/7322472/

The value is also in the personal in-the-mind knowledge and skill of knowing what is right and what is wrong. My focus is, what does the person understand and know how to do. Not what does a system or artificial intelligence understand and know how to do.
Good point. The concerning issue is that it is often so hard to understand why AI comes to the conclusion it does. It is not intellectually satisfying to know the answer without understanding why.