LLM: Getting the question right

  • Topic: AI 
  • Thread starter Thread starter .Scott
  • Start date Start date
Join the discussion
Registration is free. Ask a follow-up in this thread, or start your own.
17 replies · 1K views
Science Advisor
Homework Helper
Messages
4,072
Reaction score
2,123
TL;DR
It appears that an internal LLM did not actually answer the $1M Navier-Stoke problem.
I've made a handful of posts to PF where I have used an LLM to assist me with the post - always with attribution (always ChatGPT).
And I have remarked in a couple of cases how I needed to present several questions to the LLM before it finally caught on to what I was asking. In those cases, I have also included the exact wording of the question in my attribution.

Well, it seems we now have a rather spectacular example of that issue.

Our case in point started early this month with this Scientific American article claiming an AI solution to the $1,000,000 Navier-Stokes problem. It said:
For the second time ever, someone has solved one of the seven Millennium Prize Problems—math’s biggest targets, each worth a $1-million prize. But unlike the first time, that someone is an artificial intelligence start-up.
Today OpenAI announced that its internal model has proved that the Navier-Stokes equations, which mathematicians use to study how fluids move, are fatally flawed. The reveal comes after mathematician Tristan Buckmaster alleged that OpenAI had tackled the proof after the company became aware that Buckmaster and his colleague Levent Alpöge had been using a specific method to break a related problem. OpenAI has denied the allegations.
OpenAI’s proof shows that, on rare occasion, the Navier-Stokes equations “blow up,” meaning they dictate that a fluid’s speed becomes infinite at some points, something that is impossible in nature. The company says the proof has been certified using the programming language Lean, which all but guarantees its correctness.

That sounds pretty exciting. But on closer review, we have a new Scientific American article that discusses that earlier announcement. In part:
Two weeks ago OpenAI claimed a solution to one of the biggest open problems in math—the Navier-Stokes problem—an achievement worth a $1-million prize from the Clay Mathematics Institute. The proof ignited a powder keg of concern over artificial intelligence companies’ race to disrupt the subject.
But with the dust still far from settled, a different controversy is emerging: Did OpenAI even solve the right Navier-Stokes problem?
Generated by an internal large language model (LLM), OpenAI’s proof relies on an approach that many experts find unnatural. It solves a variant of the problem that mathematicians say is disconnected from reality and thus less interesting. In a sense, the LLM found and exploited a loophole in the framing of the question.

Don't you hate when that happens?!
 
  • Like
  • Informative
Likes   Reactions: Klystron, WWGD, russ_watters and 2 others
Physics news on Phys.org
Well as with maths we can always generalize the equation to manifold, R^n where n is a natural number and not just n=3 as is the case for Navier-Stokes PDE.
Nobody says mathematics is easy, but with LLM technology it can make the work of scholars easier, where you can concentrate on ideas rather than their implementation itself. The AI would solve the technicalities, but you better check that it does not fail somewhere in its derivations. That's still something mathematicians ought to check for. Yes, it's quite a duanting task to proof-read someone else's work; it really depends how much can you rely on a black box that you don't understand how it works...
 
loop quantum gravity said:
The AI would solve the technicalities, but you better check that it does not fail somewhere in its derivations. That's still something mathematicians ought to check for.

I am no expert, and I came across the concept of "formalization of proof" only very recently, but... There is a process called "formalization" of a proof that renders it into a very rigourous, symbolic, machine-processable form. Once that is done, another algorithm can verify whether the proof is valid.

If I understand correctly, the verification process is much less complex, rather like verifying something that involves a trapdoor function like factorizing large numbers.
 
Swamp Thing said:
I am no expert, and I came across the concept of "formalization of proof" only very recently, but... There is a process called "formalization" of a proof that renders it into a very rigourous, symbolic, machine-processable form. Once that is done, another algorithm can verify whether the proof is valid.

If I understand correctly, the verification process is much less complex, rather like verifying something that involves a trapdoor function like factorizing large numbers.
Yes I know of two such softwares; but that's the point has anyone tried to formalize proofs like of Wiles or of Perleman's, they don't use all those logical connectives or quantifiers, and sometimes they adhere to second order logic, which has its problems regarding a systematic natural deduction system. (As for second order logic I myself beyond wiki entry haven't gone too far in my readings).
But mainly, second order logic means that the quantifiers operate on sets rather than only on variables like in FOL.
I read of those Automated proof checkers, I am not sure they're operate without some mistakes; As I said if we don't understand the code and the algorithms that the BB uses we can't tell if its answer is indeed legit for 100%. It's ok to have skepticism in research, it's better than adhere with the herd without understanding why?
 
  • Like
Likes   Reactions: symbolipoint
@loop quantum gravity Wiles' proof got formalised. It was covered in Dr Samuel Allen Alexander math channel. My understanding is the one only need to understand the lean code after a proof has been formalised and then it runs through a proof checker. is not ready to be released to the masses since there are certain steps in the entire process that cannot be completely automated. Like if there are portion of the written proofs that are not clear, the AI can't tell what the intent behind an incoorect statement. E g. 'By theorem X, we can deduce...' Sonething as simple as that, the AI gets tripped up. Basically every line of a written proof needs to be spelled in sequece of logical steps.
 
elias001 said:
@loop quantum gravity Wiles' proof got formalised. It was covered in Dr Samuel Allen Alexander math channel. My understanding is the one only need to understand the lean code after a proof has been formalised and then it runs through a proof checker. is not ready to be released to the masses since there are certain steps in the entire process that cannot be completely automated. Like if there are portion of the written proofs that are not clear, the AI can't tell what the intent behind an incoorect statement. E g. 'By theorem X, we can deduce...' Sonething as simple as that, the AI gets tripped up. Basically every line of a written proof needs to be spelled in sequece of logical steps.
It first depends on what set theory we are using here.
A proof in a human community by humans doesn't need to go through all this formalization; but in an ideal world you would like to be sure that all the details are given.
As you wrote it's still unfinished.
BTW, what's so intersting in using lean and not coq?
https://en.wikipedia.org/wiki/Rocq
 
@loop quantum gravity If you do a search on why lean was adopted/picked over HOL, Isabelle or coq, you will find what you are looking for. The historical and sociological reasons are very recent, but i don't feel confident or know the technical details well enough to explain to you the technical aspect of the answer. When i looked up what you asked, what Dr Google explain to me has to do with types theory and how that fit into what the professional mathematics comunity needed. It is the types theories in all four informed those choices.
 
  • Like
Likes   Reactions: loop quantum gravity
It has been known since the introduction of LLMs that formulating the question or task correctly and with sufficient detail and guidance is important to obtain a valid answer. In addition, it is also known that taking a long time to arrive at an answer increases the likelihood of an error or confabulation.

Mathematicians are concerned about AI's proofs since they tend to be long and hard to follow. This led me to wonder what the longest human math proof was, since in my experience (as an experimental physicist) they tended to be relatively short.

The longest theorem is "The Classification of Finite Simple Groups": aka the "Enormous Theorem"

https://en.wikipedia.org/wiki/Classification_of_finite_simple_groups
every finite simple group is either cyclic, or alternating, or belongs to a broad infinite class called the groups of Lie type, or else it is one of twenty-six exceptions, called sporadic (the Tits group is sometimes regarded as a sporadic group because it is not strictly a group of Lie type,[3] in which case there would be 27 sporadic groups). The proof consists of tens of thousands of pages in several hundred journal articles written by about 100 authors, published mostly between 1955 and 2004.

Work continued to 2012 with some minor additions.

So it would seem these lengthy AI proofs may not be unreasonable.

Though the NS eq. solution may not have answered the problem's intended objective, we can reword the task and ask Astra to redo its solution. What's ten million or so more dollars to these companies that are valued at a trillion dollars?

This post was facilitated by Gemini.
 
@gleem if the AI companies want their AI systems to write mathematics better, don't worry, they have more computational resources to do that. They have more pressing priorities. Also, the professional math community don't own the rights on who can or cannot solve any open problems. They should stop clutching at pearls.
 
elias001 said:
Also, the professional math community don't own the rights on who can or cannot solve any open problems.
The mathematicians' concern is more about advancing the understanding of the science of math. Proofs should be a means to that end.

"A sentiment echoing throughout the mathematics community right now is that solving problems and generating proofs have always served as proxies for the true goal of mathematicians, which is to further human understanding. When proofs can be generated without that understanding, it undermines their value as a proxy." Terence Tao
https://terrytao.wordpress.com/2026...f-we-need-to-better-celebrate-the-rest-of-it/
 
@gleem actually what you are describing is a fairy tale. The math community wants to gate keep on who is allow to solve open problems. They don't like it that a tool has finally arrived that can do a lot of things that a human can do but do it better without their BS attitude.
 
  • Informative
Likes   Reactions: loop quantum gravity
elias001 said:
@gleem actually what you are describing is a fairy tale.
I'm sure some are fearing for the loss of their livelihood. Certainly, in teaching this is the case. Unless AI generated research is understandable, how can we humans be sure of its validity?

elias001 said:
They don't like it that a tool has finally arrived that can do a lot of things that a human can do but do it better without their BS attitude.
It seems to me that many, including Tao, are embracing AI.
 
gleem said:
I'm sure some are fearing for the loss of their livelihood. Certainly, in teaching this is the case. Unless AI generated research is understandable, how can we humans be sure of its validity?


It seems to me that many, including Tao, are embracing AI.
Why is reading OpenAi's work on NS any different than reading other eccentric mathematicians' work without the aid of AI or more accurate LLMs?
At least with LLM you can follow its derivations as they are more spelled out than in past mathematicians' work.
It's not any different than when Mathematica,Matlab and Maple first arrived in the scene of computations.
 
This is the slippery slope. From my experience LLMs are magical.

As a kid of the 1960s I was thoroughly amazed by my uncles teletype that was connected to a GE timesharing service (Dartmouth?).

It would talk to you in a terse manner. You would LOAD gunner4 and get the response READY and then RUN the program and play an artillery game.

That interaction really woke me up to computers. I learned programming years later after the magic wore off but I always imagined that a computer would someday speak to me like the ones on Star Trek.

That day arrived with advent of ChatGPT and friends. As a search tool, it saves me from browsing web pages filled with ad content and trackers.

As an analytic tool, it summarizes academic papers and provides key insights, a glossary of terms and future directions.

As a spreadsheet tool, it organizes my test results, plots the more interesting data and critiques the data.

As a brainstorming tool, it can suggest research topics, add things to investigate further and even write code for poc demos or for a project debugging error messages and even fixing the code directly.

I look at it as an intelligence telescope allowing me not to get bogged down in coding details.

AI has a dark side in that like tools of the past, we lose the skills learned on the old tools. You can see it in education, we teach arithmetic but then transition the kids to use a calculator. We make it clear to students that limited calculator functionality is allowed on tests.

I imagine there will come a time when an AI tool can be used in exams but with test specific limited capabilities. It might do your arithmetic but not your problem on a math test or do grammar cleanup but not write your essay. The higher you go the more capability is enabled.

With respect to mathematics, I believe it will be their telescope to discover new fields and ideas.

All of us must remember to exercise our brain even as we exercise our body because its easy to let AI do the heavy lifting and for us to sit back and not check what it has done.

Now we must demand that AI have grade based guardrails so that for kids its age restricted and won’t do their homework.

Currently, AI is finding needles in our haystack of knowledge. The needles are disconnected bits of knowledge that when aggregated together produce a new result.

If those bits of knowledge aren’t present the AI won’t find it and that’s where humans fit in. We discover or create knowledge.

In mathematics, when an open problem is solved often a new way of thinking or a new analytical tool is discovered. People use the new strategy with other problems. AI will likely compete here.

I feel mathematicians will pair with AI and the synergy will propel the field forward.

Only time will tell.
 
Last edited:
  • Like
  • Informative
Likes   Reactions: gleem and loop quantum gravity
@jedishrfu there are startups that are specifically combong through existing scientific publications and trying to find new discoveries from them. There seems to be new startups trying to do this or that, where the 'this or that' are tasks where no humans would want to be doing and machine learning and generative AI is perfectly good for. Also AI don't unionize, take sick leave, take political stance and force their employers to cancel contracts because ot doesn't agree with their political ideologies or definition of human rights, etc etc.
 
elias001 said:
@jedishrfu there are startups that are specifically combong through existing scientific publications and trying to find new discoveries from them. There seems to be new startups trying to do this or that, where the 'this or that' are tasks where no humans would want to be doing and machine learning and generative AI is perfectly good for. Also AI don't unionize, take sick leave, take political stance and force their employers to cancel contracts because ot doesn't agree with their political ideologies or definition of human rights, etc etc.
Remind me of myself... I wonder am I AI; well, sometimes I need some break from my vocation.
:oldbiggrin:
 
jedishrfu said:
All of us must remember to exercise our brain even as we exercise our body because its easy to let AI do the heavy lifting and for us to sit back and not check what it has done.

Indeed. There are reasons to welcome AI and to avoid it. As tempting as AI is to do some thinking for you, we must avoid its use for this purpose. Sure, productivity benefits are expected. That is great for companies at least in the short run, and maybe for personal endeavors. But what happens when you do not keep up your skills? They deteriorate. You lose adeptness and finesse in your performance. You forget. You lose interest.

A Microsoft study introduces the term "intuitive rust" to help explain the effect of letting AI to begin to replace one's skills.
AI-as-Amplifier Paradox: AI’s dual role as enhancer and eroder, simultaneously strengthening performance while eroding underlying expertise.
The study looked specifically at oncologists and programmers.

For everyday personal activities, some are using agentic AI to take care of things like scheduling or buying products online. These people are giving up their agency, and isn't that what many worry about with their jobs? Are we headed for the WALL-E world or, worse, to an Ideocratic state?

It doesn't seem reasonable to have everything that we depend on as a black box. Even today AI code and designed circuits are difficult for experts to understand. The current mantra is: if it does what we want it to do, great, let's move on. If things start going sideways, how are we going to effectively respond?
 
I witnessed a similar effect with mainframes. A Honeywell instructor came to the site to teach the systems programmers about the print system. She was the resident print system expert not because she worked on it but because the original programmers left the company. She took it upon herself to learn what the assembler code did and created a course on it.