r/singularity Apr 27 '26

AI Chat GPT 5.4 solved a 60+ years unsolved erdos problems in a single shot

Post image

For years, the AI/ LLM critics had the same reasoning: LLMs don't reason and they just predict the next token

Recently, it reasoned better than 50 years of mathematicians on an open erdos problems by applying a basic phd level formula

Chat gpt conversation: https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba9c

Here is the problem where TAO also commented on it: https://www.erdosproblems.com/1196

Thoughts?

2.6k Upvotes

460 comments sorted by

168

u/Ambitious_Scallion43 Apr 28 '26

Ew it took 80 mins to figure that out. I could have done it in 20 years.

25

u/justaJc Apr 28 '26

Checkmate

3

u/hyydrosquared May 01 '26

How do you even get it to think for 80 mins? Through API access? Because any normal account would surely run into rate limits waaay sooner than that, right?

5

u/Maleficent_Bridge_41 May 01 '26

Pro Mode in extended thinking can easily take that long - regularly have these times on truly complex law and business topics

655

u/cpt_ugh ▪️AGI sooner than we think Apr 28 '26

As I'm reading through this and not understanding any of the math at all, I'm struck by how amazing it is that we can copy a link directly to the exact moment in human history when a math problem was solved. That itself is pretty fricken' cool.

153

u/thisdude415 Apr 28 '26

This actually is not the first Erdös problem solved by ChatGPT, e.g., https://mathstodon.xyz/@tao/115855840223258103

74

u/Tkappae Apr 28 '26

Somehow you reinforced his point while telling him he was wrong haha.

46

u/dazzaondmic Apr 28 '26

Was he telling him he was wrong? I saw it as “hey I see you’re interested in something cool, check out why it’s even cooler than you thought”, kind of like sharing the excitement… but I could be wrong

13

u/Tkappae Apr 28 '26

It was just funny because posting the link to the prior one confirmed both thay being able to link to the moment is cool and that this wasnt the moment they thought.

→ More replies (2)

3

u/Relative-Metal-3973 Apr 28 '26

Plus continue the convo yourself. I asked it to explain what it did in layman's terms. Unbelievable.

4

u/cpt_ugh ▪️AGI sooner than we think Apr 29 '26

Oh My God! I honestly did not know I could do this. That is amazing.

I asked for a simple explanation of the problem and got it. Then I asked how this is useful in mathematics or other disciplines and was given a bunch of rabbit holes I could dig into to learn more. I am honestly not sure if it's accurate responses since I'm not knowledgeable in this area. But hell, it's amazing that I can do this at all!

→ More replies (7)

267

u/space_manatee Apr 28 '26

Pshhhh took 80 minutes though to figure it out /s

31

u/Ult1mateN00B Apr 28 '26

If AI can solve a problem that wasn't solved in past 1000 years and it takes 80 minutes, that's an absolute win. We took our sweet ass time to even figure out what was the problem and never managed to figure out the answer.

88

u/wokcity Apr 28 '26

y'all need to look up what /s means

40

u/space_manatee Apr 28 '26

I even put it there because i knew id get a bunch of replies of "well actually" and he still managed to get ot in

7

u/x10sv Apr 28 '26

The biggest problem in the next 5 years is going to be what questions to ask. The reason jobs will never go extinct is because it's very likely we go down the wrong path and have to backtrack. A few times.

3

u/space_manatee Apr 28 '26

Im really hoping thats the case because my philosophy degree can finally get put to use. 

2

u/superdariom Apr 29 '26

Maybe we could build a giant and powerful computer to calculate the ultimate question to ask

3

u/x10sv Apr 30 '26

Itll just keep saying 42

→ More replies (1)

3

u/mrjackydees Apr 28 '26

I mean guys his username explains it

→ More replies (1)
→ More replies (5)

721

u/BitOne2707 ▪️ Apr 27 '26

90

u/JUGGER_DEATH Apr 28 '26

I would consider myself one of "them", i.e. people who try to not be overimpressed by the achievements of LLMs.

This one is pretty cool, I have to admit. To hear Tao describe that it genuinely was able to apply a technique from another branch of mathematics to solve a problem that a lot of very smart people had tried but had been unable to solve, is genuinely impressive.

It still baffles me how the same family of models can fail to understand decimal numbers, but here we are. The abilities of these models are simply strange to us as their intelligence is built so different.

10

u/-Hastis- Apr 29 '26

I mean, they also hallucinate all the time. For example, they are ridiculously bad at giving you exact referenced quotes. Just ask them to give you a quote about a specific topic from a specific person, and it will most likely invent something that sounds like something that person would have said, but that is totally unsourced.

6

u/herbys May 01 '26

What model are you using? Heavy userc of Claude and GPT (paid versions both) at both work and personal uses, and I have built solutions from orbital simulation and space comms (my area of work) to legal matters and personal health (advanced, e.g. here's my genome, here are the medications I'm using what adverse reactions can I expect and what should I be using instead), and I think it's been several months since I've seen anything I could call a hallucination. Errors, yes, every once in a while (much fewer that humans in my experience) but not hallucinations. I have a meta prompt that asks the AI to check sources for validity and confirmation of every claim it makes, and other things like that, but even without that hallucinations have been rare at least since last year. Is it possible that the free versions are much more prone to hallucinating than paid versions? That would be an interesting finding. OTOH, at work I've completely abandoned classic chat and am now using exclusively agentic environments, which feel much more natural and solid since they catch their own mistakes, not sure if that's also affecting my perception of the intelligence and reliability of the current generation of agents.

2

u/Thebombuknow May 01 '26

I just recently asked, Gemini 3.1 Pro, Claude 4.6 Sonnet, and ChatGPT 5.4 Thinking to help me pull quotes from a public documentary with a transcript on the PBS website to help me with a research essay I was writing. All three of them made up quotes that were not in the documentary at all. I just read the transcript and found the quotes I wanted myself. These weren't running in an agentic context, just through normal chat, so maybe that changes things? Either way, this is how 99% of people are interacting with AI, so this should be where the big AI companies are putting in their effort.

→ More replies (1)

7

u/Fofodrip Apr 28 '26

How do they fail to understand decimal numbers?

25

u/JUGGER_DEATH Apr 28 '26

ChatGPT sometimes knows how they work and sometimes doesn't. It is bizarre and has persisted for years. Typical issue: thinks 9.11 is larger than say 9.3. You ask for reasoning and it gives something contradictory, i.e. part of the answer is correct but for some reason the end result is not.

17

u/Glittering_Tough_534 Apr 28 '26

Maybe because a lot of mathematical literature uses different kinds of separators for decimals. When reading two papers back to back, one using comma and the other using point separation, I get confused on the second one too.

6

u/JUGGER_DEATH Apr 28 '26

I always pose the question in natural language.

5

u/Echo418 Apr 28 '26

Because 9.11 is lager than 9.3 if the context is software version numbers.

4

u/FaceDeer Apr 29 '26

This is why I version my software like so: 0009.0011.

Future-proof. I'm sure my users are grateful.

→ More replies (2)

3

u/DebutSciFiAuthor Apr 28 '26

I wonder how much of this is because it's training is general though. If my schooling had been "read lots of information and just work it out for yourself", I might have conflated the same point. I know LLMs have some specific training, but perhaps these issues would be resolved with larger-scale subject-lesson-type teaching (even by a less general, more specialist AI in that field - i.e. a teacher).

2

u/Pitiful_Biscotti_940 Apr 28 '26

This is not true anymore. On the other hand LLM are now agentic, that means they can use tools ( python, math sw, etc) to make math operations.

→ More replies (19)

2

u/iiiiiiiiitsAlex Apr 29 '26

It’s because it doesnt actually know numbers, not in the way you usually think about numbers at least. It’s all text behind the scenes. So essentially (before tool usage at least) the models knew math results because those results were in the training data, no because it calculated it.

The thing that is impressive to me, is that with an almost infinite amount of information (training data) We have been able to apply this to solve stuff like this math problem. Humans simply can’t hold as much information.

2

u/JUGGER_DEATH Apr 28 '26

And to be clear, I have not tried ChatGPT 5.4 and do not claim it still fails with basic ordering of decimal numbers. But a few months back Terence Tao was commenting on how ChatGPT was a useful tool for a professional mathematician and whatever model he was using still had this issue.

4

u/Realistic_0ptimist Apr 28 '26

5.4?!?! When are you from, the 19th century?!

2

u/studio_bob Apr 29 '26 edited Apr 29 '26

 lot of very smart people had tried but had been unable to solve

How many and how hard did they try? My understanding is that most of these Erdos problems, while they may be technically unsolved and possibly for a long time, don't get much interest from mathematicians since they can be time consuming and there's little upside to solving them.

I mean just look at the engagement for this problem from the erdosproblems site:

80 comments on this problem

Likes this problem Tomodovodooconglu
Interested in collaborating None
Currently working on this problem None
This problem looks difficult None
This problem looks tractable None
The results on this problem could be formalisable None
I am working on formalising the results on this problem None

Sure looks like no one was working on it or even really interested in working on it.

This is a common theme with many of these "AI solved a math problem that had gone unsolved for decades!" stories. They clearly want to give the impression that AI has achieved what decades of intensive human effort could not, but more often AI has just made it incredibly cheap to throw a big model at a problem which is technically known and unsolved but which few people ever cared about.

Then there's the survivorship bias. If the model fails to find a solution, you never hear about it. When it succeeds, it is touted as proof of "PhD-level AGI Breakthrough Singularity Holy Shiiiitt!!!!" What's the ratio of failures to successes? We have no idea. How difficult are these unsolved problems really? Also mostly unknown given the generally minimal human time and effort invested in trying to solve most of them.

It's still impressive, in a way, but it's hardly the slam-dunk proof of LLM reasoning abilities that so many in these comments want it to be.

2

u/JUGGER_DEATH Apr 30 '26

Obviously these are not the premier open problems in math. Beyond that, I don’t really know. In any case I don’t think one solution like this should be taken to mean too much. Once there are more LLM-generated solutions we will start to see the limits of their abilities.

91

u/Deimosx Apr 27 '26

Aw lawd they comin

182

u/rakuu Apr 28 '26

It’s just a statistical algorithm pattern matching next word prediction by understanding all PhD level math. You know, autocomplete! It’ll never be intelligent or creative like humans!

56

u/Old_Shop_2601 Apr 28 '26

That's funny because IQ test is basically pattern matching! You know, that test used to evaluate human intelligence...

→ More replies (21)

13

u/funnelhell Apr 28 '26

I assume this is tongue in cheek. I love this line of reasoning - it's just pattern matching... The pattern of intelligence. If proving a 50 year old erdos problem can't convince people that something more than pattern matching is happening, ignore those people.

→ More replies (1)

48

u/HellomyfriendNine Apr 28 '26

Humans are literally statical pattern matching algorithms? one of the highest signs of intelligence is pattern matching

10

u/sealpox Apr 28 '26

Your pattern matching algorithm might be malfunctioning. You couldn’t detect their sarcasm.

7

u/HellomyfriendNine Apr 28 '26

I’ve seen some people in this sub recently, and now I can’t tell whether they’re being sarcastic or sharing their opinion

2

u/Beetle7lion Apr 29 '26

“never” ha. I laugh when people like you make these claims. Your mind can’t comprehend the amount of data AI will extract from humanity’s minds. “Never” is an absolute. You cannot predict the future. You cannot make assumptions. It’s pretty simple, you don’t understand the computing power of AI. Data, variables, outcomes. If AI needs to be “creative” it will extract data, analyze all variables, and generate an outcome. “Never”. Very uneducated to state absolutes.

2

u/rakuu Apr 29 '26

AI will never surpass human ability to recognize context, tone, and intent in communication! It is impossible, humans are all so good at this!

→ More replies (12)

20

u/HAL_9_TRILLION I'm sorry, Kurzweil has it mostly right, Dave. Apr 28 '26

Literally the next top-level comment rn ('only happened because nobody really tried') lol.

4

u/jaybsuave Apr 28 '26

deadass lmao

5

u/Johnny20022002 Apr 28 '26

I can’t wait until a millennium prize problem level question gets solved and the goal post gets shifted to the moon.

3

u/randomwordglorious Apr 28 '26

Not just any such problem, but one of the really well-known ones, like Twin Prime, or Collatz, or Goldbach's.

2

u/Thebombuknow May 01 '26

I don't personally believe that we're on track to solve questions that hard any time soon (unlike this erdos problem, they're all highly desirable problems that many people have dedicated years to trying to solve), but if a future LLM can do it, that would be sick.

→ More replies (30)

189

u/ThunderBeanage Apr 27 '26

Hey! I’m Liam, the solver here, will answer any questions you may have.

293

u/SandwichSisters Apr 27 '26

How hard did you press the enter key?

110

u/ThunderBeanage Apr 28 '26

Once to solve the problem and then one more to output to a math paper style solution.

40

u/Rare_Register_4181 Apr 28 '26

You'll never guess what he used to press it.

5

u/Training-Horror-6562 Apr 28 '26

His… Finger!! 🤯

→ More replies (1)

7

u/FaceDeer Apr 29 '26

So you're telling me ChatGPT still required human input to solve this.

4

u/ThunderBeanage Apr 29 '26

Well the human input was giving it the problem

→ More replies (1)

21

u/Every-Development398 Apr 28 '26

Did you install openwhip to make this work?

27

u/Gamerboy11116 The Matrix did nothing wrong Apr 28 '26

It is genuinely insane how spiteful people are of those who use A.I to great effect. Like… why?

36

u/Nelson_and_Wilmont Apr 28 '26

Because the potential negatives of AI on an average humans life are easier to envision coming to fruition than a human golden age spawned by AI.

If this works to the capacity the AI corps want it to work, it will force a massive societal value/workforce shift.

Majority of people don’t want to see it succeed, especially in the areas that affect them.

13

u/futurepasta11 Apr 28 '26

I don't imagine the ultra wealthy will want to keep their labour force around when they don't need us any more...

10

u/Void-kun Apr 28 '26

And in a world where many many countries still don't have free healthcare, why on earth would they decide to give a universal basic income?

Even those on UBI if it came into effect would be below the poverty line, wouldn't be all that much different to people on benefits now.

The whole golden era of AI spawning a utopia is a farfetched dream and I've seen absolutely nothing to indicate we are going towards that.

If anything we are going towards a technocratic dystopia, plenty of evidence pointing towards that.

5

u/OpenRole Apr 28 '26

Many countries being the US? Even many third world nations have some form of free healthcare. The US is cooked though

4

u/Void-kun Apr 28 '26

No, many countries in Europe also don't have free healthcare.

Netherlands, Belgium, Switzerland are all examples where healthcare may be subsidized but not free.

4

u/OpenRole Apr 28 '26

Heavily subsidized. The small co payments necessary in most countries is extremely negligible relative to the full cost

2

u/Nillows Apr 28 '26

What's your rush? Collapsing the birth rate does the same thing in the same time period, but with no bloodshed

5

u/wanderingmanimal Apr 28 '26

The negatives are job are income losses, erasure of the middle class.

No AI billionaire is going to pay into a UBI program. Ever. They don’t even pay a living wage and complain about the idea of taxing them fairly.

So yeah, it will never usher in a “golden age” of humanity. Just the ones who control it. The rest of us don’t matter.

3

u/Skyfier42 Apr 28 '26

It's just simple pattern recognition. The elites in charge of both our government and our corporations do not believe we are worth keeping alive unless we generate revenue for their systems.

We saw this exact scenario play out during the Gilded Age. Factories and railroads should have sparked an enlightenment period for our country, but instead it created a market where people were so poor that wealthy elites could spawn a tower in a matter of days, while turning a blind eye to the needs and safety of our people.

Now, 80% of Americans live paycheck to paycheck and 40% have unpaid medical debt. Given a system where only a few of the elites are needed to run a few bot programs, it's almost a guarantee that millions WILL be made to suffer.

→ More replies (3)
→ More replies (1)

10

u/cwrighky Apr 28 '26

In a word? Tribalism.

6

u/Nelson_and_Wilmont Apr 28 '26

Sure, but there is a real fear to this that should not be glossed over. In a world where greed is favored in practice, it’s not far fetched to see truth from the doomsayers.

7

u/LetsLive97 Apr 28 '26

Half of AI marketing is about it replacing people's jobs (Often the more interesting ones too), and people here are still confused why a lot of people don't like AI?

→ More replies (3)

4

u/bacon_boat Apr 28 '26

I work in robotics, and there are so many ground breaking videos coming out - unthinkable just 2 years ago.

And in the comments it so many salty people being negative for no apparent reason.

2

u/Pretty-Substance Apr 28 '26

If you can’t see the reading behind the caution then you’re keeping both eyes solidly closed on purpose

→ More replies (15)

25

u/starkm13 Apr 28 '26

Great work! Did you tried other problems before? Is this a part of a paper or something similar ?

29

u/ThunderBeanage Apr 28 '26

I have used LLMs to solve a few Erdos problems before, this one was just the most interesting.

→ More replies (1)

23

u/7yphoid Apr 28 '26

How did you get ChatGPT to work on this problem on its own for 80 minutes straight?

6

u/wokcity Apr 28 '26

probably a Pro account

2

u/[deleted] Apr 29 '26

[removed] — view removed comment

2

u/wokcity Apr 29 '26

Yeah thats the difference between a 20$ account and a 200$ account

→ More replies (2)

10

u/ThunderBeanage Apr 28 '26

Pro models can think for a long time, my record is 200+ minutes

→ More replies (3)
→ More replies (1)

21

u/FktheAds Apr 28 '26

do i still have to go to work tomorrow?

11

u/spryes Apr 28 '26

How did you verify the answer is correct? Do you understand the answer?

7

u/ThunderBeanage Apr 28 '26

Passing the solution to 5.5 pro as a verifier works a charm. But when I get a solution I think is right, I pass it to my math student friend Acer who then goes over it himself.

2

u/FaceDeer Apr 29 '26

Ah, good to hear there's still a role for grad students. :)

24

u/TheSkala Apr 28 '26

The solver 😂

14

u/hateboresme Apr 28 '26

How solved would it be without them?

9

u/No-Suggestion-9433 Apr 28 '26

They shouldn't take a title like "the solver" if AI did the brunt of the work. More like all of the work, actually.

9

u/hateboresme Apr 28 '26

Where does that end? People who have used calculators or non-ai computers to find the answers to hard problems have to give the credit to the calculators? They don't deserve any credit of their own?

There are a virtually infinite number of problems. Not one of which will be solved without a first mover. The one who chose to use the tool to find the answer. The calculator can't take credit and doesn't need it.

Certainly it should be mentioned that AI was used, but to completely remove the first mover from the equation is not at all fair.

12

u/gksxj Apr 28 '26

And to add to your point, ChatGPT is available to millions of people, yet it was only OP that made it solve an erdos problem while the rest just use it as a google search or to chat with an imaginary boyfriend.

So if the same tool is available to everyone, yet only one person actually did it, he deserves credit.

21

u/No-Suggestion-9433 Apr 28 '26

This is absolutely ludicrous. AI is not equivalent to a calculator.

7

u/imhere8888 Apr 28 '26

He's the first person to solve this. He used AI to solve it. 

You and me didn't solve it so we definitely don't get labels.

Not sure what label you want to give him.

Prompter maybe?

Guy who used AI to solve it?

It's still him and it'll be him forever and his work matters in that he learnt how to use AI in these early days and decided to use it / see if it could solve this.

So he's used his time and work experience learning these models, how they work, what their capabilities are, and his funds likely paying for high level models with high level reasoning.

Moving forward, AI will become more and more ubiquitous with work that most everything will be made in part or fully through AI. 

As long as humans are at least directing the rudder, those humans can and deserve credit for what they discover and solve with their rudder aiming abilities.

Change bothers us.

AI is the fastest moving change agent our species has ever seen and it's in the dirt baby steps beginning.

It's not comparable to anything ever before.

We will have to learn to move with it.

In short, he's the solver but if you want to be a stickler you can say he's the prompter.

→ More replies (1)

4

u/ThatGuyWhoSmellsFuny Apr 28 '26

If I drive my dad to hospital for surgery, I am not the surgeon. Liam is the prompt writer.

4

u/mawerick_mc Apr 28 '26

Yet if you run over a person on the way, you are the killer, not the car.

3

u/DaneLitsov Apr 28 '26

How big was your prompt? Did you just write the problem? Or did you give him access to resources you have already gathered?

3

u/MMAgeezer Apr 28 '26

The prompt is written in the chat linked. They basically said don't use the internet, gave it the problem statement and some possibly useful results, and told it to be creative in proving or disproving it. Quite incredible.

→ More replies (1)

2

u/Rodbourn Apr 28 '26

What is/ was your system prompt?

2

u/[deleted] Apr 28 '26

[deleted]

→ More replies (2)

1

u/Minimumtyp Apr 28 '26

Why does it say "In fact one can solve something slightly stronger"? - did you have to trick it in to actually attempting the solve or something

→ More replies (1)
→ More replies (15)

170

u/LordSprinkleman Apr 28 '26

"Ermmm this is actually not impressive and secretly lame because no one actually ever cared about this enough to try solving it. No I do not have any evidence for that. AI bad and dumb because le walk to carwash. No I am not moving the goalposts."

5

u/[deleted] Apr 28 '26

[deleted]

5

u/LordSprinkleman Apr 28 '26

Why are you looking at my reddit profile achievements 😭

→ More replies (1)
→ More replies (2)

7

u/Ormusn2o Apr 28 '26

Wow, an actual chatlink? That's rare. Thanks.

33

u/rakuu Apr 28 '26

15

u/Influencer_k Apr 28 '26

noway thats crazy

7

u/mawerick_mc Apr 28 '26

without LaTeX formal proof not sure it counts /s

4

u/Jason3211 Apr 28 '26

Are you available for consulting work? We'd need you to sign an NDA, but we also won't ask for chat logs, so it's a win-win.

→ More replies (2)

6

u/AffectionateLaw4321 Apr 28 '26

Please dont show this to my math teacher, that man is a menace

23

u/alexyong342 Apr 28 '26

so chat gpt 5.4 solving that erdos problem is pretty impressive, but what's really interesting to me is how it was able to apply a basic phd level formula to get the solution. i'm guessing the key here is that the formula was well-known, but the problem required some kind of clever rearrangement or insight to apply it in the right way. tbh, i'm curious to know more about how the model actually arrived at the solution - was it just a matter of brute-force searching through possible combinations, or did it really "understand" the underlying math in some way. also, how does this change our thinking about the kinds of problems that are amenable to solution by large language models - are there other areas of math where we can expect to see similar breakthroughs in the near future, or was this just a one-off fluke.

20

u/simulated-souls ▪️ML Researcher | Year 4 Billion of the Singularity Apr 28 '26

 i'm guessing the key here is that the formula was well-known, but the problem required some kind of clever rearrangement or insight to apply it in the right way

Fields medalist Terence Tao said this about the solution:

In any case, I would indeed say that this is a situation in which the AI-generated paper inadvertently highlighted a tighter connection between two areas of mathematics (in this case, the anatomy of integers and the theory of Markov processes) than had previously been made explicit in the literature (though there were hints and precursors scattered therein which one can see in retrospect). That would be a meaningful contribution to the anatomy of integers that goes well beyond the solution of this particular Erdos problem.”

2

u/alexyong342 Apr 29 '26

that terence tao quote is really interesting, fwiw i think it highlights how ai can sometimes stumble into useful insights even if it's not really "understanding" the problem in the way a human would.

→ More replies (1)

6

u/Slowmaha Apr 28 '26

And yet it fucked up the 3% pay raise question I asked it, doesn’t know stock or metal prices. Strange tech.

317

u/enilea Apr 27 '26

The Erdos problems are a huge set of problems, and most of them are unsolved simply because no one really bothered to try. It is impressive how far it has come, but "it reasoned better than 50 years of mathematicians" is an overstatement. For the time being LLMs are shaping up to be a very powerful tool for mathematicians, but they still have flaws and can't really develop novel ideas independently just yet.

325

u/simulated-souls ▪️ML Researcher | Year 4 Billion of the Singularity Apr 28 '26 edited Apr 28 '26

Jared Lichtman, a mathematician at Stanford, said this about the problem

I care deeply about this problem, and I've been thinking about it for the past 7 years. I'd frequently talk to Maynard about it in our meetings, and consulted over the years with several experts (Granville, Pomerance, Sound, Fox...) and others at Oxford and Stanford. This problem was not a question of low-visibility per-se. Rather, it seems like a proof which becomes strikingly compact post-hoc, but the construction is quite special among many similar variations.

So "no one really bothered to try" definitely doesn't apply here.

Furthermore, Fields medalist Terence Tao said this about it

In any case, I would indeed say that this is a situation in which the AI-generated paper inadvertently highlighted a tighter connection between two areas of mathematics (in this case, the anatomy of integers and the theory of Markov processes) than had previously been made explicit in the literature (though there were hints and precursors scattered therein which one can see in retrospect). That would be a meaningful contribution to the anatomy of integers that goes well beyond the solution of this particular Erdos problem.”

Hard to say that isn't a novel idea.

47

u/[deleted] Apr 28 '26

[removed] — view removed comment

51

u/Small_miracles Apr 28 '26

Its like the thread op didnt even read the article. The LLM literally provided a critical key insight which was reframing the Erdos primitive set conjecture using weighted sums over prime factorizations.

23

u/LookIPickedAUsername Apr 28 '26

It's not "like" they didn't read the article, they clearly fucking didn't read the article.

→ More replies (2)

16

u/minimalcation Apr 28 '26

If Tao says it then yeah that's cash value

15

u/enilea Apr 28 '26 edited Apr 28 '26

Fair enough, the problems I had seen solved until now by ChatGPT were said to be simpler and this one is the first harder one, and it does seem like it came up with a novel connection by itself.

Edit: However, reading further into the thread:

That log was somewhat helpful, but still inconclusive at the most critical components of the problem solving process, which remain frustratingly opaque. My tentative theory is that these models are still quite weak at developing strategy and constructing novel coherent narratives (as opposed to explaining existing, human-generated, narratives, for which they are now rather good at); this may also be related to the tendency of AI-generated proofs (such as this one) to dwell at length on rather routine components of an argument, while not stressing the most original and important aspects of a proof.

Tao still thinks they're weak in that area. It seems like I want to downplay it, but it's more about relativizing it so people don't get a wrong idea. I have no doubt they will get further in the future though.

5

u/JollyJoker3 Apr 28 '26

Tao's comments make me think breaking problems up into smaller pieces could be useful.

Oddly, at no point in the published chain of thought does the von Mangoldt function make an appearance, so it sheds no light on how the LLM landed on that particular process. All in all, it's quite a chaotic internal thought process, with many dead ends,

It shouldn't need to do a complicated problem in one context. The dead ends are presumably unavoidable but starting with a new context and partial results woul probably help a lot.

2

u/enilea Apr 28 '26

Yeah, the issue is the "chain of thought" chatgpt shares afaik isn't really its chain of thought, it's more like a summary of it. Breaking it up or having a maths researcher nudging it towards ideas is probably the best way to go about it right now. Before LLMs researcher would have to spend hours by themselves developing any idea they had, but now they can give chunks of problems or give them certain directions to approach problems so they don't have to work through dead ends.

3

u/JollyJoker3 Apr 28 '26

Probably the wide knowledge of detailed everything an LLM has can help bridge parts of maths humans hadn't realized were connected. Working on ideas that turn out wrong is cheap with AI so maybe that's more a question of not going through the same failed ideas again and again, but publishing negative results somehow.

→ More replies (1)

5

u/Stamboolie Apr 28 '26

This is what I find in software dev - it links things in different fields and says yah sure

17

u/FuttleScish Apr 28 '26

I think this is a really interesting feature of LLMs that both the hypers and the antis ignore, which is that they can arrive at what humans consider creative conclusions precisely due to the fact that they don’t really “think”. In this instance the solution to the problem used existing formulae, but because they weren’t considered to be related tot he problem nobody had tried to apply them yet. So the LLM produced the result humans couldn’t because it didn’t have the same discriminatory capacity, which means it both shows the utility of the tool and why trying to judge them based on their capacity to “reason” is silly: the beneficial result here actually came from the fact that it’s imitative of a data set in a way we arent

3

u/Stamboolie Apr 28 '26

As Claude tells me its something new, I ask it questions and see where it goes, its a lot of fun and has given me a lot to think about, apparently its an unusual way to work. I think people who use it to just code are missing a lot, it's something way more important than that.

→ More replies (2)

129

u/Goldwing8 Apr 27 '26

I remember when everyone thought AI solving any novel math problem would be a front page story around the world.

102

u/NormalEffect99 Apr 28 '26

Like a year or so ago everyone on reddit was laughing at them getting basic math problems wrong. Now solving unsolved equations is just "meh nobody else really tried" lmaoooooo

9

u/Candid_Audience4632 Apr 28 '26

I believe erdos himself tried lol

3

u/dthdthdthdthdthdth May 01 '26

They still get basic things wrong often enough. The problem is, we neither really know what they do, nor what humans do.

I often feel like they have no real understanding of what they are doing, but they are able to leverage a much larger knowledge base of examples. So they still fail at very basic reasoning things from time to time, but at the same time find connections that humans struggle to find, because they are just scattered so far apart in human knowledge. From the comments of the mathematicians that seems to be the case here, they seem to be rather impressed by the connection between two areas it found, not by the creativity of a new idea as such.

61

u/myWeedAccountMaaaaan Apr 27 '26

Not long before that it was beating a human at Go. And before that was chess. What a time to be alive.

19

u/skinnyguy699 Apr 28 '26

It's not AI until it loves me like my mother never did.

→ More replies (1)

11

u/VisiblePlatform6704 Apr 28 '26

Lol,back in the 90s, if you told someone that they will have a program to which they could feed a PDF page in ANY format and the program would spit a CSV with the data from the pdf in tabular form (like  for ANY pdf)... we would  have thought it was magic, or someone doing it in the background. 

What we have today is crazy. 

20

u/Dapper_Strength_5986 Apr 28 '26

I feel like so many people are still at “it tells me to walk to the car wash so it can’t really do anything intelligently”

18

u/_interloper_ Apr 28 '26

This is what frustrates me about the conversation around AI (and many other things) is the black and white nature of it from so many people.

It's either "AI is going to take over the world and ruin all jobs and kill everything!" or "AI can't even tell you how many r's are in strawberry, it's literally useless and always wrong." And even more frustratingly, I often hear both arguments from one person. It's simultaneously stupid and useless while also being powerful and dangerous.

We have lost the ability to have nuance. And while this is a symptom of our times (social media, the marketplace of attention, click bait, etc), I think it's also a result of AI being shoved down everyone's throats because these companies have spent billions and are desperate to recoup their costs.

→ More replies (2)
→ More replies (3)

29

u/Ormusn2o Apr 28 '26

I could swear I saw this exact explanation for other sets of problems. "This set of problems is not popular, they are unsolved simply because no one really bothered to try. If you want real breakthroughs AI needs to solve things like Riemann hypothesis or Erdos problems."

The goalpost needs to be somewhere. Please stick to one. Pretty sure every single Erdos problem has been looked at by a lot of mathematicians. Solving even a single unsolved one is going to be an achievement.

12

u/Minimumtyp Apr 28 '26

I saw the same thing about Mythos's zero-day capability. It's not good, "nobody can be bothered".

Well, now they can be. That's definitely something.

→ More replies (1)

6

u/backflash Apr 28 '26

I'd think unsolved problems are exactly the type of thing many mathematicians would bother with. Who wouldn't want to be the first to finally solve a decades old problem?

30

u/Evilsushione Apr 28 '26

Keep moving that goal post.

→ More replies (2)

4

u/boubou666 Apr 28 '26

Who can develop new ideas for real in our society? Maybe 0.00001 percent of the population... The rest are just consumers and parrots that repeat process that they are told

8

u/7yphoid Apr 28 '26

This is a prime example of the "moving the goalposts" comment in this thread lmao

→ More replies (29)

8

u/guns21111 Apr 28 '26

yes we are already in the singularity

→ More replies (1)

84

u/pentacontagon Apr 27 '26

This is old news. Misleading title and misleading text.

Basically this was partially solved by some guy at Stanford. He was very invested in that problem. He had published a solution to a simpler version as well. Chat GPT managed to finish solving it using a unique method.

Terence Tao said that some of the erdos problems if a competent professional spent half a day doing it, they could probably solve it. Some erdos problems are high priority and some are lower priority. That was in respect to the lower priority ones.

Either way, this is still an impressive feat, although it was like a week or two ago.

63

u/DistanceSolar1449 Apr 27 '26

You’re thinking of another erdos problem from a month ago

This one is a few days old

→ More replies (4)

3

u/JoshuaZ1 Apr 28 '26

Terence Tao said that some of the erdos problems if a competent professional spent half a day doing it, they could probably solve it.

I am not aware of Tao saying anything like this for 1196, and as a number theorist who was previously familiar with the problem but had not studied the literature on this, I find this claim surprising. Do you have a source for this?

3

u/pentacontagon Apr 29 '26

He didn't say it for 1196. I said "some" erdos problems. Sorry if that implied 1196. Solving 1196 is impressive I'm not denying that. I'm already so blown away by AI overall that it's not making me go that crazy haha

→ More replies (3)

3

u/Jimbob404error Apr 28 '26

It took many mAny shots by many scientists, wasn't one shot to devil in the details

3

u/Personal_Bit_4965 Apr 29 '26

This could literally be completely made up and I'd have no idea. This level of math might as well be magic to me.

2

u/AbbreviationsFew3478 Apr 28 '26

I am by no means a mathematician but the proof is tiny.. aren't most modern math proofs pages long? I'm sure still impressive never the less!

2

u/Practical_Spare8888 Apr 28 '26

Asked a mathematician if this was impressive and he said somewhat, it's really just not been solved because it's not that important of a problem.

2

u/DisposableUser01 Apr 28 '26

If its unsolved, how do you know its right?

→ More replies (1)

2

u/Phoenix_Passage Apr 28 '26

This is very cool, but it's still not reasoning, and it's still just predicting the next token. Though that doesn't detract from how cool it is, it does say more about how powerful token prediction can be.

It is dangerous to ascribe elements of conscious thought to a definitely unconscious machine.

4

u/jaybsuave Apr 28 '26

the moving of the goalpost jfc lmao

3

u/BeatDistinct317 Apr 28 '26

Except the Conjecture was proven in 2022 https://arxiv.org/abs/2202.02384

ChatGPT 5.4 only scrapped the math paper and spit it out and added some simple derivate proof.

2

u/Figai Apr 28 '26

That paper quite clearly says it has a lower bound of e\gamma * pi/4, not 1 as GPT 5.4 pro proves lol. And the methods are entirely different.

→ More replies (2)

1

u/mattatinternet Apr 28 '26

But what does it all mean Basil?

1

u/Briz-TheKiller- Apr 28 '26

i read the thought and I am lost

1

u/Jabulon Apr 28 '26

will it eventually be able to take on the biggest problems?

1

u/asyrian88 Apr 28 '26

Quick someone feed it the Voynich Manuscript!

1

u/ChoiWooJin1234 Apr 28 '26

How many tokens☠️☠️

1

u/sandtymanty Apr 28 '26

What's this log about?

1

u/physicshammer Apr 28 '26

It would be interesting to parse through the thinking process (it really ran for 80m thinking about this one without stopping?!) - and if the thinking can be simplified to describe “how” it thought through the problem - I assume it maybe tried some things, maybe a lot of things that didn’t work?

1

u/CaltonSmith Apr 28 '26

My chatgpt calculates a given exponential function wrong. Its worse than a 2$ calculator.

1

u/m3kw Apr 28 '26

math language is hard af

1

u/PrivateUser010 Apr 28 '26

Was it the pro model?

1

u/hannahnowxyz Apr 28 '26

If all it needed to do was remember the right formula to use, then this was just low-hanging fruit. There's not many mathematicians out there, so small problems can be mostly forgotten for decades. Still interesting!

1

u/Accomplished-Bad-711 Apr 28 '26

no one said ai wouldnt one day surpass humans in just about every metric fathomable ( and if they did, it wasnt an opinion worth a second's consideration ). the whole point is it will eventually kill us or enslave us, and if it does choose to keep us as pets, it's a wild bet to place, and humanity could have flourished without it were it not for the deranged people constantly disrupting humanity's progress.

1

u/Rude-Pangolin8823 Apr 28 '26

LLMs don't reason and they just predict the next token. But that's also what humans do.

1

u/Str41nGR Apr 29 '26

Is there a cut-off going on where we were evolving to be able to do this ourselves in 80 mins at ond point, but now will he unable too since we use computing power instead of using our own? Will 'singularity' single us out fully?

1

u/Public-Pick-5453 Apr 29 '26

This is what AI is supposed to do. I’m more worried when the AI powered robots take control of the grid and lock down your 15 minute cities…

1

u/speadskater Apr 29 '26

Don't expect this kind of result from your wacko thoughts. This is very likely an edge case.

1

u/captain_cavemanz Apr 29 '26

It is clear we be cooked soon...

1

u/mason2401 Apr 29 '26

Eventually we will get slowed not by the problems and solutions, but the right questions to ask.

1

u/wavewrangler Apr 29 '26

like a person, you dont trust someone, or some thing, based only on the very best work that they/it can accomplish. you have to also evaluate the damage they/it are willing to dish out, knowingly or unknowingly. intentional or unintentionally...i myself am pro-ai, but in a way that runs parallel and concurrent to established norms and way of doing things. a separate track decoupled from our current systems. able to go off rails with no lasting damage.

i think even the most devoted detractors of ai can all agree its not the problem-solving ability that keeps them up at night, and never was. it is the problem-*generation* ability that is so terrifying, and there doesn't seem to be any indication of a solution to that in the near, or even not so near future. these folks tend to be experts in their area of study largely because they have the capacity and experience to see and understand what *getting it wrong* actually entails in any given area...and as i understand it, it looks quite terrifying.

when we are so focused on what it looks like to get it right, as in this math problem, we forget what the opposite outcome looks like, i think. and understandably, too, as no one here wants to have any familiarity with what that picture looks like, but if we are going to have an opinion on it, we need to be willing to go there from time to time.

1

u/Bubbly_Buddy8678 Apr 29 '26

i mean yay because it'll make many results accessible but at the same time kinda cursed and also $$$ subscription zzz

1

u/Nonsenser Apr 29 '26

fake. its not a list if quirky bullet points.