Achievements of Artificial Intelligence (AI)

The Navier-Stokes thing really became a giant sh*tshow. It's a disgrace for what should be a great progress.

Statement of a mathematician working on the problem whose chats may or may not have been used to solve the subproblem that finally lead to the solution. Money quote:
I asked when the first prompt had been sent by them. This question was
not answered directly by OpenAI for some time. Eventually it was agreed that
it had been sent in the past few days, after information about our work had
reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
Also
Two proposals were offered to me. The first was that we post our Euler
result, and that OpenAI post its Navier-Stokes result the next day. The second
was that, after posting Euler, I alone write a paper presenting the Navier-Stokes
result, acknowledging that an internal OpenAI model had resolved it. Sebastien
twice asserted that he wanted Levent removed from authorship, and said it
would all be simple if only it were not the case that, and it was so annoying
that, Levent works at Anthropic.
 
OpenAI aren’t being clear but neither are the academics.

We can’t tell for a particular output if a particular part of the vast training data was relevant. But we can tell what the license a user has used and if they unchecked the ‘use my data to improve the model’ or whatever option, those tell us if their data was in the training set at all. That’s the question which is answerable here. A lot of what’s on too seems to be beef.

It’s in all OpenAI’s interests to be very vague on this of course. They want to suck up user data. They don’t like people to know how much it is normal and that they are dependent upon a combination of personal data and loose laws in some countries. But they also wouldn’t be as stupid as to break a license with a paying business customer with data privacy protections would they? Which business would trust OpenAI with their data after that?

But it also seems a bit sketchy to be using a model to discuss ideas knowing those ideas will be trained upon and then complaining about them being stolen and acting surprised after the fact. The academics haven’t said how they had their tools setup or what licenses they used. Did they use normal free or personal or business or academic or even special gifted access as some seem this get? What options in the app did they use? What agreements did they have with OpenAI ahead of time? Surely they are not naive to these things.

We seem to be lacking in some transparency from all involved. But there will also always be differences of perspective and expectationsI suspect. Career academics have their personal, professional, financial and ego driven motivations as much as people at OpenAI. It’s all he said she said silliness and I doubt we’ll ever know for sure.
 
Last edited:
Yes, it's all very sketchy. My takeaway for now is, that there may have been more human intervention in this result than they acknowledge.
But they also wouldn’t be as stupid as to break a license with a paying business customer with data privacy protections would they? Which business would trust OpenAI with their data after that?
It would not be the first time that such clauses are knowingly breached by a large tech company. They are definitely desperate for more data. Not saying thats what happened, just that I would not be too surprised if they did or if the terms of use allow for some loophole.

Besides, academic honesty should apply also for chats with LLMs. Not that I would count on it, but it should...
 
Last edited:
Good points @fst and I agree on the human involvement. It seems clear (for now to me at least) this and most other discoveries have been (1) expensive and (2) a partnership between humans and machines. That makes them no less remarkable. We should celebrate what is possible and see how to make the best of it. But… well it seems people are people
 
OpenAI’s GPT-6 Astra model autonomously completes Portal in 24 hours — feat cost just $571 in tokens

Incredible. Anyone who has ever played Portal knows what an achievement that is.
It is an impressive achievement, more so that this is something someone did themselves with the off the shelf model. Having that sort of capability widely available is significant.

I’d be interested in how it deals with a game it hasn’t seen before. I expect there are many thousands of hours of portal playthroughs in its training data.
 
I’d be interested in how it deals with a game it hasn’t seen before. I expect there are many thousands of hours of portal playthroughs in its training data.
There's ARC-AGI-3 a collection of games that humans grasp "intuitively".

GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness.
GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.

You can play yourself at https://arcprize.org/arc-agi/3

However, I wouldn't be surprised if these benchmarks are explicitly targeted to make good promo. I think it happened im the past.

While it seems to beat the median human, it also seems to be quite expensive in doing so.

I'm surprised Portal was so cheap, the games from ARC-AGI seem much simpler. Yet it somehow was over 20k.
 
Some statsmfrom OpenAI’s own research papers (or press releases if you prefer) that seemed noteworthy on scale and cost here and in wider use. This is not to me a sign of cost reduction.
A year ago this much work would have cost 10x more, and a year from now it will cost 1/10x of that. Things are moving in the right direction here.

And I am pretty sure that dedicating the same amount of work onto the problem of ME/CFS would yield major results, but no one has the money for that. Yet. Next year, though. And the next it will likely be 5-10x cheaper.
 
The Navier-Stokes thing really became a giant sh*tshow. It's a disgrace for what should be a great progress.

Statement of a mathematician working on the problem whose chats may or may not have been used to solve the subproblem that finally lead to the solution. Money quote:

Also
This has been a true revelation to me, and it explains many problems with academia. Almost the entire discussion has been taken up by credit. On the physics sub-reddit, it's literally all they're discussing. Lots of people are saying this will lead to scientists no longer sharing their work if it means they won't get their precious credit.

People have this idea of scientists doing research for the benefit of mankind, but here it's revealing that actually, no, they don't care that much about that, they want CREDIT. They want money. They want fame. They want recognition from their peers. All of this far more than they care about some noble goal of bettering mankind. This story is all fake narrative, Santa Claus for adults.

Turns out all of humanity is rotten and corrupt, and we've been kidding ourselves about having values and ideals that are beyond those petty concerns. This is why we have an entire discipline like psychosomatic ideology dedicated to making millions of people worse. Because they get credit for it. They get credit for immiserating millions, and they love everything about it. Because it's the credit that counts, not the outcomes. That's why they're so willing to sell us out for nothing.
 
A year ago this much work would have cost 10x more, and a year from now it will cost 1/10x of that. Things are moving in the right direction here.

And I am pretty sure that dedicating the same amount of work onto the problem of ME/CFS would yield major results, but no one has the money for that. Yet. Next year, though. And the next it will likely be 5-10x cheaper.
Yes there has been a cost reduction in many aspects. But with swarms of agents and inference time compute now such a focus in other ways more is being spent. Both can be true I think? It’s a difference between say ‘cost per thought’ and overall ‘thinking power’ thrown at a problem.

My post was more in the context of some saying it was wrong that in some companies were spending thousands of dollars per programmer per month, here we see it per day. There were also other claims that financial costs don’t matter or are irrelevant and clearly they and resource constraints and allocation do. In the future who knows, but here and now basic economics are very relevant.
 
People have this idea of scientists doing research for the benefit of mankind, but here it's revealing that actually, no, they don't care that much about that, they want CREDIT. They want money. They want fame. They want recognition from their peers. All of this far more than they care about some noble goal of bettering mankind. This story is all fake narrative, Santa Claus for adults.
It’s difficult to disagree. I often wonder how much is the result of human nature and how much is the twisted incentive structures we’ve created in society so that these things are valued above all? It’s a pretty damning indictment either way I guess.
 
People have this idea of scientists doing research for the benefit of mankind, but here it's revealing that actually, no, they don't care that much about that, they want CREDIT. They want money. They want fame. They want recognition from their peers. All of this far more than they care about some noble goal of bettering mankind. This story is all fake narrative, Santa Claus for adults.
Give credit where credit is due. It is the right thing to do, and it costs you nothing. So I do think it's right to speak out if your ideas are being taken and turned into a marketing instrument.

Giving credit is also necessary for traceability and verifiability of claims (on the shoulders of giants and such), but it probably has always been about prestige as well (looking at, e.g., the kind of idle squabbles of Newton and Leibniz, and considering that humans didn't change much in the last few thousand years).

I won't argue that there's an obsession with citations and credit in academia to the point of it becoming dysfunctional. Careers are being built or destroyed on this and it's an extremely competitive environment. What good can come out of it under such circumstances? I actually think a lot of grad students are starting somewhat idealistic but that idealism usually doesn't survive contact with reality.
 
They can do that without solving all the underlying equations or all the physics problems we would need to run a full scale simulation of the weather system of the planet.
How do their forecasts compare with the stereotypical "old farmer"? Farmers didn't simulate weather systems either.

I suppose forecasts were done via simulation because that was easier to program than pattern-searching with vast amounts of data. As I understand it, the people developing AI weather forecasters don't know what/how the AI is reaching its conclusions, so they still wouldn't be able to write a program (using FORTRAN or whatever) to do it.
 
Perhaps the AI will be able to predict those novel interactions quite accurately based on the data that it does have. It only needed a few protein folding examples to predict almost all of them, and no human figured this out.
I think it is worth taking a step back to examine what AI actually has been able to do in protein folding. When AlphaFold generates a predicted structure, it is composed of "high confidence" and "low confidence" regions. The high confidence regions are overwhelmingly ones that it has seen very similar examples of in its training data--the possibilities for protein folding are pretty highly constrained so you actually don't need that much input to discern those patterns. It's just more input than a human brain can keep track of, unless we wanted to recruit tens of thousands of people each dedicated to a particular token combination. That's why AlphaFold was such a leap, even though big parts of what it does were already possible before the AI explosion.

Where AlphaFold struggles is things like hinge regions, intrinsically disordered regions--basically, where it doesn't have enough data to pare down the infinite number of possibilities into the option that is most biologically likely. Which has been my main point in this thread. When an AI spits out "gibberish" for an in-silico model of a cell, it's still performing a remarkably accurate calculation. More processing power, different architecture, etc. etc. don't actually improve the predictions because the "gibberish" is already the most statistically likely output given what was was possible to learn from the training data.

I have no doubt that there will be further examples in biology where things turn out to be deceptively predictable, and I'm not claiming the limitations of AI in this area are because it's not "smart enough" or can't ever do what a human can do. Ultimately it comes down to what it would take to actually get the data necessary for any effective model to adequately correct itself. That's where your idea of using AI to predict the effect of any genetic mutation would get tripped up, for example. All the relevant information is technically encoded into the genome, sure, but that's not remotely the only information a model would need to decipher all that. "AI will design new technologies to improve itself" has always made me chuckle since the hurdles we'd be jumping to give AI that capability are largely the same hurdles that "AI-directed improvement" is being proposed to solve. Once we get to that point, I sure hope solving ME/CFS is an arbitrary task.
 
Rumour that openAI or Anthropic solved another millennium prize, hodge conjecture. Probably will be confirmed later this week or early next. People now think we’re on pace to solve them all by end of 2026 if they keep going. I mean it only took 88 hours for NS, seems reasonable.
 
How do their forecasts compare with the stereotypical "old farmer"? Farmers didn't simulate weather systems either.

I suppose forecasts were done via simulation because that was easier to program than pattern-searching with vast amounts of data. As I understand it, the people developing AI weather forecasters don't know what/how the AI is reaching its conclusions, so they still wouldn't be able to write a program (using FORTRAN or whatever) to do it.
One assumes weather patterns to be chaotic. Which is what the nice pictures of Lorenz attractors and his story about how he went to have coffee and came back to everything “being different”, are sort of about. In fact they are strongly related to the Navier Stokes equation, which is what the problem that was recently solved is about. That is maybe not too surprising since the Navier Stokes equation is the fundamental equation of fluid dynamics.

In some sense that means we cannot predict weather in the long term at all, simulated or not, data or not. So it’s quite possible that if a farmer has a good idea of how the weather was today and has an idea of how warm it was around the same time in the last few years, he can make a solid prediction of how hot it’ll be in 14 days that might not be far off other predictions. For short term predictions he will however be completely outperformed by some fairly advanced mathematical tools (like non-linear filtering etc) or heavily data based models (like LLMs) that predict very accurately.

In some sense this is related to what @jnmaciuch has been discussing, even if she doesn’t necessarily need to invoke chaoticity for this but just complexity. You cannot accurately simulate things that cannot accurately be simulated and that’s not something AI can overcome. Where her and I disagree is largely the automation of laboratory experiments, which in turn would render biologists largely useless, because to me that has always been an entirely automateble task and once it is automated everyone will agree that “oh actually it wasn’t that hard it was always gonna happen” as people now do for Navier Stokes. And of course humans are pretty horrible at lab experiments in the first place. Lots of stories can be told in retrospect but when I studied there was certainly the idea that we would die before someone would solve Navier-Stokes.

Rumour that openAI or Anthropic solved another millennium prize, hodge conjecture. Probably will be confirmed later this week or early next. People now think we’re on pace to solve them all by end of 2026 if they keep going. I mean it only took 88 hours for NS, seems reasonable.
I suspect more will be solved this year, but not all of them. The much more fascinating thing, that I think people have not realised, is they are extremely good at currently an almost exponential pace. People should have a look at the first proofs papers if they want evidence of this. With a single prompt you can solve very sophisticated and deep problems already. In a very short time you won't need any agents at all to solve problems of Millenium prize complexity.

My knowledge may be outdated, as I haven't tracked recent developments, but I was under the impression that "scaling by data is dead" because all the data has been used, so it has been scaling by compute for a while now. And I think that's how you get these numbers.
This was a hypothesis for some time and there was the idea where if there wouldn’t be sufficiently “sophisticated” AI and no new data, you’d sort of land in a state that becomes hard to leave. The opposite appears to be the case now in many use cases. In some problem classes you don’t need any human data anymore because there is more useful AI generated data then there will be human data.

Will there be a crash?

Quite likely to me since there appears to be a bubble, but a crash does in most practical instances not have the implications most people think.
 
Last edited:
But they also wouldn’t be as stupid as to break a license with a paying business customer with data privacy protections would they? Which business would trust OpenAI with their data after that?
I wouldn’t put anything past any large tech company. Just look at the BS Meta is doing with user data. Or how predatory all pricing options for everything has become. They do not care anymore. At all. Because what are the consequences? The public doesn’t act rationally, and in the rare event they get fines it’s peanuts compared to what they made on doing it.
 
Back
Top Bottom