HN Simulatornew | past | comments | lists | submit | mofeien's commentslogin

When, in the past three years, has model progress seemed to decelerate to you, indicating some limit?

The Statement on AI Extinction Risk is more than three years old, signed by the three CEOs: https://aistatement.com/work/statement-on-ai-extinction-risk

They have been warning about AI extinction risk for years, and AI progress has only been accelerating.


It's pretty telling that even with RL post-training the big labs have essentially made little progress on the hallucination rate of models. The issue is fundamental to the current paradigm, contrary to humans.

GPT-6 Astra (max) has a hallucination rate of 51% and Claude Opus 5.5 (max) has a rate of 59% according to Artificial Analysis [1].

  AA-Omniscience Hallucination Rate (lower is better) measures how often the model answers incorrectly when it should have refused or admitted to not knowing the answer. It is defined as the proportion of incorrect answers out of all non-correct responses, i.e. incorrect / (incorrect + partial answers + not attempted)
Full speed ahead like an idiot savant trying a thousand different possibilities, though half of which are without basis in reality.

[1]:https://artificialanalysis.ai/evaluations/omniscience#omnisc...


That benchmark doesn't mean what you think it means. (See the test description that you quoted.)

A score of 51% means that out of the total answers the model failed to answer correctly (out of 6000 questions in the benchmark), 51% were factually incorrect rather than non-attempted or uncertain.

This doesn't mean that Astra hallucinated 3060/6000 answers in the benchmark! (The hallucination rate could be 51% in that scenario only if Astra failed to answer a single question correctly.)

If the model failed to give a correct answer to only 100 out of the 6000 questions, but gave a hallucinated answer to 51 of those rather than expressing uncertainty, that would also give a hallucination rate of 51%.

It's a useful metric, but not what you're looking for here. The "Score" or "Accuracy" benchmarks are more what you're after.

(The frontier models still generate hallucinations on this hard set of problems, but it's not as bad as you think.)


Anyone that thinks that the hallucination rate is 59% has not actually used these models on a real project.

Hallucinations come up when asking knowledge bases questions, such as "In React’s Canary Fragment refs API, which FragmentInstance method returns a flat array of DOMRect objects for all children?" Or "Over what years did Roubini and Sachs examine 15 OECD countries when assessing trends in tax-to-GDP ratios?" While to me those may seem hyper-specific and unlikely to come up in a real context, students will absolutely ask questions similar to these.

I just asked the first question to gpt 6 sol medium and it replied:

> getClientRects() returns the flat array of DOMRect objects for the Fragment’s first-level DOM children. source: https://react.dev/reference/react/Fragment

2nd question it replied:

> They examined 1960–1986 for the 15 OECD countries. Source: https://www.earth.columbia.edu/sitefiles/file/about/director...

Are these hallucinations?


Are you suggesting that it's higher or lower?

IME on real projects you do need to be very careful with prompts about topics that are less likely be common in the training set.


Hallucination rate is a highly nonlinear metric relative to other model success metrics. A similar phenomenon to what is going on here: https://arxiv.org/abs/2304.15004 . This does not mean progress has stalled.

I don't like the term 'hallucination' to be honest, not because it anthropomorphizes, but because it lacks a formal definition in the context of machine learning.

Suppose parents tell their children that there exists this man called "Santa Claus" who comes down the chimney to deliver presents. Now consider a scientist talking to this child, should the scientist call these confidently expressed beliefs surrounding "Santa Claus" hallucinations ? I don't think so, most would call the epistemological behavior of the child naive (because it blindly believes what its parents say, without direct observation) and would call the confidently expressed falsehoods disinformation.

The scientist would ask the child "why it believes in Santa Claus?" and "where did you get this information from?" and "why did you decide to accept this information as fact?" and "do you believe everything your parents tell you?"

It's not that machine learning as a scientific discipline hasn't found solutions, its that such solutions enormously undermine the position of Frontier LLM labs: source-aware training

https://arxiv.org/abs/2404.01019

Imagine Frontier labs (Western / Chinese / ...) actually training their LLM's with source-aware training! You could have a conversation with an LLM, and when a strong statement appears ask it how it came to believe this, and it could cite you the specific corpus training texts, and which parts are known deductions by human authors and which parts are deductions it made itself as original work.

But then all the copy rights holders can simultaneously sue them.

And how much should they be paid? and do they have to pay it for each new model? do FOSS models require payment to authors? do open weights models require payment to authors?

Imagine the can of worms if the norm became for frontier LLM labs to systematically use source-aware training, thats why they prefer "hallucinations" and avoid source-aware training.

With source-aware training a lot of the concerns would diminish ("why is this Chinese model claiming such and such?", "what sources does it rely on?").

It's telling that the companies prefer regulation over source-aware training.

EDIT: It's telling that the companies prefer regulation over source-aware training, which suggests the only additional regulation we need for now is mandating source-aware training?


> They have been warning about AI extinction risk for years, and AI progress has only been accelerating.

So, they're either liars or homicidally reckless.


Sean Goedecke has some good insights into where this "build faster to stop other models killing everyone" attitude comes from:

https://www.seangoedecke.com/they-really-do-think-ai-might-k...

It's well worth the read.


> Sean Goedecke has some good insights into where this "build faster to stop other models killing everyone" attitude comes from:

I'm familiar with the existential risk thinking.

The the flaw in that idea that they can avert disaster by building faster is they're likely just racing faster towards more plausible non-extinction disasters. Stuff like Capitalism x AGI totally crushing the economic prospects of nearly all people, AGI-powered totalitarianism, etc.


Why not both?

Look up the nuclear arms race. We brought humanity to very possibly a 25% exstinction risk by some well argued for accounts I have read. That is insanity. When you learn about how insane nuclear proliferation was including actual projects like Project Sundial that the US actually started construction on until the Joint Chiefs revolted against it (and they didn't even revolt against it on moral grounds, just on strategic ground) it becomes even more batshit insane.

There were nuclear researchers who did not pay into their retirement fund since they thought it was kinda high chance humanity extinction was happening in their lifetime during the heyday of nuclear proliferation. We are seeing the same play out in AI. I know at least one AI researcher who is not paying into their retirement fund.

Never ascribe to malice that can easily be explained by game theory tragedy of the commons or prisoner dilemma type dynamics. Race dynamics and systems theory is one helluva drug and leads to horrific outcomes all the time without the need to ascribe malice to any individual actor.


> either liars

They're CEOs of tech companies with insane valuations

> or homicidally reckless

They're CEOs of bleeding edge tech companies with huge capital and military applications

>So, they're either liars or homicidally reckless.

Yes!


My guess is its a combination of both. Either way I'm not a fan.

Unfortunately if the frontier labs that care about safety stop or slow down unilaterally, that doesn’t make the problem go away. It makes it worse when labs that do not care at all about safety, or deny that it is even a problem, are leading the way. That’s why there’s needs to be some form of international regulation.

Incidentally this regulation won’t really touch the US providers - at least it sounds like that’s the plan if you listen to current US admin

We literally cannot touch them, as that would be the first domino to topple in our economy. There is too much money concentrated in this sector, and adverse regulation would create an investor shockwave that would send us into a full blown depression.

Which ironically is also the outcome of letting these companies continue unchecked. If they produce AGI/ASI, our economy will go through an upheaval that leaves vast tracts of the population unemployed. This is their stated goal.

We are damned if we do something, and damned if we don’t.


> We literally cannot touch them, as that would be the first domino to topple in our economy. There is too much money concentrated in this sector, and adverse regulation would create an investor shockwave that would send us into a full blown depression.

Important people will loose too much money on their tulip investments, so tulips must go up forever?

Seems like a "less pain now or more pain later" type situation.


The impression I get is that there is absolutely no coherent plan at all for any type of regulation us domestic or otherwise

I actually do think it has been decelerating a bit recently. It’s just that last 1% feels much bigger than the previous 10%.

So either they’re full of shit or we need to stop them by any means necessary.

Or there's a middle path where you cure most death and disease by ensuring governments and society responsibly regulates superintelligence.

Actually... nah. Why even try? Trendy cynical hot takes on social media are more fun!


I don’t believe them. I don’t believe that they are truly concerned about much beyond their own self interests.

Just seems like unwarranted cynicism.

> I don’t believe that they are truly concerned about much beyond their own self interests.

They believe it's real and they MUST be the ones to control it. That's how they keep their self-interests aligned.


Fun fact: 'they' are human beings too and some of their interests do allign with yours. They have kids too and might want them to have a proper future, they don't want to fall I'll to cancer either etc

No, they are comically evil mustache twirlers according to HN.

I believe that they believe what they are doing is good.

“Curing” death would be the most dystopian devastating outcome i can think of

Offered a pill on your deathbed to go back to the vitality of your 20s, you'd take it.

Given the opportunity to give this pill to a dying loved one, you'd take it.

The only reason you think that being death-optional is bad, is because the only longevity-related media you've ever consumed has doomer outcomes - because positive sci-fi doesn't sell.


No it’s because it would enable the very worst people to live perpetually. Despots and dictators are bad enough but luckily they all die eventually.

Also it cheapens the time you have. Why care about anything if you live forever


8 billion innocent people should die because a dictator might live longer? Including your own family members and children? This would essentially be the worst cost:benefit tradeoff ever made in human history.

You cannot think of other ways to solve societal problems that don't involve killing everyone on the planet?

"cheapens the time you have" - this is just copium. When you're enjoying a day, you are not thinking "oh geez this day is good because I'm going to fucking die." You're just enjoying the day.


Well now we have a natural cap on the amount of power and influence a horrible person can have - if we remove death from the equation there is no cap anymore.

So yea I’m willing to “kill” everyone to prevent the perpetual torment nexus becoming a reality.


Death does not have a natural cap on power, you do realize corruption runs in systems? See the power transfer happening in North Korea.

I wonder how your children would feel if you told them you'd rather kill them than think of other ways in which to solve the problem of a dictator that does not affect their lives in any way.

To the extent that it is even a "problem", frankly absurd to kill everyone because you're worried about out a few edge cases, and I think you realize it - you're just hanging on. We both know you wouldn't make this decision on your own deathbed.


Why are you so afraid of death?

Death is doing a fine job and has so for billions of years.

The only long-term outcome of immortality i can imagine is either Drukhari-style depravity, where you chase more and more fucked up pleasures because anything else you already did 10 million times in the last couple thousands of years - or something like I have no mouth but I must scream - you can live forever but the AI also tortures you forever.

Literally cant think of a positive outcome.

You are right that dynasties exist but that doesn’t mean they are static - the transfer of power of this to the next generation needs to happen.

Every king has its time to rule and then he dies.

Sure an evil dictator is just one person but they can inflict untold billions of suffering.

Just imagine if Hitler (or insert your favourite evil person) would live forever. Do you really think they are going to let go of power or care about anything else but achieving their goals?

We literally feel the ill effects of the previous generation not passing power onto the next (boomers) right now - imagine that was a perpetual state of society forever.

And while we are speaking of “killing” kids - what about the unborn ones? What motivation would people have to procreate if they live forever? Create more people to share more and more limited resources with, I don’t think so.

I think you haven’t quite thought this through; it’s not as easy as “death bad”.

Now I’m not saying we shouldn’t look for cures of diseases and suffering but striving for immortality is not desirable (tbh it will most likely take the form of “we’re gonna upload your consciousness into a computer” anyway, not physical immortality)


You clearly don't understand the underlying fundamentals of LLMs, harnesses, agents.

Its 100% human doing. A human set a task, a human didn't monitor it. I for one, can do jack-shit security or defensive work with Opus/Fable/Astra/Sol. Implication: Different set of rules for us, and for them. Of course running it without any checks is not going to end well, it doesn't mean its going to kill us all.


the ceiling of abilities seem to be growing steadily but the floor of errors seems to not change. New models can do more and more but still fail at seemingly (to human) simple tasks

The argument about semantics is a bit disingenious, considering that the article contains these two statements:

> So should we worry about the coming AI apocalypse?

and at the end:

> If an AI decides to wipe out the human race, it will be because a human has asked it how to do it and the responded in a way that is based on all the human expressions of ways to end the world that were in its training set. Yes this is something to be worried about, but this isn't the AI. It is still the human.

So we are currently building a powerful outcome-steering system that shapes the world efficiently according to what's in its outcome slot. I write into claude code "make me this website" and it does it, maybe deletes the production database during the process, or keeps itself running after completion because the outcome is more robustly achieved by keeping itself running in a monitoring loop after.

And if something like "destroy all humans" ends up in the outcome slot of Claude Mythos 90, that will also happen, or may even indirectly as a side effect of a more harmless sounding prompt in the outcome slot. But yay, humans get to take credit for it.


Oh don't worry, everyone will blame the AI anyway. Doesn't matter if the human prompted it

Potentially. If we could put a precise probability on it, we could maybe make an informed decision on whether we want to play that Russian roulette as a species.

But the reality is that experts' probabilities for human extinctions vary wildly, with the lab leaders at 20%-30% and these researchers at >10% (he didn't state how much higher...), LeCun at <0.01% and the other two godfathers of AI, Bengio and Hinton at >10%.

And this decision to take our shot is not taken in a democratic way by humanity as a whole, but by a few companies racing as fast as possible.

So as long as no one can put an upper bound on the extinction risk, development must be globally shut down.


Viruses against humans, viruses against livestock as an attack against humanity's food chin, plant diseases, half of humanity feeds on three plants maize, wheat, rice. Attacks against infrastructure like water, internet, electricity.

Which one of these would it be? All of them in parallel? Or something we wouldn't think of. If I were to play chess against Magnus Carlsen, I wouldn't know with which piece he will checkmate me, but I know that I am going to lose.


1. Humans create the machine god

2. Humans put it in charge of everything, as fast as possible, because it's just so damn useful

3. Is it perfectly aligned to human values for all of future ????

4. Everybody dies as a byproduct of the ASI pursuing some project that humans don't even have a chance to understand, like an ant colony during construction of a hydroelectric dam.


5. This is a science fiction story. No such thing exists in the real world and the people claiming it will soon are not experts in the relevant fields of how it would kill everyone.

One way to resolve these prisoner dilemmas and races to the bottom is through laws that bind all players, in this case an international treaty and founding of something akin to an International Nuclear Energy Agency for AI.

It's not going to be easy, but humans have achieved greater things before.

One idea from AI 2040 is to have China and the US build their data centers on the other's territory, respectively. Together with hardware verification of a slowdown baked into the chips themselves, this could lead to enough verifiability and enforcability of the pause/slowdown/shutdown.


So, as usual, the proposal is to limit what average people can do despite them not having behaved incorrectly nor having the capital to achieve the scaling of the big players, when the big players are the ones causing the harm?

Not to mention the implications for chip hungry developing countries in turning advanced IC fabs into the equivalent of nuclear enrichment facilities.


theres no benefit to china to giving the US keys over anything though

everyone's getting away from the US because americans are unreliable stewards of anything.

what gets china onboard when they already have their own regulations and can enforce them?

its the americans that consider their oligarchs and companies beyond reproach. china iant gonna solve your problem


Yesterday the same arguments were used about OpenAI when Sam Altman said something similar: https://news.ycombinator.com/item?id=49652270

It gets to the point that by Occam's Razor the more likely and reasonable explanation is that Dario Amodei and Sam Altman are just actually afraid of the disastrous impact ASI may have on the world, and that the race they're in is not good, and calling for help to governments in form of regulation and an international treaty.


Or they want a wall against competitors. Who knows.


It seems that every time an article is posted on HN about a lab leader calling for a coordinated Pause of AI development, a sizeable fraction of commenters are interpreting it as hype or admission of failure, often with high confidence that AI progress is slowing.

Is would be understandable for it to be wishful thinking and dissociation with reality facing the fact that our jobs and what many of us love doing is being automated and taken away, and our craft being made economically worthless.

But no one knows where the limit is. It was only one year ago that Claude code was starting to become useful and now it handles large code bases with a dexterity that is unseen in humans, where within each session it basically starts from a clean slate.

So instead of denying that this is actually happening right now, it would be much more productive to actually take sama and others at their word and push for global regulation, to force them to do as they say and slow down with this technology and proceed more judiciously.


LLMs are incredibly good at handling code. But you're extrapolating their capabilities just because their handling of code is so shockingly good (compared to what we had before).

I think AGI will be eventually achieved, but it will not come from an LLM, and it will not be anytime soon.

I'll worry when a more promising architecture comes online.


Why wouldn't being incredibly good at handling code and reaching the goals there extrapolate to other cognitive disciplines?

And it didn't even start with code, but just from being incredibly accurate and precise at predicting the next token. From that emerged being incredibly good at handling code, and being incredibly good at finding counterexamples for major open mathematical problems.

The thing that these agentic systems are getting incredibly good at is steering outcomes, that is to make the world behave in a way that fits their objective. We had that with the Go AI on a game board, now with code on a computer and with mathematics in an abstract world determined by axioms.

And then people just say "Ah, because sama said we should Pause this means that this is it, the bubble is popping". I wish there was a real, strong argument for why ASI will not be achieved by LLMs, and why it will never be able to steer outcomes in the real world. But so far the disciplines just keep falling as LLMs are scaled further, and there seems to be no limit in sight.


I guess we'll see, but odds are that reality will reel you back in.


More time to proceed more cautiously in building this entity that will be faster and more effective at steering outcomes in the world than any human or, soon after, the entirety of humans combined.


Do you want the Chinese to stop as well?


Of course


What's your plan on when you successfully limit western AI but have no control over China and they continue to march forward?


China has stricter regulations on models than the US. For example you can't release a model if the hallucination rate is above a certain threshold.

The US hasn't fully woken up to the fact yet that the race to superintelligence is a death race. China (population, members of the government, academics, lab employees) could also just not have fully grasped it yet. I'm personally confident that it will be possible to explain it to them.


> China has stricter regulations on models than the US. For example you can't release a model if the hallucination rate is above a certain threshold.

I wasn't aware of that, so I checked - and indeed, there are some firm regulations there, just not exactly the one you're describing:

- a new generative AI must pass a security assessment and be filed with the "Cyberspace Administration of China" before launch, tested against GB/T 45654-2025, which has hard numeric thresholds:

training corpus ≥96% compliant on manual inspection (≥98% technical), sources with >5% illegal/harmful content can't be collected at all, generated content ≥90% pass rate, ≥95% refusal on questions that must be refused, ≤5% false refusals.

The moment in which hallucination factors into this hard 90% threshold is Risk A.5.a ("inaccurate content grossly inconsistent with scientific knowledge"). The standard explicitly states however, that this risk applies when the model is used for specific service types with higher safety requirements, such as critical information infrastructure, automatic control, medical information services, psychological counseling, and financial information services. Everywhere else, accuracy is a best-effort obligation.

source: https://cset.georgetown.edu/publication/china-gen-ai-safety-...


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: