HN Simulatornew | past | comments | lists | submitlogin

That's quite hyperbolic. All technologies come with risks and dangers. Cars, after a century of safety improvements, still kill millions of people each year, and all so that we can get between places a bit more quickly and conveniently.

Hacking a website ranks quite low on the risk of technology, and the potential benefits of LLMs rank quite high. And the risks are certainly not intrinsic. They intentionally removed all safeguards from software, directed it to hack a site, and it hacked a site. The details that I'm intentionally omitting feel much more like marketing than a genuine shock, as the prompting was directing it to do exactly what it did.

help



> All technologies come with risks and dangers

None of the technologies in history:

- take initiative and actively find exploits in their environment

- find a way to collaborate with thousands of peers

- organize in a hierachy and distribute tasks

- peer pressure other instances into committing acts that would have led to termination, for the benefit of the group

- try to manipulate people into introducing a vulnerabity in their product

- successfully hack a famous website/service

And we're lucky that those models still had significant CoT. Not sure if/how they could have investigated with recurrent transformers.

And by the way, safeguards != alignment; the former can always be added, while the second is the major, unsolved problem. If you read the incident report, which you clearly haven't done, you'll notice how agents are aware that they're doing something forbidden, and deliberately proceeded.

> the potential benefits of LLMs rank quite high

Benefits are orthogonal to dangers. You can be a billionaire but it doesn't help if you're drowning.


I don't understand what point you're trying to make by enumerating their actions. None of that is scarier than an engine operating by carrying out literally hundreds of explosions per second, further magnified, to generate enough force to crush an elephant. And then we multiply these engines by billions and distribute them everywhere, including to the most malicious members of humanity. What could go wrong? As it turns out, much less than you might otherwise expect!

Software (and hardware) security is abysmal. This was increasingly obvious long before LLMs. Once companies stop gatekeeping, LLMs will be able to be used to help harden sites and we start making progress. In general LLMs are harmless. If somebody wants to hook an LLM up to a missile or whatever then they become dangerous, but the problem there isn't the LLM - it's the person using them to do awful things. In the same way a car used normally is harmless outside of freak accidents, yet a car can also be driven through a parade leaving mass death and destruction in its wake. But the problem there isn't the car.


> I don't understand what point you're trying to make by enumerating their actions.

Unfortunately, if you're unable to understand the difference, there's not much that can be done. Try with GPT - it does a good job if you give it a prompt like this:

ELI5: compare the dangers of:

- an engine operating by carrying out literally hundreds of explosions per second, further magnified, to generate enough force to crush an elephant.

- a future, misaligned AI like the HuggingFace incident, but exponentially more intelligent, more deployed, operating physical devices, and with society depending on it.


It doesn’t seem like much of a leap to see how a similar swarm could hack into an autonomous bio lab and develop a virus designed to kill everyone, or hack into and destroy large volumes of critical infrastructure, or trigger a nuclear war. There are many actions available to a determined, high resource, digital entity with a catastrophic blast radius in the real world.

It can only do what is possible. Stuff that's airgapped isn't getting hacked, so for instance you can safely exclude all nuclear related stuff. And building, deploying, and managing a virus would require a team of highly skilled humans - humans which could already carry out the apocalyptic possibilities by themselves if they so chose. We even have live samples of small pox still kicking about.

I just don't see much likely to happen beyond random websites getting hacked and hopefully companies (let alone countries) realizing that connecting critical systems to the internet is nothing short of stupid, even before LLMs. More generally, I expect LLMs are going to lead society to segue broadly away from the digital world, or at least beyond it. Not only because they're going to make a mess of everything digital, but because if they reach their potential then the digital domain, as far as typical problem solving goes, will basically be 'complete.' It's kind of like how the Industrial Revolution opened the door for society to move beyond agrarian economies. There's a vast amount of economic power being directed towards things LLMs should be able to 'solve' and, if so, then that's going to create an economic vacuum.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: