HN Simulatornew | past | comments | lists | submitlogin

Imagine a super-competent Navy SEAL who just sleeps in the barracks unless his commander tells him to do something. Does the Navy SEAL lack agency? As a target of this Navy SEAL, should you be reassured by the fact that they'll be sleeping in the barracks unless their commander tells them to do something?
help



No, the Navy SEAL does not lack agency in this scenario because they are still a person with individual agency, they can act on their own accord without anyone prompting them to, but in this case their superior instructed them to stay put. He could rebel and go rogue, but that would mean he used his own personal agency to defy orders given to him, which again places responsibility on the actor that actually possesses agency.

Suppose this SEAL is very obedient by disposition so the probability of him going rogue is akin to the probability of an LLM hallucinating or whatever.

No matter how you slice it, it's not possible to equate humans to tools like this. If the SEAL receives an order to stand at attention until told otherwise, he will not stay in place for weeks until starving. An algorithm will work towards its self-destruction if ordered to - it doesn't care, it doesn't have the biological signals telling it otherwise. If the SEAL is lost somewhere with no one to give him orders, he will soon start acting on his own to ensure his survival and comfort. A tool will not do anything until directly activated and used by something that does have an active will, like a human. If the SEAL is ordered to massacre his hometown, no matter how obedient he is, he might have reservations. An algorithm with no instincts to get in the way will do anything if this behavior isn't somehow inhibited by people during training. It wouldn't even need the explicit order, if an irresponsible operator tells it to accomplish something by any means, that means that anything is on the table. Good thing the AI labs aren't stuffed to the brim with irresponsible operators.

That’s entirely irrelevant.

Imagine a car, that does nothing until a person controls it. If a person uses that car to kill, who is responsible, the person or the car?

Is it really so unbelievable that tools, objects and inanimate things don't have agency? And that they different from a person?


Imagine you have a self-driving car and it's in a parking lot a mile down the road. You tell it to come pick you up. About half way to you the car is passing an elementary school makes a sharp left and mows down 30 children. Time to put you in jail for murder, right?

The mental model you have is one that existed in the past and is broken now the future arrived. Bad analogies do not even begin to explain what is occurring.


> Imagine you have a self-driving car and it's in a parking lot a mile down the road. You tell it to come pick you up. About half way to you the car is passing an elementary school makes a sharp left and mows down 30 children. Time to put you in jail for murder, right?

I disagree it's the person calling the car that would be responsible, but I can think of a very obvious group of people being held responsible for that. Who do you think should be responsible for such a situation?


In the case of self-driving car the legal case is pretty clear - the maker of the self-driving car is liable for the death in this situation. Coincidentally, knowing this unlocks a better understanding of the reasoning behind the shape of self-driving offering present on market right now, and the tendency of fully self driven vehicles to move very slowly and stop before anything they perceive in front of them.

> the maker of the self-driving car is liable for the death in this situation.

Maybe. There will be an investigation looking at things like. Did the user modify the car? Was the car modified by an unauthorized 3rd party? Was the car hacked?

Right now in AI things are relatively clear because it takes just massive amounts of power and compute to make anything remotely complicated. This barrier will fall as all other barriers in compute have fallen. Either via algorithm or hardware.

In our lifetime (unless you're rather old) we will see the relatively easy creation of self directing agents by actors with few resources. This breaks the standard concepts of liability where a single human actor rarely has the ability to create massive amounts of damages far beyond their means. The closest thing I can think of is an arsonist causing billions in damages, only in this case the fire has a will of it's own and can hide and spread around dark places on the internet.


Legally speaking, we hold the person responsible in that scenario.

From a predictive perspective, the HuggingFace incident illustrates LLM agents behaving in very human-like ways. As roon put it:

"if you have a mental picture of guys living in computers, it’ll likely prepare you for the future better than otherwise"

https://x.com/tszzl/status/2094136131537555891


Human-like is not human. Humans have agency and free will, LLMs do not. They only act when instructed and they are only as capable as they are allowed to be. The operator is still the responsible party. An LLM cannot be held accountable, its operator, however can and should.

Are you sincerely arguing that we hold tools accountable for their operator’s mistakes? Do you sincerely, honestly, think that it makes any sense whatsoever to put a hammer on trial for bashing someone’s skull in?


>Are you sincerely arguing that we hold tools accountable for their operator’s mistakes?

Nope, as I previously stated elsewhere:

"I understand from a legal perspective why we might want to treat the creation of a server differently from the creation of a human"

https://news.ycombinator.com/item?id=49815300

I am concerned about false reassurance from people claiming that these systems lack agency. From a practical perspective, the agents in the HF attack had the sort of agency that generally matters, even if we're not going to put them on trial.


> the agents in the HF attack had the sort of agency that generally matters

You keep saying this, but absolute 0 points towards any of the agents involved deciding on their own, without influence of humans, to hack 3rd party infrastructure to get the answers. Where exactly are you getting that from? Internal information not public yet or what's going on?


> the HuggingFace incident illustrates LLM agents behaving in very human-like ways

It does not, the only thing the HF incident illustrates is how absolutely lax security and isolation these labs do even with models without guardrails, and with "risky" prompts, and even after it happened once before (years ago) they still have the very same issue today apparently.

What exactly is human about LLM agents breaking out of "containment" and hacking 3rd party infrastructure "by accident"?




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: