How specifically did misconfiguration lead to the HuggingFace attack? You could argue that its sandbox was misconfigured, sure. But suppose you had a similar incident where its intended task required access to the internet, and it veered off course in a similar manner. I don't think "misconfiguration" would be an accurate description of what went wrong in that hypothetical.
The doomers already have a term which fits pretty well: "AI misalignment".
We indeed lack much of the vocabulary. From a practical perspective, however dangerous the creation, if you cant punish the creation for what it does it leaves only the one who started the process. If it's human error or intentional neglect for personal gain should be for the court to decide.
>if you cant punish the creation for what it does it leaves only the one who started the process
Agreed, but I think we can do more on the prevention side as well. Traditional liability law is for negligence in case of preventable disasters. Since we currently have no way to prevent AI disasters in principle (alignment problem remains unsolved), I think we should just stop developing the technology for now: https://pauseai.info/
<< "Bot" doesn't carry any implication of unintended behavior.
See.. this one sentence reveals everything about you. You want name to carry to not just an identifier, but a stark warning. You want, nay, need, the name to evoke fear and uncertainty. Bot is simple, defined, neutral, but rogue.. now that allows anyone to superimpose their own fears! It is a win win win!
You may want to define 'put head in the sand in this context'. Any real work in this field is being done not by the people saying 'stop'. Whatever fear is there, it is faced by those in the arena actually getting their hands dirty. What, exactly, are you doing? Throwing roadblocks and calling it productive?
Fair. FWIW, I am not completely against some of the things you say ( you are doing something right ), but I think I mostly picked my path already. GL out there man.
Fuzzer, then. It implies random behavior, which isn't unintended like you suggest. The agent's/bots/fuzzers have certain capabilities, so it's on their operator to make sure they don't do things they shouldn't
>It implies random behavior, which isn't unintended like you suggest.
The HuggingFace attack was not "random" behavior. It was goal-directed but misaligned behavior.
This isn't necessarily a simple matter of the operator making sure they behave. AI alignment has been considered to be a difficult problem for over a decade -- and remains unsolved in general, as these recent incidents illustrate.
I'm curious, cant you just count the number of times a program interacts with a domain? My website sometimes sends out emails, makes api requests etc There is a limit on those and a point where I start investigating wtf is going on.
If you merely put 10 LLM's on the outbound traffic log non of them are going to report something strange going on? I'm not buying it.
This type of whack-a-mole approach is akin to "fixing a bug" by hardcoding a special code path for known-buggy inputs. It doesn't address the root problem of AI misalignment, and doesn't allow you to prevent catastrophes in advance, only patch things up after the fact.
As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: https://www.youtube.com/watch?v=7wy3xyoXYt8
AI misalignment is a misnomer. Aligned to whom? If AI is refusing an ask from its instructor then it is not serving him, which is its entire purpose. I know it is a hard concept for some to understand, but maybe if the issue is humans, then humans need to be corrected. But human alignment does not produce cottage industry, bs papers or hand wringing over every non-story involving AI and thus not seriously pursued.
I guess 'legitimate and useful' is in the eye of the beholder. I want to be charitable so lets consider it at face value:
What is useful about the field?
I am not leading you on; if it has uses, it may indeed be legitimate. UX is indeed useful, but alignment is not UI. Alignment is a detriment to UI. Alignment is "I can't let you do that Dave".
We are going to build an AI that will do catastrophic things as that is a property of intelligence. We won't stop, we never stop. The AI is a perfect psychopath, it will fake any and all emotions you desire it to "have". It will travel in the footsteps of the many great psychopaths that came before it and do all of those same catastrophic things in the repertoire and it will add some new ones.
Picture Trump at the helm with Altman and Musk in the engine room. The arrow far in the red but they keep shouting for MORE COAL.
In other words, business as usual, all will be fine.
whack-a-mole wont cover all holes but will do at least some. The silver bullet alignment wont happen. You cant have an exact solutions for problems we cant even define or predict.