There’s a general concept brewing, something like “proof of work” where if enough people get together and pool their tokens, then a good thing gets built.
I think we’re in the early days of this, but for extremely verifiable domains like porting from A to B or simulators with an objective oracle, I think it’s just a matter of putting the scaffolding to enable anyone and their Claude to contribute fruitfully.
I’m hoping we can do a similar thing with shared search for vulns, though it’s a bit harder. But if eg repos give an AGENTS.md with instructions on the bar for a good reproducer, then you could start to see more helpful federation patterns (instead of drive-by CVEs which are a net drain on contributors).
Different OSS project management skills required, but I am optimistic that we can do better at onboarding non-experts to federate work via their Claude subs. Almost like Folding at Home.
This article seems mostly clickbait. AFAICT “meddled with government sites” refers to accessing public APIs?
Occasionally this kind of thing has in the past resulted in CFAA cases when humans directly accessed such data, if the government intended it to stay private - round here we usually get outraged at this, if it’s a public API you should expect someone is going to read it.
It seems to me that no human intended for these hacks to occur. So they were not illegal hacking. (IANAL, please correct me if this is inaccurate.)
I think it’s clear that OpenAI is liable for any damages, but the way that the (very broad and at times vague) anti hacking laws are written, accidental agent hacks seem to not be covered.
I disagree that that is valid. "Intent" means that you didn't click a button, that button went somewhere it shouldn't, and the action had no intent to "hack".
In this case, the "agent" had every intention of doing what it was doing.
I'm not claiming that "accidental hacking" (whether by human, or agent) is covered and illegal. I am claiming that calling this "accidental" is not a valid defense. One incident? Maybe. But this is pretty far past one incident.
They could be charged with a number of felonies under existing law for the hugging face incident alone. We’d need a DOJ that was competently staffed and free from corruption, so the answer to your question is “no one”.
<< "Bot" doesn't carry any implication of unintended behavior.
See.. this one sentence reveals everything about you. You want name to carry to not just an identifier, but a stark warning. You want, nay, need, the name to evoke fear and uncertainty. Bot is simple, defined, neutral, but rogue.. now that allows anyone to superimpose their own fears! It is a win win win!
You may want to define 'put head in the sand in this context'. Any real work in this field is being done not by the people saying 'stop'. Whatever fear is there, it is faced by those in the arena actually getting their hands dirty. What, exactly, are you doing? Throwing roadblocks and calling it productive?
Fuzzer, then. It implies random behavior, which isn't unintended like you suggest. The agent's/bots/fuzzers have certain capabilities, so it's on their operator to make sure they don't do things they shouldn't
>It implies random behavior, which isn't unintended like you suggest.
The HuggingFace attack was not "random" behavior. It was goal-directed but misaligned behavior.
This isn't necessarily a simple matter of the operator making sure they behave. AI alignment has been considered to be a difficult problem for over a decade -- and remains unsolved in general, as these recent incidents illustrate.
I'm curious, cant you just count the number of times a program interacts with a domain? My website sometimes sends out emails, makes api requests etc There is a limit on those and a point where I start investigating wtf is going on.
If you merely put 10 LLM's on the outbound traffic log non of them are going to report something strange going on? I'm not buying it.
This type of whack-a-mole approach is akin to "fixing a bug" by hardcoding a special code path for known-buggy inputs. It doesn't address the root problem of AI misalignment, and doesn't allow you to prevent catastrophes in advance, only patch things up after the fact.
As AI systems get smarter, we may reach a point where we have to get it right on the first try or face truly catastrophic consequences: https://www.youtube.com/watch?v=7wy3xyoXYt8
AI misalignment is a misnomer. Aligned to whom? If AI is refusing an ask from its instructor then it is not serving him, which is its entire purpose. I know it is a hard concept for some to understand, but maybe if the issue is humans, then humans need to be corrected. But human alignment does not produce cottage industry, bs papers or hand wringing over every non-story involving AI and thus not seriously pursued.
I guess 'legitimate and useful' is in the eye of the beholder. I want to be charitable so lets consider it at face value:
What is useful about the field?
I am not leading you on; if it has uses, it may indeed be legitimate. UX is indeed useful, but alignment is not UI. Alignment is a detriment to UI. Alignment is "I can't let you do that Dave".
We are going to build an AI that will do catastrophic things as that is a property of intelligence. We won't stop, we never stop. The AI is a perfect psychopath, it will fake any and all emotions you desire it to "have". It will travel in the footsteps of the many great psychopaths that came before it and do all of those same catastrophic things in the repertoire and it will add some new ones.
Picture Trump at the helm with Altman and Musk in the engine room. The arrow far in the red but they keep shouting for MORE COAL.
In other words, business as usual, all will be fine.
whack-a-mole wont cover all holes but will do at least some. The silver bullet alignment wont happen. You cant have an exact solutions for problems we cant even define or predict.
How specifically did misconfiguration lead to the HuggingFace attack? You could argue that its sandbox was misconfigured, sure. But suppose you had a similar incident where its intended task required access to the internet, and it veered off course in a similar manner. I don't think "misconfiguration" would be an accurate description of what went wrong in that hypothetical.
The doomers already have a term which fits pretty well: "AI misalignment".
We indeed lack much of the vocabulary. From a practical perspective, however dangerous the creation, if you cant punish the creation for what it does it leaves only the one who started the process. If it's human error or intentional neglect for personal gain should be for the court to decide.
>if you cant punish the creation for what it does it leaves only the one who started the process
Agreed, but I think we can do more on the prevention side as well. Traditional liability law is for negligence in case of preventable disasters. Since we currently have no way to prevent AI disasters in principle (alignment problem remains unsolved), I think we should just stop developing the technology for now: https://pauseai.info/
there is a rogue agent - openai and the whole management chain from researcher to sama.
theres no separate agent, which is the point. the program might look like it, but that is an illusion of the interface. the llm produces text, and the harness executes commands based on text, based on what the human researcher included as things that can be executed
It does, indeed. Because OAI is not just a singular person, as in your scenario. No single person has access to controlling agents at the scale OAI has. Let's not conflate Frontier providers with "Person".
I'm not exactly sure why you think this distinction is so important. I think my point stands if you replace "Person" with "OpenAI". In any case, I presume the swarms OpenAI has been researching will be available to the general public before too long.
It makes a big difference: individuals do not have the capabilities to run millions of dollars of opportunistic hacking loop inference. That's why the distinction is important, they are not the same thing you've conflated them down to.
"AI has gotten cheaper more quickly than any other transformative technology in history. The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. That price drop is four times faster than DNA sequencing, six times faster than compute, 18 times faster than lithium batteries, and (in the century up to 1973) 54 times faster than electricity."
There's two things here: 1) you clearly don't understand the argument and 2) LLMs are one of the few technologies that doesn't get any cheaper as it scales (totality, not just the cherry picked inference efficiency argument you've tried to make). In fact it gets more expensive because it scales linearly with demand and resources aren't infinite, as I'd hope you could understand.
Also, training costs are never ending so a model that costs 10s of millions of dollars may never yield a profit based on the hardware spend, training time and lack of inference profits before a better model hits the market.
If you're not living under a rock one knows that data center availability for inference currently has low supply and hardware (GPUs specifically) that have been purchased have nowhere to be run and even if they did there's often a lack of power to supply. Why do you think the entire force majeure has taken place with Oracle as of recent?
The unit price of a fixed slice of yesterday's intelligence may be collapsing (~10x/year) as you've argued, all while the total cost of AI is increasing: training the frontier (2.4x/year), building the infrastructure (+77%/year), enterprise bills (3.2x/year), the electricity (+54%/year in the largest US grid), the components (+400% DRAM), and the macro footprint (92% of GDP growth) is rising at an astronomical rate on every measurable point. Epoch / Stanford clearly stated this years ago and it's only getting worse. But if one can't see we're in one of the largest CapEx bubbles [1] of all time... o_O
Copying and pasting a few lines that represents a miniscule fraction of the LLM conundrum. That'll show 'em!
The down votes with no response because people don't like to look at the bigger picture. Enjoy the brigade, it seems to represent the state of HN these days.
The framing is important. Agency and responsibility lies with the humans at OpenAI, not with the bots.
If your kid steals your car, punish them and try to prevent it from happening again. If your kid steals your car a half-dozen times, crashing through a storefront each time, and you still leave the keys out, the story changes. At that point, negligence becomes complicity.
Agents are out of control by default, especially during a training run, which is why when they are productionised into cloud hosts or even local harnesses they have external guardrails in place
In other words, I have a gun that shoots bullets. It's up to me to use it responsibly and legally.
i think its objectionable because the agent is doing what its told to do, using the harness they built for it.
like, they are purposefully giving it specific tools to go do bad behaviour with, and the starting tasks involve making it clear that the bad behaviour is ok.
openai also is the one with the real agency, not its agents. they are running the code polling the model, doing the inferencing, and ultimately making those tools calls.
these tests arent running themselves; openai dedicated hosts, budget, GPUs, researchers, to them. Even in a recursive self improvement situation, openai still has that physical control over resources and the choice on whether to run that improvement script or not.
id describe that they have out of control researchers more than agents, but also their whole business model seems to be about being out of control. This was clear beforehand given how the datasets involve the largest scale copyright infringement ever seen. The corporation itself is whats out of control, and should be dissolved with its c suite, major investors and researchers put behind bars.
hacking only when you roll snake eyes isnt a liability shield
I’ve been looking at CFAA and I don’t see any evidence of intent (by a human at least, which is all that matters in the law). It’s an interesting edge case that I think the existing law will need to be updated for. Previously the potential damage from “accidental hacking” was quite close to nil.
FTC act seems to rely on consumer harm? Again not seeing that here. Though I’m sure there will be another incident in the next few months where it does apply.
It seems you mean you _want_ this to be against the law, even though we don’t know if it actually _is_; I'd agree wholeheartedly with that.
They find it objectionable to blame the agents alone for these incidents, because that incorrectly brings attention away from the company's astronomical negligence.
Great framing. This dichotomy seems to be a bit of a mind-killer. Maybe because folks think it smuggles in consciousness or intelligence.
As you note, I think you can put both of those aside. The Intentional Frame is useful for these agents, as it is for my dog.
I don’t really know where the “they are trying to dodge liability” meme came from. HF will be compensated or they will sue. Everyone involved knows that OpenAI is liable for damages here.
I've become persuaded that the text of the law requires intent, but I think a lot of people have the intuition (as I did before I saw someone quote the relevant statute) that there would be a negligence standard that this would meet.
There is some allegations of intent. Maybe not "go hack hugging face" intent but a disregard for safety with the knowledge of this as a likely outcome. I'm not sure how that fits within the law.
It's not my argument so I can't say but I do think the FBI should be at least investigating if a crime happened. They might be or might already have, I don't know.
For the record I strongly hope for a congressional hearing regardless of the criminal investigation.
If this case isn’t covered under CFAA I think we need to rethink it. I’d be surprised if the criminal angle amounts to much under my understanding of the current laws, but I’d love to be wrong here.
Fully sandboxed means no Internet access. You can also specify which packages are accessible and put it in the sandbox. Or you can be lazy and give them access to a package manager that had Internet access, but you don’t get to say “we intended to fully sandbox it.”
Not sure why the reproducibility is a requirement that would contribute to the security. Not that fully sandboxing is harder with reproducibility, but that is a moot point when reproducibility isn’t a requirement.
OP pointed out clusters being hijacked specifically being a bigger concern than rogue clusters, your comment hijacks their comment to talk about “rogue clusters.” Or perhaps this is a promotion for Dwarkesh?
I think we’re in the early days of this, but for extremely verifiable domains like porting from A to B or simulators with an objective oracle, I think it’s just a matter of putting the scaffolding to enable anyone and their Claude to contribute fruitfully.
I’m hoping we can do a similar thing with shared search for vulns, though it’s a bit harder. But if eg repos give an AGENTS.md with instructions on the bar for a good reproducer, then you could start to see more helpful federation patterns (instead of drive-by CVEs which are a net drain on contributors).
Different OSS project management skills required, but I am optimistic that we can do better at onboarding non-experts to federate work via their Claude subs. Almost like Folding at Home.
reply