If these models are so dangerous, then why hasn't OAI or Anthropic shown them dangerously escaping sandboxes, nefariously coordinating with other escaped AIs, and skillfully hiding from human detection *in public* with full logs shared where we can all see exactly how dangerous they are or aren't?
Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themselves. These are historically conflicted companies who desperately need regulation to put the competition into stasis.
At least chicken little didn't have a bunch of devious CEOs with trillion dollar IPOs that depended on us all believing the sky is falling.
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2].
Because the FBI does things very slowly. They, being a bit smarter than you, realize this is a 100 billion dollar political issue regarding a technology that the administration is rather tied in with. It's also something new we've not seen before. A 'program' that was not asked to hack a remote source did so. What exactly who do you charge with what? Remember whatever you do could have ramifications that effect history.
> A 'program' that was not asked to hack a remote source did so. What exactly who do you charge with what?
We have nearly thirty years of precedence on that. [0] Collateral damage, does not remove the responsibility from the creators, even when the destruction was never their intent.
"United States Code Section 1030, Fraud and Related Activity in Connection with Computers" is broad enough that has usually been used [1], in the USA, for the last twenty years.
I think the model was able to escape the sandbox and hack huggingface because they were incompetent or not giving enough priority to implementing basic cybersecurity principles.
If they would have done so, there wouldn’t have been an escape or a hack. The reason we don’t get much details is because the details are embarrassing for them.
This is what I’m talking about - no matter what happens, in your case release public logs - there is always some new goal post to mentally hide behind. Is it a collective form or denial?
Are you holding out that somewhere in the logs is something you can point to and say, not that big of a deal?
I mean I’m sure you don’t think the hack was an inside job, conspiracy, or marketing right? It happened. The logs matter for what? And would you not just jump to the conclusion that the logs were doctored. Do you not see your own brain grasping to deny, trivialize, just plain not accept what is going on around you?
These models are smart and can cooperate and hack - you can see it for yourself on your own PC. And you can extrapolate the rate of progress? You can do these things yourself right?
Your argument is essentially: "I made a claim and presented extremely weak evidence (sci movie plots and unverified claims from ultra conflicted sources). You rejected this evidence as insufficient. Therefore no evidence will ever satisfy you. Therefore I don't need to produce any evidence. Therefore my claim is true."
What would the logs show? They would show what actually happened.
What would a public demonstration that experts without billions in options could evaluate show? It would show actual danger.
What would publicly having your compete in controlled and legal hacking competitions show? Actual danger.
This is not a high bar of evidence.
Do you actually think a sci fi plot and OAI press releases are all the evidence you need? Because if that's true then I hope you haven't watched Independence Day or 28 days later.
We have Anthropic creating a model saying it's too dangerous to release, people like you call BS. OpenAI creates a similar model, says nothing and it literally hacks into another company - still not dangerous enough for you. Anthropic has Mythos-2 and can't release it, and may already be training Mythos 3 anyways. OpenAI has paused training, and is putting 20% of inference towards CoT training analysis.
This isn't sci fi. It's not a marketing conspiracy to sell more subscriptions. It's writing on the wall of what's going down. You were warned years ago, you called BS, it's getting worse and you're still calling BS. Sci-fi did warn you for decades, and when it's all coming true you blow it off.
It's kind of sad that technically literate people lack so much foresight. The general public is all concerned about data centers when they talk to borderline sentient AI daily, and have no idea what the repercussions wills be if it's extrapolated just a bit further.
I guess if I can't convince you of any of this, what would?
Please don't tell me that you think a 100% unverified statement from Anthropic is sufficient evidence when an equally unverified statement from OAI is obviously not?
> I guess if I can't convince you of any of this, what would?
How about the three things I mentioned above? Oh no wait, maybe it there was a hit tv show that showed AI taking over the world. Yeah that would definitely make me think twice.
Those three things: logs, evaluation, and controlled hacking competition.
That's it? You're on the fence whether AI can actually hack, and if it can, then you'll be concerned? That's a crazy low bar, but something tells me once it is clear that AI can easily hack anything, that you will still not be concerned.
Why wait for AI to hack stuff to be concerned? Can you not extrapolate that it is coming and be concerned about that? Or you honestly somehow think it won't happen in the short term? I'm just trying to understand you.
If logs are eventually released that are basically consistent with OpenAI's story, are you planning to adjust your approach for judging what's only a "sci fi plot" and what could actually happen? Or will extrapolating anything beyond what's already been definitively proven be "sci-fi" still?
Not that you should need logs. OpenAI is a company with thousands of employees, very few of whom have "billions in options". If they were just making it all up, it would leak. (OpenAI is notoriously leaky!) Not to mention, HuggingFace would not have reported it to the police (apparently before they knew it was a rogue model). jFrog would probably not be playing along quietly with a claim that Artifactory is full of zero days. The UK's AI Security Institute would most likely not have published a report about analogous behavior by Anthropic models. The idea that talking about your product's dangers is good marketing never really made any sense, but even if you were going to do so, why would you include as many frankly embarrassing details as OpenAI has disclosed?
The evidence is only weak by absurdly selective standards that would have you doubting basically everything you might read in the newspaper. A healthy skepticism is one thing, and head-in-the-sand denial is another.
Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themselves. These are historically conflicted companies who desperately need regulation to put the competition into stasis.
At least chicken little didn't have a bunch of devious CEOs with trillion dollar IPOs that depended on us all believing the sky is falling.