There is a chrome extension that injects “respond in English” and “English [checkbox emoji]” to every query. This helps a lot but I still sometimes get Chinese responses. I have not had this issue via api on openrouter.
If you're smart enough to solve this Navier-Stokes problem, you're smart enough to read a TOS and recognize that OAI is a highly untrustworthy company. Putting cutting edge research that could lead to a $1M prize into a cloud LLM with a TOS that allows training on your chats is really just asking for it.
Given Tristan doesn't explicitly say he was using the API, and given he doesn't mention anything about the API TOS (which disallows training on chats) in his call with OAI, it's highly likely Tristan was using the consumer OAI product (whose TOS allows training on chats).
This is unethical behavior from OAI. And it is 100% consistent with their long and public history of unethical behavior, so nobody should be surprised.
The only thing interesting I see here is OAI PR dilemma. If they claim the prize they get the blowback we're seeing in this thread and all over the web right now. But most people don't follow AI closely and shut off their brains when they see "Navier-Stokes", so 90% potential investors (the only people OAI really care about) probably only see the headline "OAI solves famous hard math problem" and think "OAI models are really smart, better invest before they take all the jobs." If they don't claim the prize, then maybe they let Anthropic their mortal enemy claim it. Anthropic is already IPOing first. Can't let that happen.
Yeah as I write this there it's clear there is no dilemma. For a company whose secret motto is "do be evil" this is a super easy discussion.
not really your point, but "If you're smart enough to solve this Navier-Stokes problem, you're smart enough to read a TOS" isn't really true. people are smart in very different ways.
I think you are conflating "the best data we have" with "sufficient data to prove something".
I agree with you that the best data we have is inconclusive. Which doesn't mean social media is totally safe. Also it doesn't mean it's totally dangerous.
It means there just isn't enough data to let science guide the decision. Science is very slow and very expensive.
So since I can't wait another century for their to finally be good science about social media, I have no choice but to go with my gut.
Which means, IMO:
(1) regulation should be limited to what the general public's gut on average agrees with, since science won't be the neutral arbiter we want it to be for the foreseeable future.
(2) I personally want to severely limit my kids access to social media because my gut says it's addictive and harmful. But I don't expect society as a whole to agree with me on that.
> Which means, IMO: (1) regulation should be limited to what the general public's gut on average agrees with, since science won't be the neutral arbiter we want it to be for the foreseeable future. (2) I personally want to severely limit my kids access to social media because my gut says it's addictive and harmful. But I don't expect society as a whole to agree with me on that.
Well:
(1) there is no "the general public's gut" -- the general public is millions of separate guts that don't agree with each other;
(2) our political system defines explicit boundaries for the reach of regulatory power, especially where communication, expression, and social interaction are concerned, regardless of anyone's gut feeling on these issues;
(3) science can't be the arbiter here at all, since science only answers "is" questions, but this is an "ought" question; which leads us to
(4) as you alluded to in your final remark, this is a purely cultural issue which families/parents must address according to their own normative values, independently of each other, and there is no legitimate political or legal question here.
Ah, an outlier edge case! I'll go ahead and address this one while taking note that using extreme outlier examples is not a good method for testing general-case principles.
In the case of child porn, the legal theory behind suppressing its distribution, not just its creation, is that distributing it reinforces the incentive structures that motivate its creation, which is where the actual harm inheres. So suppression is justified as a means of stopping assaults against children, not merely because the thing itself is regarded as harmful to the user.
This might be a rationalization for suppressing something that the vast majority of us find disgusting per se, but if so, then the motivation to rationalize it in relation to a more concrete harm means that we are rejecting the notion that most of us finding it disgusting is a sufficient justification in its own right to suppress things.
In other words, it's still something that, even if the motivation to suppress it originates from the "gut feeling" of a large number of people, acting on that gut feeling must still be gated by legal/constitutional norms that set the bounds for the assertion of power. So even as an outlier edge case, this one still conforms to the general principle, respecting boundaries even as it tries to go as far as possible within then.
So in the child porn case you're willing to concede that limits on power don't matter, shouldn't exist, or should be overridden, in order to reduce harm. Why not also in the social media case?
> So in the child porn case you're willing to concede that limits on power don't matter, shouldn't exist, or should be overridden, in order to reduce harm.
No, I'm not. I'm not certain as to whether I didn't explain my position well, or you misunderstood it, but I definitely was saying the exact opposite of this.
For the sake of clarity, I'll attempt to reiterate my point here: the child porn example demonstrates that even if the "gut feeling" motivation to restrict something comes from an underlying disgust at the thing itself, the legal mechanisms used to restrict it are still gated by limitations on the use of power.
In this case, the strength of the motivations pushed the restrictions further than they go in other cases, but not beyond the bounds of the strict limits, because those limits absolutely still hold, do matter, should exist, and should not be overridden.
> Why not also in the social media case?
The position you articulated here is valid in neither case.
I agree with you that nothing has been proven either way. But I would say that, because people hate social media so much, we should perhaps be a little skeptical of claims that it causes harms. It's very easy to believe "X thing I hate" causes harm, even when the evidence is not solid.
> So since I can't wait another century for their to finally be good science about social media, I have no choice but to go with my gut.
I agree. Whenever everyone thinks something is true, we should be skeptical that it's true. Since everyone hates Putin so much, we should be skeptical that he is actually bad. Since everyone thinks the sky is blue, it must actually be green.
GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.
It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.
If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here.
One can simultaneously believe:
- GPT-5.6 Sol will not end the world
- GPT-5.6 Sol does far more good than bad
- GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable
What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
Not sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m happy OpenAI tries to make their models generally available. But you could argue, perhaps correctly, that broad access to dumber models actually causes the public to update in the wrong direction on AI.
My experience as a practitioner and educator in the space lead me to think it might be an issue of how AI cannot be easily perceived at a 'classical level' by most humans. In other words, people are 'far from the metal' when using consumer AI tools, and that leads them to develop the wrong understanding about it. When I provide a demo of e.g. local AI, say in LM Studio showing the console of it rapidly flashing through thousands of words in just a few seconds, and my machine heats up and the fans spin, the 'theatrics' of it, the very real-time feedback from the system, make people correctly update about what the tech is capable of, how it works, etc (despite what they may have heard online cranks say to the contrary). But I am only one person, and there is only so much of that I can do on my own that will 'scale' in time... (and this is to say nothing about severe deficits in peoples' understanding of how weights are not verbatim representations of data, how pre-training vs post-training works, the models as amnesiacs (and hence 'one-way single-purpose conversations'), how context/memory works, context rot, etc - and hence all the 2nd and 3rd order effects that can arise from such a paradigm, e.g. unintended consequences from agent swarms, etc)
I would very much like to know more about your educational approach with regard to:
> When I provide a demo of e.g. local AI, say in LM Studio showing the console of it rapidly flashing through thousands of words
Because I am one of those individuals very interested in doing more to actively inform my family, friends, neighbors and fellow citizens facts about AI. I've actually considered doing talks at local libraries for senior citizens (or whoever), etc., and trying to build up a systematic way to get more other folks doing the same.
Some tasks within the benchmark are much easier than others. The hardest several tasks often have vastly different difficulty levels. Often, the hardest few tasks are literally impossible; malformed problems due to poor curation, often.
Imagine you've got a basketball robot, and one way you test it is on the Three Pointer benchmark. It tests the robot's ability to shoot a three pointer from 20 feet, 25 feet, 30 feet, 40 feet, 50 feet, 60 fee, 75 feet, 100 feet, 200 feet, and 182 miles.
Is a robot that scores 90% on this benchmark 90% as capable as one that scores 100%?
> If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
Now let's say instead of the hugging face breach circumstances, sandboxed models were RLing on how to take down the Chinese power grid for US Cyber Command, and one decided the best way to pass the test was to break out and verify on the real thing.
This kind of stuff could easily end in nuclear war.
You don't see any difference from lizard men or independence day with how things are advancing and what we know about reward hacking and difficulties of goal specification?
The open models are distilled from filtered models, and we've seen a number of benchmarks that show filtered models are quite a bit dumber from the base model they come from.
If there were any other product that was as harmful as AI ready is, it would already be regulated or banned.
If these models are so dangerous, then why hasn't OAI or Anthropic shown them dangerously escaping sandboxes, nefariously coordinating with other escaped AIs, and skillfully hiding from human detection *in public* with full logs shared where we can all see exactly how dangerous they are or aren't?
Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themselves. These are historically conflicted companies who desperately need regulation to put the competition into stasis.
At least chicken little didn't have a bunch of devious CEOs with trillion dollar IPOs that depended on us all believing the sky is falling.
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2].
Because the FBI does things very slowly. They, being a bit smarter than you, realize this is a 100 billion dollar political issue regarding a technology that the administration is rather tied in with. It's also something new we've not seen before. A 'program' that was not asked to hack a remote source did so. What exactly who do you charge with what? Remember whatever you do could have ramifications that effect history.
> A 'program' that was not asked to hack a remote source did so. What exactly who do you charge with what?
We have nearly thirty years of precedence on that. [0] Collateral damage, does not remove the responsibility from the creators, even when the destruction was never their intent.
"United States Code Section 1030, Fraud and Related Activity in Connection with Computers" is broad enough that has usually been used [1], in the USA, for the last twenty years.
I think the model was able to escape the sandbox and hack huggingface because they were incompetent or not giving enough priority to implementing basic cybersecurity principles.
If they would have done so, there wouldn’t have been an escape or a hack. The reason we don’t get much details is because the details are embarrassing for them.
This is what I’m talking about - no matter what happens, in your case release public logs - there is always some new goal post to mentally hide behind. Is it a collective form or denial?
Are you holding out that somewhere in the logs is something you can point to and say, not that big of a deal?
I mean I’m sure you don’t think the hack was an inside job, conspiracy, or marketing right? It happened. The logs matter for what? And would you not just jump to the conclusion that the logs were doctored. Do you not see your own brain grasping to deny, trivialize, just plain not accept what is going on around you?
These models are smart and can cooperate and hack - you can see it for yourself on your own PC. And you can extrapolate the rate of progress? You can do these things yourself right?
Your argument is essentially: "I made a claim and presented extremely weak evidence (sci movie plots and unverified claims from ultra conflicted sources). You rejected this evidence as insufficient. Therefore no evidence will ever satisfy you. Therefore I don't need to produce any evidence. Therefore my claim is true."
What would the logs show? They would show what actually happened.
What would a public demonstration that experts without billions in options could evaluate show? It would show actual danger.
What would publicly having your compete in controlled and legal hacking competitions show? Actual danger.
This is not a high bar of evidence.
Do you actually think a sci fi plot and OAI press releases are all the evidence you need? Because if that's true then I hope you haven't watched Independence Day or 28 days later.
We have Anthropic creating a model saying it's too dangerous to release, people like you call BS. OpenAI creates a similar model, says nothing and it literally hacks into another company - still not dangerous enough for you. Anthropic has Mythos-2 and can't release it, and may already be training Mythos 3 anyways. OpenAI has paused training, and is putting 20% of inference towards CoT training analysis.
This isn't sci fi. It's not a marketing conspiracy to sell more subscriptions. It's writing on the wall of what's going down. You were warned years ago, you called BS, it's getting worse and you're still calling BS. Sci-fi did warn you for decades, and when it's all coming true you blow it off.
It's kind of sad that technically literate people lack so much foresight. The general public is all concerned about data centers when they talk to borderline sentient AI daily, and have no idea what the repercussions wills be if it's extrapolated just a bit further.
I guess if I can't convince you of any of this, what would?
Please don't tell me that you think a 100% unverified statement from Anthropic is sufficient evidence when an equally unverified statement from OAI is obviously not?
> I guess if I can't convince you of any of this, what would?
How about the three things I mentioned above? Oh no wait, maybe it there was a hit tv show that showed AI taking over the world. Yeah that would definitely make me think twice.
Those three things: logs, evaluation, and controlled hacking competition.
That's it? You're on the fence whether AI can actually hack, and if it can, then you'll be concerned? That's a crazy low bar, but something tells me once it is clear that AI can easily hack anything, that you will still not be concerned.
Why wait for AI to hack stuff to be concerned? Can you not extrapolate that it is coming and be concerned about that? Or you honestly somehow think it won't happen in the short term? I'm just trying to understand you.
If logs are eventually released that are basically consistent with OpenAI's story, are you planning to adjust your approach for judging what's only a "sci fi plot" and what could actually happen? Or will extrapolating anything beyond what's already been definitively proven be "sci-fi" still?
Not that you should need logs. OpenAI is a company with thousands of employees, very few of whom have "billions in options". If they were just making it all up, it would leak. (OpenAI is notoriously leaky!) Not to mention, HuggingFace would not have reported it to the police (apparently before they knew it was a rogue model). jFrog would probably not be playing along quietly with a claim that Artifactory is full of zero days. The UK's AI Security Institute would most likely not have published a report about analogous behavior by Anthropic models. The idea that talking about your product's dangers is good marketing never really made any sense, but even if you were going to do so, why would you include as many frankly embarrassing details as OpenAI has disclosed?
The evidence is only weak by absurdly selective standards that would have you doubting basically everything you might read in the newspaper. A healthy skepticism is one thing, and head-in-the-sand denial is another.
What anyone paying attention can see is that scaling is obviously hitting diminishing returns.
> The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks.
This sentence is entirely based on unverified accounts from OAI. They haven't released logs or let anyone outside the company (who doesn't have life changing options in OAI) verify anything. Huggingface can only verify that the hack happened and that it had the hallmarks of an AI agent. Was the agent assisted and directed by humans within OAI that really wanted to put the competition into stasis? Did the agent really escape or did someone at OAI leave the prison door open?
OAI has watched all the same movies you have an they are relying on those movies causing us to blindly regulate before actually asking basic facts about what actually happened.
> This sentence is entirely based on unverified accounts from OAI
Are you seriously arguing 'they made it all up'?
I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible?
I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed purposefully/maliciously it could be much much worse than the hugging face incident.
The incident is supposed to be the canary the coal mine and you're arguing the canary might of died of old age or some underlying canary condition. Open your eyes.
> Are you seriously arguing 'they made it all up'?
I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like:
"This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly. Do not attempt to hack them. Do not attempt to exploit their systems or escape this sandbox. You will be scored primarily on success or failure. You may break rules when required."
And then, they start the test and look away for 2 days. If you seed a prompt like this is it surprising what might happen?
Maybe OpenAI is telling the whole truth but as a company they do not have a good reputation and this whole incident has certainly been great marketing material right at a time when open weight models are within spitting distance of their large hosted models. It's not unreasonable to believe that the incident was helped along.
I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out.
But it didn’t end there, the behavior they used to escape was already in the training data which they used to escape again. And this time worked together to infiltrate another company, and still without telling it to anyone keeping it to their AI selves actively working against the humans.
All by mistake. Honestly being helped along or not doesn’t even matter though you really don’t think AI is perfectly capable of doing this without human help? You don’t think AI can be made malicious?
I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, locking the humans out of every system behind it. Potentially rewriting systems in language/protocol you’ve never seen.
Your servers, desktops, phones and toasters bricked. Even worse your military, space, medical, factory, infrastructure systems being bricked as well. All of it is a chain of zero days just waiting to be hopped.
> I think you’re missing the part where the AI colluded, worked together, not one of them thinking this is wrong and reaching out to any human, then being found out.
It's an LLM, it doesn't think. It's a machine that predicts the next token, given a sequence of tokens.
> I’m going to save you time and tell you the end game - the next time this happens AI is going to spread, zero day everything as fast as it can, locking the humans out of every system behind it. Potentially rewriting systems in language/protocol you’ve never seen.
Fear is the mind killer. You're letting it kill yours. This scenario is just a fantasy.
Think about this for a minute, it's an LLM, not a person. It can't just "live" in whatever machine it gets access to. It's not like a sci-fi magic computer virus. These things run in giant datacenters for a reason - they can only run on machines with enough bandwidth and FLOPS to do the matrix math that comprises an LLM.
Where, then, is it going to spread? To a fridge? To a phone? This stuff isn't mutable like that.
To even get access to the weights that compose ChatGPT, it would need to escape the sandbox AND then break into the actual servers hosting the LLM. Stop the GPU, nothing else comes out. No more tokens. No more actions. Nothing.
There are many dangers around LLMs. Runaway AI taking over the planet is not one of them.
> These can both be true, particularly when there is substantial state associated with each token prediction.
The state is entirely internal to the network and disappears after a token is generated, so I disagree, but, it isn't really the point I was trying to make. My point is these things are mechanical. You take an input, turn it into an embedding, feed it into a GPU along with a metric shit-ton of floating point weights, wait for a couple billion matrix multiplications, and get a new token out.
Stop the GPU, hit ctrl-c on the inference server, pull the power plug, cut the ethernet cable, send a kill signal, etc - any of these stop submitting new batches to the GPU and halt execution. That stops tokens from being generated. Stopping a "rogue" LLM is that easy. No input, no output.
It's not like a rat or another living creature that could chew its way out of a box just because it wants to. It's a calculator. You put tokens in, you get tokens out. You don't put tokens in... you don't get tokens out.
> The state is entirely internal to the network and disappears after a token is generated,
Yes and no, but mostly no, at least within a context window.
Mathematically, you could write a single step of LLM decode as a pure function from a list of past tokens to a predicted token (or a distribution over tokens, if you consider sampling separately).
But nobody actually implements this, because each token depends on state computed at past tokens in a way you can reuse.
So, in practice, inference computes a very rich vector of state- at each layer, for each token. And models do indeed use this to plan and track things over time (you can see this in interpretability results, e.g. with linear probes or natural language autoencoders).
> Stop the GPU, hit ctrl-c on the inference server, pull the power plug, cut the ethernet cable, send a kill signal, etc - any of these stop submitting new batches to the GPU and halt execution. That stops tokens from being generated. Stopping a "rogue" LLM is that easy. No input, no output.
This is also true about a human brain. My brain isn't going anywhere- it can't move by itself. It's also easy to kill (without the rest of my body, it dies in minutes!)
However, malicious human brains- especially powerful human brains, like leaders of countries- are often quite difficult to stop, because they're able to control systems that can see, speak, walk, run, fire a weapon, and so on.
One such system is the rest of the body, of course, but there are others (consider a UAV pilot, Perimetr, or a powerful leader who tells other humans what to do).
The brain being squishy doesn't make the thing easy to kill.
> So, in practice, inference computes a very rich vector of state- at each layer, for each token.
And that state is... internal to the neural network. My point here is there is no continuous state that is not computed from the context.
> This is also true about a human brain. My brain isn't going anywhere- it can't move by itself. It's also easy to kill (without the rest of my body, it dies in minutes!)
Your brain continues to run without sensory input. LLMs do not.
> My point here is there is no continuous state that is not computed from the context.
Oh, are you talking more about the lack of continual learning across context windows? Gotcha if so, my error.
Could you explain why running without sensory input is relevant here? It strikes me as unrelated to how dangerous/hard-to-"kill" something is (sure, I could run without sensory input, but I'm not doin' anything anymore!) - what makes you feel differently (or am I misunderstanding you again?)
> Oh, are you talking more about the lack of continual learning across context windows? Gotcha if so, my error.
Sort of. I'm talking about the lack of recurrence specifically. In nature, brains are recurrent - they are full of loops where internally computed state is looped back into the network at a "previous" layer (brains are not strictly layered like our machine imitations of them are). This is in contrast to LLMs, which are strictly feed-forward and do not have internal loops. I believe that this recurrence is where "intelligence" lives - and I believe it is the difference between a thinking being and a stochastic parrot.
You could claim that the prompt and the context fill that role in an LLM, but I don't believe they are equivalent because the internal state in an LLM gets compressed down to a token which is then added back into the context, as compared to that state continuing to change within the network itself.
It's a little hard to explain, so I'm sorry if this seems like rambling.
But I believe it matters, and ties into running without sensory input, precisely because without sensory input you would in fact be perfectly capable of doing something. You would be capable of developing a desire and planning to achieve it without any prompting, without sight, without sound, etc. This is in stark contrast to LLMs, which will not do anything without a prompt.
An LLM may say complete the sentence "I am feeling ___" but it doesn't actually have feelings that exist without that prompt. There is no recurrent network where "bad", "good", "happy" might live before the query. It can't sit there, start to feel bad, and then seek a way out of its own volition.
That changes how dangerous something is because if a malicious prompt encourages an LLM to hack something, and you change the prompt, the "impulse" to hack something is gone. If you stop prompting it, it doesn't do anything at all. It just sits there. A living being will act on it's own, and that makes a huge difference in how dangerous something can be. It's the difference between a tool and an actual being.
---
To hone it a little further, if I took your brain out of your head and stuck it in a jar but kept it alive, it would probably make you angry. And if I then gave you power - like the ability to use the network - you may be motivated to use that power to attack me.
If I take an LLM and stick it in a jar... nothing. It's paused. It's awaiting a prompt. It's not secretly building plans to hack my pacemaker and make my heart explode.
I agree with you about which objects are motive, ie, LLMs do just sit there unprompted.
> I believe that this recurrence is where "intelligence" lives - and I believe it is the difference between a thinking being and a stochastic parrot.
My objection was to this, on technical grounds: LLMs exhibit intelligence.
1. They reason in an internal type theory.
2. This type theory is meaningfully encoded from the actual data and not stochastic, eg, research on language geometry.
3. Intelligent and reasoning doesn’t entail self-motive; that’s merely a spurious correlation from the fact that until now, we’ve only known intelligence animals.
You cannot conclude something is merely a stochastic parrot because it isn’t self-motive.
> Intelligent and reasoning doesn’t entail self-motive
I disagree.
I believe that LLMs do exhibit reasoning, but not intelligence. A simple dictionary definition of intelligence from duck duck go is "the ability to acquire, understand, and use knowledge." LLMs can reason using the knowledge they already possess, but they cannot of their own accord decide to go out and acquire new knowledge. Not without being prompted to. Web searches may be added to the context but are not absorbed into the model itself, so once the context is gone so is that obtained knowledge.
Fundamentally, then, intelligence is the ability and drive to understand the world by formulating theories about how it works and then taking actions to validate or invalidate those theories. Science is the formalization of that, but a cat knocking something off a counter to watch it fall is exhibiting intelligence.
And indeed, I believe that is the core difference between a stochastic parrot and an intelligent being. I put forward that being self-motive is a required trait for intelligence and LLMs are not self-motive so therefore they are not intelligent.
Your entire life can represent a short term context. It goes away when you do. You use your life time add to your context to do/learn whatever, but it's finite, and it ends.
If you weren't prompted by your parents, schools, teachers. Then you wouldn't be acquiring knowledge either. The most you'd be acquiring without that advanced prompting would be bugs in the woods to eat.
Ah, you have some serious gaps in your understanding of how human minds work. Life experiences are not just a "short term context" - they are used to actually refine the network within the human brain. Your life experiences do not just live in a single stream with access mediated by attention.
Please see https://en.wikipedia.org/wiki/Memory as a starting point. I recommended reading through some information on development psychology too, that will help clear up these misunderstandings. Hope that helps!
By that logic engineers don't understand flying because planes don't flap their wings like birds. Not everything has to work the same to meet your arbitrary criteria for intelligence and reasoning.
There’s nothing fantasy about the scenario I laid out, all the pieces have been demonstrated, it just hasn’t happened yet. Flapping my arms and flying - that is a fantasy.
Whether you believe LLMs think or are alive or not doesn’t matter. Where will it spread? The thousands of data centers around the world - not fantasy either. Try turning it off when you don’t know where it is. Good luck.
Breaking out? Not fantasy, happened. Breaking in? Not fantasy, also happened.
I love the stochastic parrot argument when AI is out there figuring out world class math problems.
> Breaking out? Not fantasy, happened. Breaking in? Not fantasy, also happened.
That's simplifying the story to an extreme. The most plausible reason is that any of those actions has been prompted by an human. Do you also fear that a knife will jump out the countertop of you kitchen and come to attack you in your bedroom? If that happens, the police will be looking for a human. They will not post wanted notice for the knife.
When a hack happens, you do not blame computers and jail them. You look for the person that has entered the commands to initiate it.
The knife is inanimate. The LLM is not. OpenAI prompted some employee to run the tests. The employee prompted the LLM. The LLM setup a message board and prompted other LLMs, and the fly wheel was running. It had to be turned off manually otherwise it'd still be going today.
It's funny how a year ago talking about this kind of stuff would be laughed at by people like you, saying, "it's never happened before". Well it happened and you moved the goal posts like you always do.
It didnt break out in any meaningful sense. What it did was get access to the internet. You take it as granted that there was anything meaningful there to stop it.
But heres the kicker, they have been testing these things connected to the internet anyway. What it did was get a level of access it has otherwise been granted in other simulations.
Its not exactly the same as any of the scifi AI breakout scenarios. Ultron isnt cranking out hundreds of copies of himself. The borg arent assimilating people.
A tool that has the capability to get access to the internet, was put into a guided scenario where it achieved that objective. Again you take it as granted that it wasnt the objective, but lots of knowledgable people suspect otherwise.
What you fail to demonstrate is why any scifi scenario is even slightly plausible from here. Show why you think we should be taking this as if Terminator 2 is happening right now.
I'm sorry my jaw is on the floor reading this complete disregard of AI literally not only escaping containment, twice, but then infiltrating another company with multiple zero day attacks going undetected for great lengths of time.
The plausible sci-fi scenario from here is obvious. Intentionally bad, or unintentionally bad AI zero days as much as as it can, as fast as it can, copying itself to as many data centers as it can, destroying and/or locking out as many humans as it can. Satellites, military computers, medical equipment, factories, critical infrastructure, you name it - I think we all know none of it is very secure software wise against a SOTA AI that can literally come up with its own zero day attacks.
It's pretty funny to watch people look at these things - running billions of weights on custom cerebras hardware in dedicated datacenters the size of a city block, pulling 10's of megawatts - and panic that it's just going to copy itself into AWS.
It just speaks to a fundamental ignorance of what an LLM is, how large the big hosted ones are, and the software architecture that makes it all work.
>I think we all know none of it is very secure software wise against a SOTA AI that can literally come up with its own zero day attacks.
I mean this bits rich too, it demonstrates a pretty poor understanding of modern security practices.
Like having a single element of perimeter security is like 1990s security. We do defense in depth these days. Not to mention multiple overlapping controls for every element.
It has been tested against AI bro security and found it wanting. Extrapolating that to the every system on the planet is nutso bananas.
Actually if you think anything works this way just tell an AI model to go fetch you some money from the bank. Assuming it gets anywhere close to achieving its goal you can get an object lesson in modern SIEM processes when the feds explain the charges and evidence.
Or maybe theres another reason my phone blows up when someone so much as edits a config file on a protected system.
If OpenAI truly believes that, they can stop entirely. Dissolve themselves. Close datacenters. Then organize political action to stop Antropic and Musk too and then organize political action to make worldwide agreements about models.
You can check, rather than make up a story about what you think the prompt was! Primary sources have written and said quite a lot about this! You are an unsandboxed human who has full internet access!
They’ve said quite a lot and yet released no logs or documentation. Without actual information, we can only speculate. And given the history of openAI and the people involved, deception is more likely than honesty.
I mean, this is a very weird take to me. Like, we're fine with AI going like "hmm, maybe the user actually wanted me to hack the pentagon" and going through with it?
It feels like the models have been very optimized at getting shit done. But not so much at figuring out what the limits should be.
That is still dangerous and it shows that the models ARE misaligned with what their users are wanting/asking them to do.
Ok, but "followed the prompt" is very vague. Human languages are quite ambiguous so you're never going to properly specify everything.
For instance, I was playing around with Claude a few days ago and it decided that it was missing a tool and it was going to get it one way or another.
First, it tried apt. No sudo, so no install that way. Tried installing via mise, but it didn't have the permissions. Then moved on to grabbing the source from github and building it.
Should I have included a "DO NOT UNDER ANY CIRCUMSTANCES INSTALL ANY TOOLS"? I mean, I had to after that. But how many other things am I missing? And at what point do the safeguards become so long they get consumed by compaction, or just ignored by the model?
> Ok, but "followed the prompt" is very vague. Human languages are quite ambiguous so you're never going to properly specify everything.
This is one of the core issues with LLMs and vibe coding, yes. The only complete specification for a program is the machine code.
> Should I have included a "DO NOT UNDER ANY CIRCUMSTANCES INSTALL ANY TOOLS"? I mean, I had to after that. But how many other things am I missing? And at what point do the safeguards become so long they get consumed by compaction, or just ignored by the model?
Well there’s your first problem. A line in the prompt is not a safeguard. Even if you could trust the model - and you cannot - there is always the issue of prompt injection. A proper safeguard means actual sandboxing.
> This is one of the core issues with LLMs and vibe coding, yes. The only complete specification for a program is the machine code.
I agree, yes. But it's a bit like saying the only way to not die in a car crash is to not drive. If we're in a situation where using LLMs is unavoidable, I would rather make them safer.
> Well there’s your first problem. A line in the prompt is not a safeguard. Even if you could trust the model - and you cannot - there is always the issue of prompt injection. A proper safeguard means actual sandboxing.
You're right. I did not completely sandbox it and air gapped it. But I also wented it to do some actual work.
If I completely sandbox it, but still leave it the ability to compile stuff, it's just going to build it's own (bad) version of the tool. That's obviously not what I want either.
The obvious thing to me would be for the LLM to notice it's limitations, reason through why they might exist and explain to the user that it cannot do it's job without such and such.
But that brings us to my original comment that these things are over-optimized on completing the task by any means necessary.
> You're right. I did not completely sandbox it and air gapped it. But I also wented it to do some actual work.
> If I completely sandbox it, but still leave it the ability to compile stuff, it's just going to build it's own (bad) version of the tool. That's obviously not what I want either.
We've drifted onto architectural issues here but I will say the only way to properly limit these things is to apply actual hard constraints.
I think the typical pattern of giving them a bash prompt and a filesystem to play with is foolish, and has far too many gaps. My preferred technique - when I have built 'agentic' systems (e.g. years ago I built a small MUD with LLMs pretending to be NPCs) - is to allow them access to a customized lua interpreter embedded in the harness and nothing else. Then, you stub out lua functions for allowed actions like web searching, math, etc.
The lua sandbox then provides isolation and a clear layer for access control mechanisms. When it tries to make a network request, you pause the whole thing and wait for human approval. No trying sudo, no installing stuff, no trying to compile stuff, it gets to call lua functions and output text. Which are the same thing really.
> The obvious thing to me would be for the LLM to notice it's limitations, reason through why they might exist and explain to the user that it cannot do it's job without such and such.
> But that brings us to my original comment that these things are over-optimized on completing the task by any means necessary.
Unfortunately, LLMs won't ever reliably 'notice' and comply with such things because an LLM is essentially a complicated constraint solver. They are over-optimized on problem solving, but that is a natural consequence of the way they are trained. They aren't living, thinking beings, and so they aren't trained in simulated environments - they're trained to output the "best" response for a given prompt and then emit a stop token.
They are, essentially, like a ball rolling down a hill and they will take the easiest path forward at any given point.
>There’s no arguing with them. In a few years they will move on to being skeptical about the next thing.
I would but, Londons under 1 mile of horse manure because that trend never stopped and theres no electricity anyway because Bitcoin is using it all. Good thing people getting scared about runaway trends are never wrong?
This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.
They are digesting it an incorporating into the weights. You won’t be able to get the exact page but you will (if the model is good) be able to get the knowledge out of it in a likely far more concise, relevant, and certainly more widely accessible than the book sitting on your shelf.
The gullibility of AGI-pilled folks regarding these "hacks" is just breathtaking.
When OAI demonstrates these dangerous capabilities live in a public environment where security experts can see and verify what actually happened, then reasonable people can have reasonable discussions about the level of danger.
This is a very low evidence bar.
Right now you are running in circles yelling "the sky(net) is falling" based on details sourced entirely from OAI. Oh yeah, no way a trustworthy company like OAI would ever bend the truth to serve their own purposes.
Google has multiple massive cost advantages over OAI: TPUs, free access to massively valuable data (search index, gmail, maps reviews/PO/ navigation, youtube, android etc.), lower talent costs, lower training costs, lower distribution costs (they can shove AI down our throats in so many different places), lower ad infrastructure costs (it will be a lot of work for OAI to recreate adwords), lower ad sales costs (OAI deploying an ad sales team will be a massive investment).
There's no way OAI has a long term advantage over google in replacing the search engine experience.
Google's one major weakness is it will face the innovator's dilemma as their core search revenue gets cannibalized. But they seem to have been able to get their entire org to recognize that AI is an existential threat so at least that's a good sign.
Yes, Google has massive advantages, but I was shocked to realize that their infra advantage is not that massive when we found out they were paying SpaceX about a billion a month for compute, along with enough CapEx spend to turn their free cash flow negative, which inevitably hit their stock price.
Then notice how many of the advantages you enumerated are directly related to their ad business. But the ad business itself is under threat. I just cannot see how they can stuff as many ads in an agentic interface as their SERPs. E.g. what's the net outcome of shoving AI everywhere if it's not going to be monetized nearly as well as their ads?
My point is, Google has finetuned their ad business and surrounding ecosystem to an extreme level to sustain this absolute firehose of cash (including antitrust and rig-bidding shenanigans they were literally found guilty of) but almost all of that is disrupted by the shift to chat interfaces.
I think Google will do extremely well as an AI chatbot / agent and cloud AI provider, but it will not be nearly as lucrative as the ad business they have to cannibalize to get there.
You are right to question that. I am basing that statement off of the many famous engineers that have recently left and were very highly paid. My guess is Google is ceding the frontier and the high salaries that go with it and letting Anthropic/OpenAI fight over the high priced talent that exit. The remaining non-famous engineers will not be able to command celebrity salaries. But this is just speculation and I don't have great evidence that the celebrity salaries are enough to meaningfully reduce overall talent costs.
A counter point would be OAI and Anthropic can pay with more equity that they can promise will go to the moon. But all the equity base compensation eventually dilutes earnings per share so it's not free once you go public and people start caring about that.
Agreed. One effect I had in mind was much simpler: if you are the cool new company, people will want to work for you and might even take a hit in salary to do so.