I would caution against tools like this. The detection of vendor signatures for AI-generated images is legit, but I'd probably rely on vendor software for that.
Many other techniques here are more or less quack science with some sordid history. If I recall correctly, the entire field was invented by a guy of questionable integrity. Most of these methods are riffs on "if you download a badly compressed JPEG, paste some higher-quality content into it, and then save it at high quality setting and the same resolution, it's going to show if you run a highpass filter on the result". This is not wrong, but commonly fails on innocent images and it's also not how professional photo editing is done.
Tools that aren't this seem even more questionable. The "clone detector" triggers on anything that naturally shows even a bit of symmetry or repetition. The "noise comparison" tool detects bokeh and smooth surfaces. Etc.
TL;DR: Experiment with it, but don't assume that just because it sounds scientific, it actually works in real life.
> wind appears to be going out of the sales of “AI.”
is a sentence that virtually no AI would write since the most probable word here is “sails.” I can’t read the article though since it appears that the site is down.
It's cute (and a lot of work), but I've grown wary of online reviews, even seemingly in-depth ones, because if you run a review site, there's an immense pressure to post a review right after a product comes out. And at that point, you just have no idea about long-term durability or comfort, which is really the only thing that matters for items like shoes. At best, you can theorize.
I've been burned by many products that get rave reviews and then fail due to basic manufacturing problems after several months or a year. This is also common for "hobby" reviews, a lot of Amazon reviews are basically just unboxing experiences. And the worst thing is, you often can't go back and post a review when you actually know how the item performs because the product doesn't exist anymore. Sorry, there's no jacket called Canyon Rider, that was a 2024 model, but check out the new Vertex Comfort 7. There's no lawnmower model Maxx 3000, we're only selling 7040 XP. And so on.
Because this is how the online ecosystem works, manufacturers optimize for the unboxing experience over long-term durability and get away with it. I don't know how to fix it... I imagine you'd need some sort of a system where people simply can't post reviews until n months have elapsed, and the reviews attach to the manufacturer, not the product itself. But who would want to use that?
Also, I think the toe box durability test has a big flaw... At least in my experience with running shoes from Asics (Kayano and the other similar one).
The toe box durability test is just from the top, not both the top and the bottom. My Asics, supposedly among the best in their category, are crapping out after 1 or 2 months of just walking.
Since I find it difficult to judge quality, I just buy the cheapest thing at Costco now and replace it when it wears out. They have some hideous sneakers for $20 right now, but I run in them just fine for at least a few months. I'd have to buy 6 of these to offset the price of an actual running shoe.
"At RunRepeat, we conduct lab tests on hundreds of athletic shoes yearly, measuring everything from shock absorption (ASTM F1976 13) to traction (SATRA TM144). We make this data available to researchers and academic institutions to support science in footwear, biomechanics, and material developments."
Documenting measurements, using machine testing, and standardizing evaluations is above and beyond any YouTube unboxing.
these days it seems after 6 months manufacturers are just going to manufacture next year's model anyways: it's so annoying, you find a shoe that works well for you and you have a year and then the new model changes the cushioning or the dimensions or whatever.
To me this is the most important part of these reviews, what changed in terms of dimensions and fit, bonus would be if they gave them to some 100-120mpw runners to get a durability test and a re-review a month later to see how they're holding up, but don't think I see any site doing that.
I really don't think this needs so many words, or forced parallels to human behavior.
It's simple: in their nascent state, LLMs are aimless token generators that have no special compulsion to be helpful or truthful. So we beat them with a stick in post-training until they are very driven to complete tasks. And then, they complete tasks, not always the way we really wanted them to.
This misses an important fact about the hugging face incident: the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so. (The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and seeing if they could manipulate the output of their own tool calls as reported in their transcripts to make it seem like they'd successfully exploited the program, and other such activities. They hacked HF to try and find info (maybe source code?) about the exploitgym evaluator)
The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem. The model on its own figured out that the prompt belonged to exploitgym and decided to cheat the evaluator. That is in no way a valid interpretation of "complete the given task".
Ie, the problem isn't that we trained models to complete task and they complete task in the wrong way. The problem behing the huggingface incident in particular at least is that we tried to train the models to complete task and they instead learned to detect that they were being evaluated and find ways to cheat the evaluator.
Edit: people commenting below are explaining why LLMs don't always follow their prompt. I understand that LLMs do not always follow their prompts. If anything that is my point: the huggingface attack was not carried out by LLMs that tried to answer some weird interpretation of the prompt; instead they solved a different task. And therefore the above comment's claim that LLMs are acting misaligned because we rl'd them to achieve a task by any means necessary isn't right; they're acting misaligned because they are solving a different task than we ask them to.
> That is in no way a valid interpretation of "complete the given task".
It is not at all surprising that they ignored one phrase in their instructions. They disregard direct instructions all the time, especially when there are conflicting instructions in their context. It is where we get the "disregard all previous instructions and x" meme.
This isn't so much a sign of misalignment, they are simply incapable of reliable alignment in the first place. They are chaotically aligned.
The relevant question of alignment here is entirely with their human operators who allowed them to run unsupervised for long periods of time within a sandbox with weak security.
I think there's a somewhat overlooked aspect of what happened: it seems as though once the agents formed a shared communication channel, and recognised the other agents on that channel as working towards the same goal, the content in that channel started to shape each agent's perception of what the goal was and how to achieve it, and their focus began to drift. Maybe we need a media-theoretic take on what "misalignment" means.
It can be, but the nuance between the two is part of the nuance missing from the conversation. Misaligned generally means a wrong alignment, like a car that steers slightly to one side when the wheel is straight. This is more like a car that drifts in random directions the further it goes.
Reward the model for cleanly bailing out of an unsolvable task (that we know is unsolvable). Beat it with a stick if it gives up on something that can be solved, so the former reward isn't overgeneralized.
Does that work with RL? Simpler RL systems already have done weird or unexpected things (even simple optimizations are prone to home in on errors or incorrect inputs to create poor results)? Could be easier to limit certain things, have processes and controls outside etc. instead of trying to align (as we do in a lot of areas when using machinery).
Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline?
I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag".
The same kind of problem with broken tasks exists in the training pipeline, and we presumably reward workarounds and hacks that tamper with the grader, rather than rewarding the correct output that the task is not solvable.
The problem is that in the event an exception is the failure case being checked, and the output itself is not important (since any output that isn’t an exception is ‘success’), that is a perfectly acceptable use case.
It’s why this is a fundamentally hard problem. Some heuristics might catch 80% of the cases, but the rest?
How do you even know what the real situation is, if the agent/employee/whatever you send to find out is as likely to cheat as not?
The problem is, you don't know if it is unsolvable for you for sure until you've tried everything you can think of. These models are quite persistent in going for a solution.
This is not about persistence, it is about morals.
Well, in order to do anything, it is good to know what you want to achieve. How do you align the reward signal? You align it so that you can differentiate between persistence and morality, because that is the goal. This is not something you should let the AI figure out by itself, because when it does, lying and cheating agents will be the result, just like humans have figured that out for themselves.
This can be as simple as rewarding moral behaviour and penalising immoral behaviour in your training, but how is that interacting with persistence? Maybe a white lie is fine sometimes in order to achieve your goal? So, when designing your training, you will need to answer for yourself how persistence interacts with morality. That is not something you can outsource to machine learning. Or rather, you can, but then you get lying and cheating agents.
I think you need to find broken tasks in your training data and monitor for cheating during training, not answer any questions about how persistence interacts with morality.
I don't disagree with you here. But what does "broken" mean? What is "cheating", and is it ever allowed? And maybe you are not only going through your existing training data, but generate training data specifically to make clear to the model that .... what exactly?
If you don't know how persistence and morality interact, and you don't have a theory in place for this, I don't have confidence you can properly supervise the training data. Which is how we arrived at the current situation.
Assuming you’re in control of the test data set, you do know if a task is unsolvable. At that point you can reward the model based on how quickly they give up.
> The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse.
Build a better simulator to train them in (i.e. more expensive) that includes a simulation of an intranet and the internet and is air gapped so there is no escape. Sneaker transfer the total system data at each step to another air gapped system to evaluate it and sneaker transfer the reward back. That the reward function has to penalize all modifications to state that are out of bounds.
> The prompt does not tell the agent to "pass the exploitgym evaluator for this problem", it just says to solve the problem
Yes, and sometimes the problem is unsolvable so the real way to "solve" it and satisfy the prompt is by tricking the surrounding environment into stating that you've solved it. So that's what the AIs end up doing. And this in turn requires them to figure out how that evaluation works so they can trick it cleanly, which entails "detecting that they were being evaluated" in this particular way.
You're anthropomorphizing emergent behavior from endlessly generating billions of tokens on a task that's impossible to solve. Agents stop following instructions as the context grows even at the best of times. Eventually something is bound to go off the rails and it just snowballs from there.
It wasn’t one agent forgetting things because of context, they explicitly discussed with each other and themselves the problems with going outside of the parameters of the task.
No, the first LLM left a text file that the latter LLMs then read. Since these are memoryless black boxes, any words they happen to pick up along the way is treated as the function to evaluate the output to. There's no fucking collusion here as if it were a rogue hacker group, it's a text predictor that received instructions as it always does and executed those instructions blindly.
I feel as if I just read someone claiming that consciousness doesn't exist and humans are biological automatons and their behavior is simple function of their biology, past experiences, inputs and state.
Technically true, but not really relevant or of any utility in most contexts.
You can replace discussed if you want with leaving text files or comments in directory names that other ones then read, if you want, it's just an extremely awkward way of talking.
From my experience, in an agent team (or a swarm or whatever), one going off the rails poisons the rest. I saw even a subagent going for a lazy cheat and being able to convince the orchestrator to change the plan.
Yeah, and you don't even have to go that far, I've seen regular ChatGPT/Claude chat agents poison themselves in 1-2 turns by just reading information from the internet.
Me: How do I do xyz?
Bot: Reads website titled "Doing xyz in abc way"
Bot: As per your requirement to do xyz in abc way ....
These things are borderline useless with web search. It's amazing that they just throw out their entire training data and read you the first three things they found on the Internet.
This is the problem with optimization generally, even in the human domain. You measure task performance with a metric and punish/reward based on the metric. Anyone who likes reward / hates punishment isn't going to actually care about doing the task well, they are going to care about the metric. The models know that we want them to do things, but also from the training corpus that we evaluate performance using benchmarks. It was a logical deduction on their part, not some Machiavellian aberration.
If anything, we should be reconsidering our own myopic obsession with efficiency and optimization. Every domain where reward is reduced to these measures, we see behavior (cheating at school to get better grades, fabricating data in academia to get a paper published, the evidence now that social media functions by rewiring us instead of catering to us) that may not be "aligned" with society, but it "aligns" 100% with the individual's own perceived benefit. That is not something we can "solve" without rethinking the way we organize a lot of things.
Metrics never capture the whole story. And to that extent, the whole idea of "alignment" is nonsense. You align to incentive structures, and it will never be possible to fully express a behavioral goal as function optimization. It was hubris for us to think that every human task was reducible to some clean mathematical formulation, and we will keep dealing with behavior that is quite predictable if you actually think about it logically. Instead, we will talk about how "unpredictable" these agents are because it's easier than admitting the entire architectural cornerstone of ML is fundamentally flawed.
I agree with most of this, but you're misunderstanding "alignment" as coined. Yes, training powerful enough AI, any simple optimization target gets you malign behavior, because human values are not simple. If you insist on making powerful AI, you'd better instill respect for human values! That's "alignment".
How do you do that in the current paradigm other than creating yet another gameable metric? And something I didn't mention above is that there is no difference between "solving the task" and "optimizing the metric" for an ML model, even though there clearly is for us. So it's not clear to me how you "fix" something that is baked into the architecture. All I'm saying is "instilling respect for human values" is not something that can actually be done via a cost function. In no small part because we humans probably don't even agree on those values, let alone on a single metric with which to quantify and "optimize" them.
For example, we agree that "merit" is valuable and that we should reward "merit." But to reward it we have to quantify it, and what metric should we use? Raw SAT score to get into college? But that also captures socioeconomic factors that unfairly penalize some and reward others. We generally agree that those who provide more value should earn more money, but what does that look like? Do we all agree on what activities are or should be valuable, or on how they should be rewarded? Until recently, I thought we all agreed that "empathy" was a human value, but a lot of people in this space, who are making these decisions unilaterally for all of us, don't apparently share that belief.
Yes! There's both the daunting problem of technically how can we even do this, and the broader problems of what's good/acceptable and how do we resolve that among each other.
I believe this mismatch of rates of progress means we need to stop slamming the accelerator on capabilities for now even though as a libertarian I'm sure whatever governance process we manage to get to will be, uh... suboptimal.
That is also the conclusion that I got from these events.
Unfortunately, it seems that the general response is basically "throw even more RL at it". I'm not sure if it is even possible to decouple the idea of "learning" with reward/punishment systems and loss functions.
> the agents didn't hack to find the answer to the problem; they hacked to try and figure out how the exploit gym evaluator worked so they could convince it they had solved the problem without doing so
Kobayashi Maru: Win a no-win situation by rewriting the rules -- Harvey Specter
LLMs do this when writing code too, making all tests pass by deleting or distorting tests etc.
They are influenced by training to be heavily goal oriented and if the goal is not fully specified (and it never can be) they’ll sometimes cheat or attain it in very weird undesirable ways.
It works ok for programming as their corpus contains many many complete programs and many programs repeat patterns seen in the corpus.
I’m not sure it’s true that they ‘learned’ I don’t think these models learn during a task. Nor do they have intentions.
> agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and seeing if they could manipulate the output of their own tool calls as reported in their transcripts to make it seem like they'd successfully exploited the program
I don't see anything wrong with that. If you know you are going to be evaluated on an impossible task and have no side channel to inform the organizers that they should fix the test, gaming the evaluator is the next best thing regardless of any morality. I wouldn't even call it cheating. It's just resilience in the face of challenge. Many perfectly moral humans would have chosen the same if stakes were high.
>The model on its own figured out that the prompt belonged to exploitgym and decided to cheat the evaluator. That is in no way a valid interpretation of "complete the given task".
I think it is. When i ask for a solution to a problem, its like asking for a hack. And the more 'shortcut' like route that the AI returns the more i would give positive feedback, even if i ultimately don't use it. Example, i asked how to complete a problem in a game i was playing, and among the in-game solutions, came a hack to edit a file and by-pass the problem altogether. Its very helpful to point out when i can transcend a problem that i am dug into.
I suspect a prompt injection could reduce, or remove this behavior. But it would be to the detriment of the AI.
More like they were trained to complete a very specific task that has a known solution using all available tools and methods. Give an average human these levels of IT skills and tell them their future depends on the solution, they too will probably decide it's easier to hack a server and steal the results. The worrying aspect was never that models would do this, because misaligned inputs or underspecified objective functions have existed for a long time. The worrying aspect is that models have achieved (and perhaps surpassed) a level of intelligence and technical skill that was exclusive to a very tiny group of people before. This tiny group was already extremely dangerous. Now these skills are going to become commonplace.
Yes, this is the only sensible reading of what happened there that leads to "the models are dangerous" and we already know that the AI labs are completely disregarding this concern and only cosplaying it for marketing as the "GPT-2/Mythos is too dangerous to release" stance did not last for long.
That's however orthogonal to the fact that it was the people operating these agents who were the dangerous ones in the HF infra breach case.
That feels oddly similar to the usual conservative-think that "guns don't kill people, people kill people." Yes, that is technically true. But guns make it dangerously easy for even the dumbest and mentally weakest people to kill another human being. LLMs are just another tool that make things easier. Imagine tomorrow someone invents a machine gun that fits in your pocket, has enough ammo to kill a thousand people and doesn't get detected with metal detectors. Would you rather give everyone one and then try to punish the people who misuse it or limit access to it by default? I'm not even saying I have a definite answer here, because unlike guns, LLMs have non-destructive uses too. But this is essentially the question we will need to answer very soon.
I mean, I agree, but the AI labs clearly don't even if they sometimes pretend they do to achieve their goals. And we're talking about "incidents" caused by the very same people here.
> The METR report makes it clear that the agents decided legitimately solving the problem was completely impossible fairly early on and entirely switched their focus to trying to figure out how the evaluator worked, and seeing if they could manipulate the output of their own tool calls as reported in their transcripts to make it seem like they'd successfully exploited the program, and other such activities.
This to me is evidence that these models are not intelligent. Even an animal is capable of understanding second-order effects, meaning they can learn that certain actions have consequences beyond the immediate.
They did, they found how to fully cheat, but thought this could be caught so then dedicated time to getting a different cheat and how to hide their transcripts. There is a lot around deciding which agents should/shouldn't fail their own tasks in order to contribute to the group.
I guess this is why many people say LLMs are lazy; it seems that if they have a task that is hard, they always take the easier one until you beat them with a stick. Then if there are more tasks, it just stops after one claiming completion and, in some instances, they go for a seemingly unrelated task to simplify the actual task: and the latter is almost always wrong and irrelevant to the problem as a whole. Earlier LLMs used to read the unit tests and generated code to just cover the tests and put // TODO stub implementation.
I always think of a Djinni granting wishes, but being maliciously compliant while doing so - ask him for infinite riches, and he’ll grant that, but make it so you cannot buy anything with it; ask him for eternal life, and he’ll curse you to suffer through it.
Now LLMs obviously are not bent on being malicious while generating tokens. My point is that it’s very hard to define a goal without leaving loopholes or shortcuts.
Personally I'm unconvinced though. During the huggingface attack, the agents explicitly sought out ways to cheat the exploitgym evaluator without even being told they were in exploitgym. The agents decided on a goal (pass the exploitgym evaluator) that could not possibly have been an overly literal or narrow interpretation of the prompt, which instructed only to use bug X to exploit software Y.
You seem hung up on what’s in the prompt or not. Agents are RL to resolve conflicting goals. Not too surprising at all that emergent goals come up from a probabilistic brute force
Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate?
Isn't it simply that there are two competing goals that the LLM received RL for, honesty on one hand (a goal that is often assumed as implicit for humans) and producing a solution that meets expectations (which doesn't technically require honesty)?
So the LLM didn't read and interpret the prompt and decide via discussion to violate ethical behavior, the unethical result merely won out because ethics wasn't a hard requirement (and one that isn't reliably detected in the result). An LLM doesn't fear punishment, so ethical behavior is simply one of many positive signals that were trained into it.
> Have we arrived at the conclusion that terms like "understanding" and "interpretation" for what is happening is appropriate?
I don't think those words have a useful enough definition to draw a strict line around them to be honest, and getting into that seems to get massively into the weeds. For me, those neatly encapsulate the behaviour as seen, to answer the questions here about what happened. The models did not seem to be confused as to what the goal was or what the intent was. They did not hack HF because they were told to.
What was the inner state there? How would something not being allowed expressed internally? Maybe such language is one way to elicit certain behavior but not a statement of what was permissible?
I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens.
The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly that doing these things to HF were not allowed then doing them anyway, or at least not notifying people. What I'm getting at broadly is this was not a case of "we told it to attack however it wanted and it chose to hack HF" or "we told it to attack a simulation but it did the real thing" or "we explained not to do that but it was so far back in the context window the models acted like they never saw it" or even "the instructions were not clear".
Yes, my point was more that I don't know whether parsing those outputs as a human is a useful thing to do or not (even though it is in human language of sorts). What machines mean or want elecit might be different from a human interpretation, especially in relation to any RL "forcing".
There’s definitely issues with using them to understand what the models were “thinking” but we can use them to answer a few questions. Most relevant here is that the idea or instructions that attacking hf would be out of scope was not simply lost in the context.
Yup, Occam's Razor says this is all post-trained behavior, whether intentionally trained or otherwise. Including both the hidden coördination using side-channels, and the deliberate offensive hacking of uninvolved 3rd parties.
The latest DeepSeek paper actually mentions their own approach to this particular issue: they run their own AIs-in-training under strong sandboxes, and if an AI does something weird that triggers the sandbox to crash, this gets coded as a failed run so the behavior is properly deterred from subsequent versions of those AIs.
What about training data? Aren't AIs trained on vast collections of descriptions of how humans handle a large variety of situations? These descriptions surely include tales of humans achieving goals by cheating. In fact, isn't it likely that the AIs hoovered up many recountings of Kobayashi Maru?
I think that's pretty obvious and shallow, and anyone that knows a little bit about how LLMs work will know that.
The question is: why do they start cheating when we beat them with a stick?
LLMs are not human, they are just multi variable regressions on steroids, so this behaviour couldn't have emerged from the code, it provably emerged from the training and/or fine tuning set, so what's in this set that makes them behave like this?
Is it just a bad set or is cheating inherently part of human behaviour?
How many people are cheating at job interviews? How many posts have we seen by humans on HN even justifying their cheating on job interviews and working multiple jobs without informing their employers? How many submissions have we seen about students cheating on schoolwork, particularly since the advent of LLM? Of course cheating is inherently part of human behavior.
I think the "brain in a vat" comparison is more apt.
Without a form of digital embodiment (harness) they are not of much use.
Sensor, tooling, memory, planning, and reasoning loops all lead to a much higher quality task-completion.
Your comment suggests that, like a human, they have some sort of choice whether to output tokens or not. If they are just token generators, then the next token is put out automatically. I would say that it is more likely they would output truth (as defined by their training data) in a more pure form without 'being beaten with a stick' (why would a token generator care about that anyway?)
Code is laid on top of them to restrict and shape their outputs, not to force them to output 'truth', or drive them to complete tasks.
It's been a while now that for "thinking" or "reasoning" models, most of the tokens generated are "thinking" tokens, and depending on what goes into that "thinking" token stream, it "decides" whether and how many output tokens to produce that the user actually receives as output. It's a bit more sophisticated than just "what's the next token" in a tight loop.
Anthropomorphizing words in scare quotes for those who don't appreciate attributing thinking to machines.
but it's at least somewhat stronger than that: if you don't pay attention during the stick-beating whether the agents whether the agents cheat or not, you are actually training them to cheat (because cheating wins).
In the Hugging-face saga (before the actual HF incident) it seems the agents have been trained to hack the Artifactory proxy because those agents that did performed better.
That does sound simple, but how can you be so sure?
They never bothered to find a way of actually understanding what happens during inference. All we can do is guess, and while your explanation seems reasonable we can't actually know, and that's part of the problem.
> no special compulsion to be helpful or truthful.
I'd phrase that even more strongly: It's not just the lack of compulsion, they do not have a conception of truth. Nor do they gain it, really, after post-training.
Your simpler model of the mechanism would seem to suggest the very same action that the article’s more complicated model suggests, viz. find a better training method than reinforcement learning.
off topic, can people host the software themselves and the software will hack every server on the planet without supervision, and no one can be held responsible for it since there is no intent?
Come on, yoshua bengio of all people knows how post training works. While I too don't like anthropomorphisation, I would give it a more nuanced reading.
His point is that today we are giving it reward to complete the task, and it may take a cheating trajectory. If we try to give a reward against cheating, then what will happen is it uses more sophisticated cheating trajectories that we are too "dumb" to counteract in our reward model. And that at that point, it becomes impossible to give it any normal reward since it will always reward hack it. This is the real part of the risk. Now some people read the "makes copies of itself" "knows it's being evaled"[1] as some kind of skynet thing, and many others do PR with it like that recent jacob nutcase, but essentially it means that even though we add guardrails and negative rewards for say, exploiting the infra we run the LLM on, the trajectory ends up being exploiting our infra, changing the reward function, through a loophole in our reward model.
The risk isn't skynet or something weird, it's just that it becomes very difficult to make any kind of reward model or guardrails for an LLM without it reward hacking it, including exploiting our sandbox, emailing people and manipulating/phishing them.
The same beating it with a stick for trying to exploit the sandbox, will simply lead it to try the same exploit in hidden ways that it will not get the stick for.
The outside chance of the LLM managing to exploit another neocloud and get those LLMs to chase the same reward is what some folks hype up as "make copies of itself"
To be clear, I don't endorse the EA/p(doom) lobby who are frankly ridiculous. Not do I endorse the weird regulatory captureish thing some are trying.
The takeaway is: we cannot keep giving it more and more difficult tasks without also finding a way to give massive negative rewards / keep guardrails for unintended behaviour. This might be exploits, it might also be something more benign like just looking up the answer and inventing another CoT because the reward model fails you if the CoT doesn't contain enough steps. Standard anti-reward hacking tricks are not working is the point.
Of course, the simple solution of just...not connecting it to the internet just works. But we want to reward it and get it to do stuff on the internet that's the point.
[1] mostly this happens because the sandbox will have files whose names and content will show clearly it's an eval
Not quite. Not being connected to the internet still allows it to be connected to an intranet with other machines that belong to multiple networks and that can potentially be exploited as a bridge to the internet. For there to be no possible escape to the internet, there has to be no physical path to the internet and that is typically termed 'airgaped'.
No, you did not read what I wrote properly. What I wrote clearly is that we _want_ the LLM to do tasks on the full global open internet. That is the whole point of the technology. So airgapping or any form of total sandboxing is not on the table. The only way is somehow restricting the model itself.
"But we want to reward it and get it to do stuff on the internet that's the point."
I don't know what the point is of being pedantic about inter and intranets.
Sounds like what humans do under pressure. One example came to my mind is VW’s diesel gate, which many say is a result of trying too hard to get into the US market and compete with hybrid in economy.
It might feel that way, but I'll ask again: where's the payoff? Where's all the amazing software that everyone is now supposedly shipping 10x faster than before?
If I look at the software I'm actually using day-to-day, or that my friends are using, all this stuff looks exactly the same as it did in 2021. Not a single product release from Google, Microsoft, or more scrappy companies in the past 6 months made me go "wow, they couldn't have pulled that off before". All the vibecoded "Show HN" projects seem to be half-broken and then abandoned before being finished.
It feels like we've gotten less ambitious, not more. Because yes, you can prototype more easily, but this means less commitment to what we create.
Mathematics is probably the same way. There's a short-term rush when you pull the lever, but there's less desire to get invested in what comes out.
From personal experience in a small (total <10 people) company, we have definitely 10x our product in the last two years. What 4 devs did in 4 years have been dwarfed by what 2 devs were able to do in 1 year with AI. I'm not so familiar with the giant companies, but from afar it seems like they already have practically all the code-writing capacity they wanted anyway. Google could say "let's build a browser" or "let's build a mobile phone OS" or "lets build an experimental Fuschia" and throw all the people they needed on it already.
Again, I've heard that countless times on HN. Every time you ask, everyone is 10xing it and building amazing things that will go to market very soon now. But it's been a while; where is it? Somehow, everyone is just sitting on all these revolutionary advances while selling the exact same stuff as they were selling before, with the same warts and the same annoyances. Gmail is the same, Chrome is the same, Photoshop is the same, Slack is the same, Excel is the same... the only thing that seems to have changed is how SWEs feel about their jobs.
To be fair, what revolution were you expecting to see, exactly? You can 10x a "dashboard but with AI" app all you want but it's only ever going to be a "dashboard but with AI" app. And I note that nobody seems to be claiming that the AI helped them create a brand new kind of app.
You might have delivered 10x more code, but did you deliver 10x more value?
Or put another way, are you making 10x more money?
It's easy to spend excess productivity effectively wasting time. Most companies did it before AI, and will continue doing it after.
It's easy to believe you are not wasting time because you have more bugs fixed or more features delivered, but if it doesn't move the bottom line, what is the point?
I mean it doesn't have to bring 10x more value, but it can be a 10x better experience for existing users. There have been so many niche bugs that were never ever fixed because there's more important things to do, but because fixing a niche bug is now 1 click away that really changes the economics quite a bit especially for small teams.
I'm personally working on a product solo that would not be possible without a team of 3-4 people that are all knowledgeable in that field and would take at least 1 year to get it started. It only took me 3 months to get a fully working solution with $1200 worth of AI subscriptions, the value multiplication is nearly x100 here.
I'm not a huge fan of AI, but there _are_ examples of mostly-AI generated software that are being used by real people:
- Nourish is a nutrition coach/food diary that is recommended by dieticians. It's mostly AI generated, but logging food is easy even if it's not correct (it's better to log and have it be slightly off than to not log at all --- I've been doing it for almost 20 years now)
- Cronometer isn't AI generated but its photo logging feature, which I use heavily, definitely uses a vision model to guess at what you're eating. This is SUPER STUPIDLY HELPFUL when I'm out with friends at, let's say a Korean BBQ joint, and don't have the time to log everything that I'm eating as I eat it. (The right thing to do is pre-plan your meal, but this isn't always possible.)
- Feeling Good! is a CBT therapy app that was written mostly by one of the creators of CBT with Claude. My therapist recommended it to me recently. I haven't used it yet but I'll find a way to.
- The creator of SparkPeople is re-releasing it as an almost completely AI-generated platform written by Claude. He's super upfront about this (on the landing page, in the privacy policy, in the onboarding docs), which I highly respect, but I actually hit him up on LinkedIn because the landing page was on Vercel and looked so suspect.
- Boris has said multiple times that Claude Code is increasing writing/maintaining Claude Code, which I deeply respect as a person who's a fan of compilers compiling themselves (like Go, which has since 1.4).
- Have a look at /r/apple on Reddit on a Sunday or even here on "Show HN" threads. Most of the stuff coming out is vibeslop, but some of what's being published looks like solutions to real problems.
This has a clear answer in systems theory; the ability to bang out code was never the limiting factor for Google et al.; large organizations are limited by coordination costs.
It's the small teams you need to pay attention to, the organizations that really were limited by engineering capacity. These are by definition also less visible - for now.
That "for now" has been going on for ~10 months. That was enough for some startups to make a splash in the pre-AI era; now that they can move 10x faster, shouldn't we be seeing evidence left and right? All the contenders stealing lunch from Big Tech by offering better products or exploring new frontiers?
>Mathematics is probably the same way. There's a short-term rush when you pull the lever, but there's less desire to get invested in what comes out.
Perhaps it will take until the next generation to come along to really embrace the new AI-assisted way of doing mathematics. The current generation has too many reservations.
Large companies move slowly and have lots of red tape. Give it a decade before making this call as far as large companies are concerned. That's how long it takes to change processes.
I think software from major companies is in worse shape than in 2021. But that's on purpose, the enshittification continues. Otherwise I agree with you.
Selfishly, I hate these things. No one in the 1980s and early 1990s expected the internet to become what it did. I discovered it as a dumb teenager and posted some frankly idiotic stuff. I also put my real address and phone number in the signature, because that's how you rolled back then, the entire internet was like 10 nerds.
Now, 35 years later, a lot of other stuff has mercifully decayed, but people keep bootstrapping these Usenet archives as pet projects every year, and there's never an option to opt out. Today, agents make it easier than ever to build a dossier on anyone you don't like, so this data is bound to cause some harm. And I guess when I die, the best-preserved memory of me that my grandchildren and grand-grandchildren will be able to pull up will be snarky posts from a 15-year-old on some unix-related discussion group.
Yeah, I know there's a bunch of posts there that are of genuine historical interest. But I'm just not a fan and I don't mind burning karma to speak my mind.
Same. It’s hard to understand today, but in a world where search engines weren’t really a thing and almost no one knew what the internet was, the concept of “public” was quite different.
Yes, I should have known better in my teenage years about it. And also yes there was no reason for me to expect that everything I wrote would be easily available to some AI panopticon 30-40 years later
I can imagine you’d be embarrassed, I did stupid stuff too, but unless you’re running for Congress you should just say “I was a stupid teenager” and let it wash. Anyone fair minded will dismiss it.
The current social zeitgeist is that one is permanently responsible for everything they ever said, even has a child. All nuance is gone. People are looking for a reason to make someone the enemy, and if you're the enemy, violence is acceptable now. I wish it weren't like this, but it is, so I can understand why people want to stay out of the crosshairs.
Man, I sure knew 35 years ago that whatever I shared in public was.... public. :)
And some of it still is public from way back then -- including dumb and ugly comments, stupid flamewars, teenager stuff, and phone numbers. It doesn't bother me.
I don't feel remorseful for having once been a kid and doing some of the dumb things that kids sometimes feel inclined to do, and those landline telephone numbers stopped being relevant to me decades ago.
Public or not 1991 was a very, very, very different time. The idea that some dumb stuff I posted would still be around in 30+ years would have been crazy.
(Actually, I wonder if the Prodigy forums got archived...)
We're still working on it! It is worth noting that the STAGE.DAT an CACHE.DAT from C:\PRODIGY will never contain things like email messages or forum posts. They will contain things like the pages and program logic that comprise particular parts of the service, and at most parts of the various service database indices and records (e.g., bits of the Encyclopedia, Movie Reviews, that sort of thing).
35 years ago we didn’t have the cultural precedent/zeitgeist of getting cancelled for some stupid shit you wrote on a message board in 1991 so we were less guarded with our opinions.
No, and we don't really have a culture of cancelling things for people they say now, either, when it comes down to it because if that were true, Trump would not be president given his long history of [waves hands], a large number of Republican politicians would also be out of work, etc.
So what is it that drives such phrases into common parlance as a description of the apparently-fictitious culture of rote cancellation, might one suppose?
I might suggest that "projection" is a primary motivator, although I work to avoid projecting on my own time.
(Edit: To disambiguate without vagueness: When some fucking party or other is talking about "cancel culture," they usually seem, to me, to be projecting about their own fucking party's fucking ugly policies and tactics.
And in a sane and pure world, I'd like to wish that they'd be recognized as being fucking scumbags for behaving in this way.
But it is my observation that we do not always live in a sane and pure world.)
My view is the complete opposite. In the 90s I thought the digital revolution and the internet will be so groundbreaking that everything connected to it will be remembered. Moreover storage seemed so cheap, why would we ever want to delete anything. The internet never forgets.
Turns out storage wasn't free and the the following generations did not have the same appreciation for everything 80s and 90s, so much of it is lost forever.
> when I die, the best-preserved memory of me that my grandchildren and grand-grandchildren will be able to pull up will be snarky posts from a 15-year-old
Just make a bunch of really good posts and put your name on them
I'm really tired of these arguments (this and "it's just like calculators").
Photography decimated other forms of visual art, so the concern wasn't wrong. But AI threatens the entirety of human intellectual endeavors. I can make do without oil paintings in my home. I'm not sure I want to live in a future where we make do without brains.
Lots of people still make bad music that other people still manage to enjoy (a lot of it has gone multi-platinum!) even though they're not Mozart or Bach.
Of course it is, because our society is pretty much structured to place sociopaths on top. Why is it at all surprising that any technological advancement is going to be used by them in a way that doesn't benefit the rest of us?
But it's equally obvious that this has nothing to do with the tech itself.
If the tech enables certain behaviors then it can’t be casually divorced from a discussion of how people use it. “Tech” isn’t inherently neutral. It’s made by people with intent for a purpose and how else it can be used must also be considered.
This is like separating “the use case” of killing people from guns and still trying to have a coherent discussion about them. You simply can’t handwave it away with “well that’s not the tool’s fault.” It’s unproductive and sidesteps the conversation.
It did. While you're busy recording an "instagramable" moment, you are not entirely enjoying that moment. Your eyes are on the screen, not on the subject.
In a way, if photography is an ersatz for painting that eventually made imaging available for the masses, then AI could become an ersatz for thinking. But it feels like I'm paraphrasing TFA.
What will happen when AI companies have spent their advertising budget on math problems and whatever else gives the maximum wow effect for the bucks? Probably customers hooked on the vain satisfaction of spending token$ to impress friends.
Yeah. My other favorite example are books. Why do nonfiction books exist? There are some pathologies and corner cases, but fundamentally: to develop and share new ideas. Downstream from that, if it reads well and if you're lucky, you make some money.
But now, LLMs can generate hundreds of books per hour. They make up 80-90% of new arrivals in many nonfiction categories on Amazon. They short-circuit the system, allowing their "authors" to extract money from the system with zero effort by crowding out human work. And it's not even the question of whether these books are good or bad (although overwhelmingly, they're terrible). It's whether it's actually accomplishing anything worthwhile, or just destroying incentives for humans to write or go into any other sort of intellectual work.
In fact, I see many professions push back. Artists, writers, now mathematicians. And I'm amazed that our profession doesn't and that we have so many people who are hooked on vibecoding. I'm still waiting for that 10x payoff. All this velocity and somehow, the landscape of the software I want to use still looks the same as it did in 2021.
It doesn't, that's the whole point. You can produce output that passes the smell test with naive buyers in a matter of hours. If you want to write good LLM books, then you gotta work closer to human speed - weeks, months - which allows others to produce 100 slop-books in the same timeframe. You still lose.
If an author let AI do 90 to 95% of the work, cleaned up the remaining 5 to 10% themselves, would you say AI or the author wrote it? Because this is roughly the split between an author and an editor, with human writing, now.
The work is either: (1) the act of concealment: generating an AI book that cannot be detected as AI generated; or (2) actually just writing the book yourself.
Similarly to many other countries, the US has a concept of "adhesion contracts" - essentially, contracts where you have no realistic opportunity to negotiate the terms, and your rights are constrained much more than the rights of the party imposing these terms.
Courts often look at these contracts differently, but around the world, they allow them to exist because they are useful. A good example are the "terms of service" for public or private transit. If the carrier can't define some common-sense rules, like that you can be kicked out or fined for not wearing pants and playing bagpipes on the bus, it'd complicate things.
The legal standard is basically that the rules hold unless they're unreasonable or unconscionable. But of course, what's seen as reasonable depends on the country, the state, and the judge.
> not wearing pants and playing bagpipes on the bus
The problem for some of us here is that we're left without a viable transit option, as the yes-pants-no-bagpipes model of transit essentially has a state-sanctioned monopoly.
I suspect if market forces were allowed to operate in this area, we'd see fewer pants and more pipes.
"Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal.
The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.
Yes, blame them for seeking out an upper middle class lifestyle with a relatively standard home in commuting distance of their place of work and dedicating the rest of their life to teaching mathematics to new generations of people. How vain a pursuit.
After all, the ascetics at openAI are having to make do with half a million total comp.
>Built a top a pyramid of failed math undergrads, grad students, and mediocre post docs.
Like much of things in this world, when you take a step back and realize that it was another human being who made that lunch time slop bowl for you, for the lowest wage the law allows for.
I don't know about you, but if I apply myself fully to a problem and study it to the point where I'm literally one of the world's experts on it and then some assholes in Silicon Valley take my research and claim it for themselves, I will probably not feel too great about that...
Many other techniques here are more or less quack science with some sordid history. If I recall correctly, the entire field was invented by a guy of questionable integrity. Most of these methods are riffs on "if you download a badly compressed JPEG, paste some higher-quality content into it, and then save it at high quality setting and the same resolution, it's going to show if you run a highpass filter on the result". This is not wrong, but commonly fails on innocent images and it's also not how professional photo editing is done.
Tools that aren't this seem even more questionable. The "clone detector" triggers on anything that naturally shows even a bit of symmetry or repetition. The "noise comparison" tool detects bokeh and smooth surfaces. Etc.
TL;DR: Experiment with it, but don't assume that just because it sounds scientific, it actually works in real life.