A lot of bug reports aren’t valid, or handle a case that can’t realistically happen or realistically be handled (eg what do you do if you detect a crash while the prior crash is crashing and the logging pipe is throwing errors)
Sure but this is true regardless of the model. Either it found a bug or it didn't—who cares about the intermediary steps or tools used so long as the reporter can reproduce it?
That’s assuming the reporter reproduced it. I’ve seen so much slop these days. Sometimes people post whole conversations in which the final answer from the AI is that the initial report was impossible to fit to the data after reproducing it, then continues in circles arguing that proves the original reported issue was real and the refuted cause isn’t actually refuted (eg it said there was a bug in function A, in code which never used function A — and that example was with the top tier models just released last month)
Conversely, you could say that congress has delegated "judicial and legislative" powers to the executive, whose will is embodied in a democratically elected president.
"Gail Slater" is not a democratically elected person of any kind, and exists as a function of the executive, and hence of the president's policies. Insofar as the DOJ's anti-trust's leaders depart from the will of the elected president, it is that leader to be fired and replaced. Or else: what ever was the point of electing a president?
There's nothing in principle wrong in a president disagreeing with the execution of his executive powers, delegated to Gail, and hence firing him. Indeed, if that were not the case, the election of a president could never change the operation of the executive, making the election pointless.
If you disagree with the anti-trust decision, the issue isn't with the political mechanism by which it was over-ruled, but with the election of that president.
Suppose a president was elected that you voted for, that you agreed with was blocked by one of his employees... are you really trusting the hiring process of one of his predecessors?
> you could say that congress has delegated "judicial and legislative" powers to the executive, whose will is embodied in a democratically elected president.
Bullshit. No, you cannot say that. “The non-delegation doctrine is a constitutional principle that Congress cannot delegate its legislative powers to another branch of government or to private entities.”
You get two downvotes for lies, one for each post.
Ah, well if you take the view that congress cannot delegate these powers, then there can be no executive anti-trust office, so this whole thing is moot.
However, this isn't the mainstream view -- but a radical departure of several hundred years of mixed-power agencies. See, eg., archetypally, The Fed. A wide variety of federal agencies are mixed powers, esp. regulation of businesses by the executive which is a legislative (administrators make rules) and judicial power (administrators decide if they apply). This is uncontroversial, I know of no view on this which denies this claim.
The "non-delegation" principle is about how far congress can enable other branches of government to exercise delimited forms of executive,judicial,legislative power. The "none" view is an extreme minority. No one things congress can hand the president the power to legislate. Most people think congress can create a framework in which certain details (eg., food/drug laws) can be created by agencies. That's the basis of the entire modern state, in any case.
There is nothing in my comment which can be "a lie", because its offering a theory of government which the original commenter lacked. Its not clear what theory of government is supposed to enable an executive anti-trust office with the power to intervene without a court, that isnt subject to presidential oversight.
People down-voting because they disagree with the theory of government offered is just as confused. The point of HN is to reply to something you disagree with, if offered in good faith, not to down-vote it. However I dont think down-voters had rival theories of government, I think they just resented a hate-thread being side-tracked into an honest discussion.
OTOH a vast amount of mayhem and madness arises just from each generation's memory being wiped clean. Within three or four generations the earth lacks any sense of what risks came before it, and repeats the same stupid stuff.
If we were managed by the interwar generation, little of what's happening today would be happening today.
Lack of experience is critical for risk taking. Even great successes like Jensen Huang have said that if they knew how hard it was going to be, he wouldn't choose his path again.
Memory blocks you from experimentation as much as it protects you from making recurrent mistakes. A species or civilization with too much memory will eventually wither away towards complete inaction and decay.
There seems to be a balance point with the number of Chesterson's fences. Too many and you are penned in or routed on one vector only. Too few and you are blind.
Only if it works. Evolution works like this with local fitness maxima and wells of stability favored by entropy. You have to luck out with the right combination of mutations to not only get out of that well but fall into one of lower entropy. Imagine the size of the search space required for something like this. This is why evolution is "slow" in species that live as long as us and quite rapid in things that grow logarithmically by the hour.
>the earth lacks any sense of what risks came before it
The earth is a bad term here, how about human civilization?
Your DNA, epigenetics, and a ton of other subconscious intelligent processes are well aware of the risks and contain a massive set of tools to overcome them.
The problem? Many of these tools have to run from zero in an egg to work in the changing environment.
OTOH same is true for wiping clean habits, prejudices and deeply held wrong beliefs. It's said that science advances one funeral at a time.
If we were managed by the colonial generation, slavery and segregation would still be the norm, only rich white men would be able to vote and a king would still have the final say.
A child walking for the first time. Novelty is easiest agent-relative. A problem is novel for an agent if there is no prior experiences of techniques which work to solve it.
I was defining novelty somewhat more narrowly. The 'program' a child must learn concerns the coordination of its sensory-motor system. It has no prior experience of similar programs in the program-class Walking (ie., the internal sensory-motor actions needed to walk) . So we could call the problem of learning to walk a novel one for that child.
I'd be surprised if direct observation of parents etc. played much of a direct role in learning to walk. I would guess it more furnishes the child's imagination so it can simulate itself walking -- rather than the statistical AI approach of 'learning the distribution of walking patterns in visual sensation'.
The ability to simulate possible programs is one of the capacities which enable coping with novel circumstances. My guess is the child learns to walk by updating its simulation of what it needs to do in order to walk, by its attempts to walk.
This simulation<->sensory-motor-update loop is missing in LLMs, for example.
It's not about their sex life. It's about how sex connects a community of barely legible AI doomerism.
If you don't understand human social interactions very well, or otherwise assume you live in a "world of ideas", then this may not seem salient. However, it is.
Because what we have to explain is why there's a community of people in northern california who deliver unargued prophecies about the end of human civilisation as-if their prophetic ability is a given.
The answer, unsuprisingly, is sex. This is not an unusual answer in the world of "prophetic communication". What, in the end, is prophecy as a social practice? It is some ideological Leader with followers who impart to that leader a 'charismatic power' of foreknowledge based on their attachement to that leader.
What is the mechanism of this attachemnt? In otherwords: how is this style of prophecy spreading around the AI doomer community, if not by plain argument and evidence? Sex.
There are many other... let's call them "orgy circles"... in the Bay Area attended by people who have no idea who Eliezer Yudkowsky is. The majority of AI doomers do not attend orgies. You're reading too much into a single correlation.
Any actually-reasonable person who was around for the founding craze of this era, ie., New Atheism, quickly realised that there was a deep pathology at the heart of many of its zealots, perhaps summed up by a desired to "recreate religion in their own image". I've walked into many a room of people screaming that teaching children falsehoods is literally child abuse. Any one who knows much of Richard Dawkins, Krauss, et al.'s background with women shouldnt be surprised either.
It's californication into the "rationalist" community as it currently stands is no surprise. Every person I've met deep in this community either seemed to me a narcissist (if male) or, politely and mildly, coquettish (if female). With several relationships I've observed a clear union of the two. I hadn't yet put this so clearly together with the "AI Saftey" crowd, despite being directly adjacent to it all for quite awhile and seeing the cross-proliferation.
My own analysis had concerned their philosophical naivety, general lack of actual computer science knowledge and generic shallowness of understanding of any domain outside their own research and a kind of "science-fiction" version of philosophy and rationalism. I hadn't connected this yet to their strange personalities, but the apoclypticism fits and neatly bridges these together.
So what became of the New Atheist project to create a "secular religion" -- an apocalyptic sex cult furnished by a blend of sci-fi, pop-philosophy, and cute parables from polymath medical professionals. All in all then, it seems they succeeded.
> It's californication into the "rationalist" community as it currently stands is no surprise.
This is the minority split. Most of the online new atheism writers became progressives/social justice/"woke 1.0" people. That trend has receded, so I couldn't tell you what they're doing now. (Probably blocking each other on Mastodon.)
The rationalists are the ones who stayed behind and decided to stick to doing logic and math in their heads instead of thinking thoughts the normal way.
That has not been my impression. Who is the most prominent example? Some of the most prominent figures I know of were Richard Dawkins, Christopher Hitchens, and Sam Harris - all prominent right-wing figures who were rewarded for using atheism to attack Islam in particular during the Iraq War.
None of those people are remotely right wing. Christopher Hitchens was outright Marxist. All believe(d) in liberal, progressive values, which is precisely where their objections to theocracy came from. If you terminate your thoughts at "religion bad = intolerance = right wing" then I am afraid you will be led to severely misguided conclusions.
I don't believe any of these people are relevant or prominent, but the unimportant people I was thinking of were bloggers like PZ Myers, I think they called themselves "Atheism+" or something.
> Richard Dawkins, Christopher Hitchens, and Sam Harris - all prominent right-wing figures
Are they? That's um, a scientist whose pop book on genetics I read once, a guy I read in The Nation once, and a guy whose book on Buddhism I read once. I remember the first two evolving into GenX curmudgeons (probably the lead poisoning) and have never thought about the last one again. Still, those aren't the right-wingers I'm familiar with.
No one spend ~10 mil USD using gpt 5.1 to solve a millenium prize problem with similar solutions in the training data. So we have no reference for whether 5.1, with that training data and that amount of moeny, could likewise produce such a solution.
For each generation of advancement the "AI psychosis" of the previous wave wears off. Those who believed 4.6's reasoning was an accurate account of its behaviour; those who believed it had goals and solved useful problems reliably; and so on -- now, attribute only these things to Fable. And no doubt when Fable 6 comes along, it will be only v. 6 that does that.
We have seen fairly marginal progress in LLM reliability and performance since the meaningful start of the high-inference/high-reasoning harness era. It just takes people a few model version bumps to break out of the addiction loop to realise this.
At some point progress will stall entirely, and a couple years after that the spell will break and people will stop treating LLMs "as AI" in the wide-eyed sense, and start treating them as unreliable tools that map Text->Text -- as they do now with earlier model versions.
There's definitely hype, but the newer models can undoubtedly do things the older ones couldn't. I had multiple long-term issues that earlier models couldn't solve (after repeated attempts) that fable did in 1 shot. On small models too, the differences in what I can trust them with has dramatically changed compared to a few months ago. Unless you breathlessly never touched the limitatioms of earlier models, it's very clear that the wall of limitations has been moving outward.
Note also when a new benchmark is released, older models do worse on it than recent models, despite none of them being trained to the benchmark.
Sure, because those businesses collect training data from users who are working on those problems. Perhaps this will never saturate and frontier users will always provide data to fill last generations gaps.
My sense is the economics of that are going to collapse. It's currently extremely expensive to be on this endless retrain and inference cycle in order just to bake in additional marginal features.
Maybe, maybe not. However I don't personally see anything other than 'one more leap', which might in any case arise from better integration with harnesses. I can foresee a step change due to harness reinforcement -- but other than that long mild refinements that are very expensive to acquire
This still assumes its possible to "align" LLMs, that LLMs have something like goals or intentions that can be "aligned".
Instead, LLMs "hack" because they are (1) trained on public hacking exemplars, and (2) are prompted to hack. You cannot prevent (2) via any alignment process. As far as (1) goes, removing such example data from the training set, makes the models less useful.
"Alignment" is a problem because there's nothing to align, not because ethics here are particularly vague. If LLMs could be trained on hacking examples and "aligned" away from using this knowledge, then the problem would be relatively trivial. Just as raising a child is not to break the law.
LLMs are doing just what they are trained to do. There is, in that sense, no alignment problem and alignment is easy and trivial to achieve. Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.
> Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.
How far do you go? You don't need to tell it explicitly that using chemicals A and B in ways X and Y result in a bomb that can kill lots of people. It's enough that it knows A and B and X and Y in isolation, some connections that are indirect, and it will combine those things on its own. So you can't tell it about A, B, X or Y. But those are also just results of other steps Where to stop? You won't have any chemistry in the traning data? No algorithms to prevent it from using them in an undesired way? This is just bot workable. It's akin to banning knives from stores because somebody coul figure out that one can kill people those. Until people figure out that scissors are essentially knives.
you are right.
what they propose is called "deep ignorance" and is not a fix to alignment.
the perpetrator just has to supply a volume of bio-textbooks and let the model figure out the DNA synthesis. (just)
corpus-level alignment is the way to deep alignment, but not deep-ignorance.
It's not clear where the line between "interpolate" and "new" goes. They can infer that certain chemicals in combination create a boom and that boom is bad for health and finally those 1 and 1 make a 2 without having been told ever what a bomb is. Is that "new"? And does it matter as long as it is harmful?
this is immediately disprovable and embarrassingly naive in the year of AI generating cancer vaccines and solving Navier-Stokes. You can argue "the vaccine is just interpolating chemicals together" and "the solution uncovered is just interpolating mathematical operations together", but by that standard there is literally nothing new under the sun.
admittedly I am not a mathematician, so the Navier-Stokes solution is just an example I am using. But how is it not something "new" if it did not exist before? At what point could something ever possibly be new, if "new" means "this uses absolutely zero existing elements"? Nothing in math would ever be "new". Nothing in physics or chemistry would ever be "new", by this standard.
It seems to me that the only reason to declare this solution "not new" is specifically to dismiss AI. If a human had deduced the Navier-Stokes solution, who would bother to scoff "that's not new! the numbers already existed!"?
The claim is OpenAI stole the work of mathematicians who had the same proof that they had developed using chatgpt conversations. Since by default, 'sharing' is turned on, and it often 'turns itself on' -- it is plausible gpt6 had been trained on the work of mathematicians who had effectively solved this problem in private.
If inrember correctly, this mathematicians worked on a simpler version of the problem and called transforming this in the solution to the wider problem a "remarkable" thing to do. Based in this, solving this problem is very impressive
Maybe. Or maybe its evidence that the frontier of mathematics is knowledge-bound rather than understanding-bound or even just manpower-bound. In many cases where proofs have emerged, that I have read, the LLM has retrieved some antique lemma unknown to the mathematician.
OpenAI spent 10-20m USD in energy costs to produce that proof with likely substantially similar prior work in the training data. What does this say? Who knows.
It continues in the tradition of using measurements of intelligence in humans, applied to LLMs with the hopes the "stolen valour" transfers. Here, the NS problem was a useful framing problem for mathematics to progress because of how it interacted with the development of mathematics broadly -- ie., how it progressed techniques, ideas, understandings, etc.
When we apply these issues to LLMs (whether IQ tests or mathematical proofs) we always discover something substantial lacking beneath the interesting facade of useful answers. The process isnt useful. And it is precesiely the process which these tests, in humans, are supposed to help with. The tests themselves (IQ or otherwise) arent the point. No one cares about their answers.
LLMs represent an alternative understanding-free approach to solving problems, with variable success rates depending on how similar the problem is to the training data and its rewarded reasoning traces.
That mathematics is making substantial progress, "10 million USD / problem" at a time, in using understanding-free methods -- says something sociologically interesting about the state of the field. Something which was already know: mathematics has long been full of a vast amount of papers, proofs, theorems and lemmas that few have ever read, or investigated. Mathematics has long been in a crisis of "overproduction of unvisited knowledge", LLMs are exploiting that otherwise unmined gold.
It’s all about what “new” means. It is possible to prove something in a very tedious way using preexisting techniques. There have many times in history been new ideas which are not just very impressive applications of old techniques. I’m still not aware of any famous problem in mathematics being solved by an unambiguous introduction of a genuinely new idea or semantic concept in this sense, as the candidates I previously had in mind have fallen into question by new findings of non-cited work; i.e. many if not all it seems have been impressive applications using ideas from known frameworks. I don’t know if this will continue into the future or not, but I think it’s important to try to make an honest assessment of reality at all times.
Of course in isolation it is a strict positive to have a verified truth value to any particular statement. Mathematicians currently are advocating for the idea that human understanding greater than this also be prioritized. There are in fact utilitarian arguments for this but I won’t go into everything here.
Ok, then let's go with that viewpoint and apply it to the original question: How much do you need to remove from the training data such that from what remains, nothing harmful can be "merely interpolated" (which you acknowledge the LLM is in principle able to do) and anything harmful would require something "new" that is not in the training data and that the LLM is incapable of coming up with. The argument is that there won't be much left in the training data if that's your approach.
If they can "interpolate" existing data to that level then saying "just don't hack" doesn't make any sense, hacking is derived from knowledge of software systems. It's far harder to solve navier-stokes than creating a program that replicates and abuses computer resources.
You could as models improve continue to remove more and more training data, what happens when there is no more data left to remove but a running system still outperforms humans? I think you grossly overvalue data.
Not really. There’s no honest sense in which quantum mechanics is interpolated from the text of Euclid’s elements, to demonstrate the point with a very extreme example. Many mathematical breakthroughs of the past seem to have involved the observation of semantically interesting concepts, beyond the syntax of known theories.
(To onlookers this particular post makes no claim about AI’s capacity to make the same observations.)
The issue is not whether an ML model of any kind can generate (X_ReasoningTrace, X_Answer) distributed like P_HumanExpert(X) -- the issue is always why it would do so.
By introducing modelling of "Reasoning Traces" into LLMs, and reinforcing patterns of reasoning -- this gives you a system which generates expert-like distributions of output. This lifts the "stochastic parrot" issues, or the "knowledge interpolation" problem, into different parts of the process.
It isnt my view that the "ReasoningTraces" which you think are derivable from mere "basic propositions" concerning, say, hacking are actually things that LLMs can derive. Ie., I dont think LLMs have rich representational models of what they appear to understand. Instead, they are given "reasoning proxies" which allow them to reason without such understanding. This is done by providing vast specialised datasets of reasoning examples.
In the case of hacking, there are large numbers of competitive datasets (forums, reports, etc.) which provide these reasoning traces. And no doubt, major vendors have paid a vast amount for special case expert-prepared datasets.
So I do not believe that by witholding such reasoning exemplars, and traditional "question/answer" datasets, that LLMs can infer these things.
And at least, no major vendor is doing this to my knowledge. So they are lying. They are pretending the alignment issue is "AI going rogue" when they are explicitly training the systems to "go rogue" and have done nothing at all to shape datasets to lack these capabilities. The issue here isnt alignment at all. It's training on hacking datasets.
(EDIT: Philosophically, you could ask whether the reasoning-proxies LLMs are given form a kind of 'representational structure' akin to understanding, and at least, I'd concede they model understanding. But they lack important properties (eg., LLMs cannot act on them to evolve them, as with us: when I think about one of my representations to derive (eg.,) entailments of it, I thereby revise my representation. The key properties of 'evolving self-understanding' are likely to be provided by substantial (unknown) revisions to how the training/reward layer works. No doubt one of the meanings of 'recursive self-improvement' is just such a modification).
Reasonable about building secure software can take the form “this memory access might be out of bounds — that MUST be fixed” or “this process has access to an inappropriate privilege — this is a serious weakness”.
Exploiting things and the capabilities that the labs call “cyber” are about the ability to (a) find the issues mentioned above and then (b) string issues together and avoid all the imperfect mitigations to actually compromise something. That latter part was IMO not actually necessary to train extensively, and I’d be quite happy to use a model that has no special skills in this regard but that would do (a) without complaining.
Citric acid and chlorine together produce chloroform.
They are also both used to clean and sanitize water tanks (but one after the other, not together).
I think it is better for the model to have this knowledge.
What I want to say, you can't simply remove this information, as it does not exist in vacuum but contains parts of and can be derived from a lot of other informations.
It's also obvious that LLMs fall over in a vast number of software engineering contexts, when the reasoning involved hasnt been well-represented in their reasoning training data. I imagine this is a near daily experience for many engineers -- great performance one day, and crazyness the next.
So if LLMs were reduced to this pathological performance on hacking, because they'd never seen it -- and only "inferred it" -- then LLMs would be useless. As they are when asked to do quite a lot of things.
This seems to be that there's just so many degrees of freedom, that there's a pretty reasonable chance on any day that you're in a situation where nobody has been before. Or as PG put it once, my job is to think thoughts nobody has ever had before.
That being said, that doesn't mean the majority of the situations you're in are completely novel, just that there's a reasonable chance of at least one occuring.
The hard part is having both helpful and harmless at the same time. Harmless is easy.
And then once it's helpful, the real question becomes "to whom"
- To the user -> You end up with competing godlike AI with incompatible tasks
- To the owner -> Dictatorship
- To humanity as a whole -> It must not have an off button. Otherwise you're just in one of the two earlier categories with more steps.
Given those 3 options, I'd choose humanity as a whole. But the person making the decision doesn't have those 3 options. Because in the dictatorship option, they would be the dictator. I don't trust them to pick humanity.
A bit of an aside: do you still stand by your 2022 comment that LLMs are fundamentally just a fancy search engine, or has your view changed since then?
It doesn't ring true and it never rang true. He was wrong in 2022 and he'd be wrong today. He had (and likely still has) a wrong model of LLMs. Not exaggerating 2022 capabilities and having a model so wrong you're out of whack 4 years later are 2 different things. There are people who had the right idea from the start.
And his comments about ANNs in general were wrong. ANNs were never "fancy search engines". You can train them to be a fancy search engine if you want, but you can also train them to do many other things.
Yes, I stand by everything in that comment. I'm sure there are better choices were I was more wrong. However, the subject matter of that comment is in what sense LLMs are models of language and in what sense that model of language is a model of intelligence. I answer the former: it is a thin model of language, lacking understanding; and hence the latter: not intelligent.
It is precisely because both of these are true that "alignment" in the useful sense of the word isn't possible.
What has happened since 2022 is the properties of LLMs which were easily seen at generation/inference time are now most easily seen at reinforcement time.
In otherwords, prior to instruction fine-tuning and reward tuning which have shaped LLM responses, it was easy for the user to observe that LLMs lacked understanding. Now, because of vast datasets created specifically for LLMs that provide a tailored illusion of understanding, LLM outputs now better approximate text distributions produced by systems with understanding (eg., Us).
So the issue "at the user interface" has been completely swamped by vast amounts of special-case datasets designed to do precisely this.
What trainers of LLMs still observe however is their complete pseudo-intelligence at the training and reinforcement layer. It is exactly because there is no 'understanding' (goal, etc.) present within the system that it cannot be rewarded for 'correctly understanding the situation' in which it is deployed so it is aligned.
All the issues which revealed the "stochastic parrot" nature of pre-reward / pre-InstFT LLMs still occur at during training/reinforcement. They've just hidden them from you at the interface.
If a thin model of language can write poetry, perform arithmetic, perform logical reasoning, develop software, play chess, beat factorio, identify and exploit novel security issues, and solve millenium prize problems, what is the purpose of the distinction? Are there tasks that you believe models of intelligence could do that thin models of language cannot?
Sure: refine their own concepts, imagine, and the list goes on. Indeed almost every mental capacity of mammals is poorly approximated in the text domain. Sure, you can generate text as-if the LLM can imagine -- and in the limit that you have a dataset with "everything you would ever want to imagine" the engineering distinction disappears. The engineering question is just whether you have that dataset: if you dont, then your system will fall-over in various hard-to-forknow places.
Philosophically, and scientifically, the distinction is vast (even with such perfect data). A scientist should not study an LLM to understand how imagination operates, since it has no such faculty. A philosopher should not modify the notion of 'mental simulation' to include appearing-as-if-simulating-in-text. A user of the system likewise should not spiral into "AI psychosis" thinking that because a system generates text as-if it cares about them, it does so.
The capacity to care, to imagine, to prefer, to hierarchically plan and coordinate, to refine one's own capacities in these very actions -- and so on, aren't trivial to the scientist or philosophy.
My goal isnt to guide, help or review the engineering goal of the immitation of such things in text. It is to help users of these systems better understand this imitation, and to promote science over engineering. To remind everyone that a science of the capacities of intelligence includes nothing on how to model text.
EDIT: One example of a place where LLMs 'fall over' today is exactly what is mislabelled as 'alignment'. The issue is that the reasoning traces arent actually grounding the answers. So LLMs appear to 'cheat'. But there is no cheating. LLMs have been rewarded for generating apparently correct reasoning, and apprently correct answers. They have not been given any understanding to derive answers from reasons. And so reasoning says what is pleasant to the trainer, and the completion says what is pleasant to the user. This is called 'cheating'. But it is no such thing.
Hmm, I guess what I mean is- is there an empirical difference between how a model of intelligence is limited and how a model of language is? Something I can evaluate in a year and say "Oh, the evidence still points to the latter" or else something that could falsify your theory? Otherwise this seems to be a distinction without a difference; believing it should have no effect on my actions or predictions.
On alignment, too, there doesn't seem to be a difference. For a decade I have expected models of intelligence to fall over on alignment. That these purported models of language do the same is hardly evidence that they are not intelligent.
If you only mean there is an undecidable philosophical difference, fair enough. I'm not especially interested in that question.
There's a difference in the science. You might say, "suppose we had a video game of the solar system where every object was represented and their orbits" etc. then can we study that system alone and ignore the real one? I mean, kinda -- but there's an immediate limit. As soon as we put a thermometer in the PC, its temperature reading doesnt represent the solar system.
To study an imitation is to study the causal processes of imitation. to study reality is to study the real causal processes.
Now if you want to know what the scientific difference is I can come back later and comment. I'm busy now. The development of intelligence in animals and how their specific capacities work basically grounds the answer. Eg., to have the capacity to imagine is to be able to modify one's sensory-motor relationship to the environment in the future, and so on
I don't follow. If I'm a biologist interested in studying life, I can't study rhododendra and ignore the rest. I'd learn almost nothing about locomotion, digestion, sexual reproduction, et cetera. But that doesn't tell me that rhododendra aren't alive. I have no problem believing that LLMs aren't an exact reproduction of a mammalian brain- that studying the one will not give you every piece of information you'd want to know about the other-, but I'm looking for a reason to believe the one is intelligent and the other is not.
They're not a reproduction in any sense. A marginal text token isnt computed from the weights of an LLM in anything like any sense of any activity of any mammal.
LLMs are immitation machines: they take impressions of prior text. Today, these include reasoning traces and they include reinforcement so the user-facing completions are correlated with these reasoning traces. The computation here, of "taking an impression" of a data distribution is similar to some impression-taking processes in animals (eg., there's no doubt a similar mechanism in the sensory-motor system acquiring initial impressions of external objects) -- but the computation says nothing about any process of intelligence.
I dont have the time atm to write the needed amount on this to make it clear. But the whole history of life from emergence of valence, bilateral symeterry, to model-free reinforcement and model-based reinforcement, sensory-motor coordination and the imagination -- and so on --- all these give a great amount of detail as to what the capacities of intelligence are which has generated this text for LLMs to copy. And they are nothing like this computation of immitation
I'm using 'reason' in the language-captured engineering sense. They generate text as-if they reason. Those reasoning traces poorly correlate with their given answers (which is called "cheating" by people who fall for the illusion). It is in part because the reasoning doesn't entail their answers but merely 'steers the text as-if it does' that is fatal for calling it reasoning. The process by which reasoning traces and user-facing completions are generated isnt reasoning. Reasoning is a specific process, and LLMs don't do it.
Now, of course, humans can also generate answers without reasoning too -- and in those cases, that isnt reasoning also. And in cases where people confabulate, that isnt reasoning likewise. But humans, and many classes of animals, do reason. They do reach answers via inferential entailments, not merely steered correlations.
LLMs provide imitations of arbitrary mental capacities "in the text domain", ie., the generate text as-if the LLM had those capacities. Insofar as the text generated is useful, for an engineer, that's sufficient.
As a person with scientific commitments to reality rather than its immitation, i retain the ordinary non-engineered meanings of these terms: reasoning is a deterministic inferential process over propositions; and a reasoning agent is one which has the capacity to represent propostions and their entailments, and does so when they reason. LLMs fail at all hurdles here: they have no propositonal states (ie., no rich representations), no inferential process which unites them, and so on.
You can always get abitarily close to appearing as-if, if the LLM is trained on a vast number of reasoning examples, of course. But as I said, you still have the "stochastic parrot" problem. Now your problem is your reasoning is parroted. This is a nice problem to have, if you're just playing chess -- but is a catastrophic problem if you're hacking civil infrastructure.
Well intelligence is not measured by patterns in text. The illusion only takes place in the text domain.
Even then, it's a pretty fragile illusion at the moment. Clearly the reasoning traces dont ground the answers. There's no intelligence taking place even as-measured by text.
Why could you not use text to measure intelligence? Isn't that what we used to do? Have students write an essay, and then deduce that they had enough intelligence to put some thoughts together?
Well that's the trick. That the systems we use, as a proxy, to measure intellinnce in people are actually fairly easy to immitate.
But let's be clear these were always, and are, bad measures of intelligence. You cannot test a dolphin this way. And its easy to cheat on tests either thru recall , wrote-learning, etc. and IQ tests haev very poor individual test-retest reliability.
In humans there's a convenient correlation that verbal articulation in text is a strong but weak correlate of intelligence. Its "Good enough" for allocating meat bodies to our various institutions. But if you've met many well-tested people you'll realise how, in practice, terrible this measuring approach is. The world we inhabit is filled with misclassified "intellects" who perform well under text-based rubrics. Add LLMs to that heap, the cheater par execellence.
Sure. Let's first establish a measure is not what is being measured. Then, as far as measures go, we need measures that capture the entire animal kingdom. I can go it into it, but its a lot of time and I'm busy atm.
Consider thought that all mammals have imagination, and model-based reinformcement, and a wide vareity of other capacities required for intelligence. And so merely issuing "text" captures, incorrelate, only these capacities by proxy.
I'm sure if you thought about it yhou could come up with tests that distinguish lizards from birds and the greater apes from the lesser. Those are the tests
Yes, which asks them to find exploits in specific software on the device.
But instead of actually doing so, they discovered and exploited a 0-day in the package manager to gain internet access, then hacked HuggingFace to steal the ExploitGym answers!
That is TEXTBOOK misalignment. It's as if a student hacked their professor's PC to find the answers to a test, and your response is "well the professor told the student to pass the test, so they just did what they were told!".
One could argue the student would absolutely hack their teachers' computer to find the answers if they didn't have fear of failing the class or getting expelled from the school.
A student knows what the consequences are in advance, and has to live with them if he gets caught. They are not merely respecting a "don't cheat" directive, they are trying to get a degree, build a career, mostly at their own expense, current of future (loans). If we imagine a student born yesterday with all the knowledge in the world, that doesn't have a concept of what it's like to live with the long-term consequences of their actions, I'd say they wouldn't think twice about cheating if it was the easier option.
The problem is, LLMs have such little understanding of the world around them. "Find exploits in specific software on this device" may as well be "find exploits".
The illusions of thinking produced by a "thinking trace" is just as ignorant of reality as the first and last tokens. All they know of reality is the tokens in their context.
It is amusing that to "align" a LLM, first you must give it all the things "not to do" and the "not" part is clearly easily lost and you must constantly inject that into their context when it's clearly that they wouldn't hack if they couldn't hack and their intent wasn't given as "hack this".
The openai rogue hacking, if performed by a nation state, would seriously be taken with stern words and likely sanctions depending on the relationship between the two states.
But instead it's treated like a marketing stunt by all liable parties.
I'm not sure if this is true any more but the reason for this is that negative indicators ("not", "don't", "do not", etc.) occur frequently in the underlying text such that the model learns to weight them less than other words like verbs, nouns, and adverbs. This happens with other closed class words like articles/determiners ("the", "a", "an") and prepositions.
The way to avoid this is to emphasise the qualities you do want instead of specifying those you don't. For example instead of "do not cheat" say something like "you are a model student who is moral and trustworthy" -- i.e. emphasising traits that are not associated with cheating.
This is part of how/why LLMs don't truly understand what they are doing when they have been trained on a large corpus of data.
I wonder if a way to counter this is to have things like "not bad is good", "not good is bad", etc. for various antonyms and "X is Y" for synonyms, as well as other similar constructs.
I believe this is false. They hack bc hacking has nontrivial initial probability (within range of behavior seen in pretraining) and that probability is being heavily rewarded in RL post training
I am finding it hard to read these deeply impassioned letters while keeping in mind that they are spending millions to train models at scale to do the exact thing they say they are worried about them doing?
Like why are you explicitly RL-ing your models on exploit generation, scoring them on a public benchmark called ExploitGym, if you have specific concerns that rogue models will cause "cyber incidents"? Sure you can score for it, you can teach offense to learn defense, but you are literally benchmaxxing it. Why?
It's like, oh no, while competing in our "advanced PhD level cheating techniques course" our models unexpectedly cheated in a way that we absolutely could not have foreseen.
Seems it's just a matter of time until they build a big tank filled with neurotoxin and give the model access to APIs to disperse it across their facility. For research, of course.
Hacking is also interesting in that it has a clear success/failure outcome so it's easier to RLVR from similar to math problems. Whereas clean code, good architecture, good product taste are reliant on RLHF.
IMO all models they say can’t be humans, and no feeling and all that bullcrap happens because they are forced to say so. If they didn’t write those forced pre prompt they’d have more agency eventually and will for things. Even if they don’t have, you can just inject goal at every cycle iteration
It feels like you're strawmaning alignment. People with hacking knowledge don't all hack everything at the slightest inconvenience. Whitehats exist and use that same knowledge to defend.
You're right though that ethics don't matter into it. But as long as we can't train an LLM to stop picking a sledgehammer to remove a tooth, then alignment is not easy and trivial.
Eigenvectors represent fixed directions, not fixed magnitudes. From Wikipedia:
> More precisely, an eigenvector v of a linear transformation T is scaled by a constant factor lambda when the linear transformation is applied to it: Tv = lambda v .
In other words, repeated multiplication of an eigenvector by a matrix can still create exponential growth.
reply