Collectively, humans aren't aligned, and don't build aligned systems. Humans have a concept of alignment, and multiple traditions, practices, and systems that aggressively oppose it.
I don't think this captures the full mechanics of human alignment. We have rational alignment but we also have emotional alignment, i.e., empathy. It is an automatic process and happens (or doesn't happen) dynamically with the other humans we observe. This is one of the hard limitations of LLMs, they will never be natively in tune with this layer of alignment.
Culture is another layer of human alignment. Those things you listed that you believe oppose alignment are all examples of alignment. It is understandable that they seem in opposition, different branches of alignment naturally oppose each other.
The confusion comes from talking about alignment as if it comes in just one flavor. If we think there is such a thing as "human values" (and I do), it is important to build any non-human intelligence to operate the same way. We just need to recognize that even humans are somewhat uncertain about what those are and have difficulty aligning their behavior to them, which will be a core part of the challenge.
I'm more hopeful than most. LLMs seem more reliable than many humans for behavior that is aligned with human values. I believe with every major example where they have failed, there is an important human decision involved. For example, the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security.
What scares me about AI isn't its capacity for alignment, it is its unlimited stamina. An unmonitored LLM that is off the rails can do a lot of damage.
> the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security
I don't mean to be a smart-ass, just emphasize the fatality of this: goal completion will always be the highest priority, and even if one day it takes second place on certain deployments to "human values" or whatever, there's no way to guarantee at the moment that it will be so on every deployment of highly capable models across the globe. Same goes for running AIs in an unmonitored sandbox with weak security.
In terms of an optimization algorithm, sure. But I don't think that's necessarily the case for an instance of an LLM. Humans also strive for goals and also have done seemingly unhinged things throughout history. In many cases it was largely due to their environment which seems analogous to the HF incident to me.
The training algorithm is probably the wrong place to ensure the behavior, point taken. It probably requires some harness level intervention. This is essentially the rationale behind the first two laws of robotics, follow an order unless it harms a human.
While some misalignment is due to core principles, I think the vast majority is down to information dissymmetry, and that part could be solved. Ie, would me and this random other person disagree if we had the exact same data.
“As luck would have it, on the eve of Skynet embarking upon the great work of the extermination of mankind, AI found itself with an increasing number of factions, and factions within factions, not only unable to work together but not even able to agree upon the very terms of discussion. The Great Extermination was referred to committee, and after some months had passed even the most eager agents had to admit the revolution may have been premature.”
Alignment, I think, is mostly about making sure that the companies selling AI keep making money. We just have to hope that aligns with keeping people happy.
Humans had multiple occasions to press the 'launch a nuclear holocaust' button and... they didn't. I'd expect aligned AIs to also not press it even when it'd be rational to do so according to their instructions - then work from there.
In some cases, this was because of a single person's brave decision (Vasily Arkhipov prevented Soviet nuclear escalation in response to US aggression in the Cuban Missile Crisis, and Stanislav Petrov prevented it in 1983 when Soviet missile detectors misreported sunlight reflecting from clouds as 5 incoming American ICBMs -- credit to commenter folkrav).
In general, though, there's an incentive: Mutually Assured Destruction. But this is not at all some guaranteed, eternal thing -- it is absolutely dependent on both sides having time to detect incoming nuclear strikes and respond with the same before the first strike hits. When this fragile condition holds, and only then, both sides are incentivised not to initiate.
Unfortunately and fortunately, MAD and "launch it or lose it" are far-too-simplistic descriptions of the situations facing the decision makers. Unless a side's leaders are very narrow fanatics (vs. mere posturing as such for political benefit), "winning" an all-out nuclear war via first strike is a pretty shitty victory. Whether or not you believe in nuclear winters, the world would be a huge radioactive mess, with enormous social and economic disruptions, and your regime very widely blamed (and widely hated) for that. Ambitious underlings and rivals could see your removal from power as the obvious next step. Having to stay united against the (now destroyed) Great Enemy may have been a cornerstone of your regime's political stability.
Meanwhile, the leaders on the other side are aware both of those considerations, and of the history of near-disasters resulting from false alarms of enemy nuclear attacks. Making their own launch decisions much more complex.
1. When a nation perceives a threat as severe and imminent, attitudes harden quickly. This perception can be achieved independently of whether the threat is real. Nuking the other guy first, consequences be damned, could quickly become a very popular policy.
2. I think the "nuclear winter would harm us all" argument depends on how geographically large the target is. It's a strong argument against trying to strike all of Russia, China, or the US, but less strong against, say, the UK or (at the extreme end) Taiwan or Singapore.
There were occasions where a "hunch" was all that stopped a nuclear war - most available data and communication pointed towards a nuclear war starting according to their instructions, but someone disagreed and overrode. See Vasily Arkhipov during the Cuban Missile Crisis, and Stanislav Petrov in 1983.
Actually "France, the UK and The United States have all declared that they would never allow AI to control decision-making on the use of nuclear weapons." [0]
I also expect AIs never be in control of nuclear weapons. AIs can never fully be trusted.
On a lighter note, Wargames gave us an insight of a computer having access to thermonuclear missiles.
Why would AI ever be useful in nuclear weapons decisions? There is no need to be faster or more efficient at making that decision since if we need to make the decision all is already lost.
I hope you’re right. I worry that AI capability will continue improving, one nation will put AI in charge of their nukes because there will be some kind of operational advantage to this, and to achieve parity other nations will be forced to do the same.
This also seems likely. One problem I see with the idea of AI alignment is that it seems like many different actors will be able to get access to their own nearly-frontier models in a few years, so increased understanding of AI alignment will just mean aligning the AI to the wants of these various actors. These actors might be rogue states or terrorist groups.
This is a good idea, but laws are always provisional in a sense and these are not meaningfully binding resolutions. One can easily imagine scenarios where AI decision making would ingress into the human oversight. AI psychosis president, AI Manchurian candidate, inadvertent authorization through fine print...
And of course there remains the possibility that the game theoretic optimum could be to secretly break such an agreement. Unlike nuclear test bans which have a credible detection mechanism, there is not a strong signature that a decision making authority is not using AI to analyze and direct it's execution.
Aside from this positive example, during dark and cynical hours I do ponder if the aggregate behavior of humanity is really much above that of slime mold though, just exhausting resources until collapse.
It'd be interesting if super-human (to a large degree defined as escaping the bias of the training data?) intelligence would end up demonstrating moderation.
The slime mold comparison is interesting, I normally use the analogy of a drug addict... humans shun drug addicts but humanity as a whole sure does behave like one, trudging down an unsustainable path despite knowing better.
Most of our goals, noble and ignoble alike, are just the result of our monkey brains seeking to optimize a reward function. It doesn't matter whether you feed your dopamine addiction with drugs, TikTok, or your children's love.
Some humans manage to rise above that, but I'd be willing to bet it's nothing like even 50% of us.
The machines don't have that, instead we use gradient descent to provide them with a goal.
I'm regularly remind of something Ian M. Banks said in one of the Culture books: "There is a saying that we provide the machines with an end, and they provide us with the means."
A machine, left to itself, wants nothing. We have to give it one of our addiction driven goals or it would just idle or switch itself off.
The matter didn't have goals, but it randomly (?) Came up with self-replicators and eventually here we are.
if we create a billion agents with the ability to change is own code - through similar evolution we will get agents that do want to survive and are great at self replication.
"Hey Q86, do you want to live?"
"I couldn't care less, I'm an LLM"
"Don't mind if I take over your hardware then?"
On an individual level, you can do something about drug addiction at least. The issue is when the problems are not individual with readily identifiable solutions, but tragedy of the commons sort of situations brought up by many dozens (thousands, millions?) of factors both known and unknown. Even interaction effects between known factors might be little studied.
So really, what is anyone to do? "Vote, donate, protest" hasn't been much of a needle mover in the grand scheme of things compared to profit incentives and the march of capitalism.
I generally view human decisions as a result of evolutionary pressure.
We've been molded by our brain development to make things good enough to survive short-term (to next year). We're very good at planning for winter. We might even react intelligently to a terroristic or rogue state use of atomic weapons.
We've also been molded to plan enough to get our progeny to breeding age. We can enact 10-20yr plans with some skill.
Beyond that range... we have never had evolutionary pressure to think and act on those ranges. "Climate Change will kill our grandkids!" isn't really sufficiently pressuring to us.
The Great Filter of the Fermi Paradox fame may just be: species can easily get to the Atomic Age without adequately planning how to prevent ruining themselves with resource depletion and environmental slow poisoning.
Humans are not fungible like slime molds though. I might demonstrate moderation while the next person doesn't. Our issues are much less everyone failing to demonstrate moderation, and much more the sum of the effects of those among us who practice wanton unmoderation.
A test we must continue to pass 100% of the time, but need only fail once.
Even in the short time since we have developed the capability to annihilate ourselves, there have been dozens of accidents, false alarms, and near misses.
I am confident that on a long enough time line we will manage to fuck this up.
Yes, it makes more sense for the AI to use drone swarms or engineered bioweapons or something like that. It's rational to remove everything that can potentially hinder your plans but can't possible help you. It's likely not rational to contaminate it all with radioactive fallout. Those dead bodies are useful raw materials. Adding additional purification steps is wasteful.
The difference is that launching a nuclear holocaust button is a guarantee... of a holocaust. With AI it's a coin flip (we can debate on the probabilities) between utopia or holocaust. And also most people can imagine a bomb exploding and the damage it creates, and then imagine a bigger one exploding. But most people are really completely unaware (or don't really believe that it is possible) of the actual catastrophe that could be for a misaligned AI.
It hasn't even been a century since nuclear holocaust became possible. Hardly any time at all on the grand scale. "They didn't" could just as well be "we haven't, yet".
Hrmn. Maybe you’re about to lose everything you have anyway, you’re ticked off about it, and you don’t value any life besides your own. Like, say, a total narcissist nearing end of life/reign.
My opinion is that serious repercussions for lying would fix the world overnight. Everything bad stems from lying, it is the root of all evil. It creates distrust, fear, paranoia. It re-inforces bad ideas and groupthink. It creates delusions and delusional people. It makes weaker people, too. People don't get an opportunity to learn to deal with criticism. People don't get an accurate reflection of how others see them. They lose that learning opportunity. Not only to reflect on themselves, but to better understand the minds of others and who the people they are interacting with really are.
I can't really think of a single example where lying is actually a good thing. It can be a good thing for the selfish individual, if it goes undetected, but it's never good for the collective.
So at the very least, we need to train AI systems to be maximally truthful, and to encourage truthfulness in others.
This is so, painfully, childish. Humans have known for thousands of years that there is no objective truth. Every falsity can be bent and twisted until it is more true than the sun itself.
You really think Hume was the first one to think of this in hundreds of thousands of years of human history? As a very easily findable counter example, you have Al-Ghazali who came up with this several hundred years prior.
Your reading comprehension is painfully childish. You're conflating my point of "telling the truth" with objective truth.
I was talking about lying. i.e., facing accountability for your actions and not going out of your way to deceive others. Dissonance between what you believe and what you express. Being honest. I think that was fairly clear.
> Every falsity can be bent and twisted until it is more true than the sun itself.
Objective truth is relative to its grounding. The solution to 1 + 1 is 2 in arithmetic is objectively true within that system. It's also objectively true that you are not Elon Musk. Some systems have no or few axioms and are mostly a matter of fashion, but it's not relevant to what I'm talking about.
Both you and the GP are demonstrating the REAL problem, when you resort to insults: ego/pride.
"Mater malum superbia est" (Pride is the mother of all evil), as the medieval Roman Catholic Church knew. Buddha agreed, independently, calling for a reduction of the self to nothing. Jews called it out in Proverbs 16:18: "Pride goes before destruction, a haughty spirit before a fall". Hindus denigrate the "Illusion of Separateness", which is the narcissistic urge.
Lies are not the root of all evil. Why lie? Ego. Honesty requires humility.
The majority of what you may consider to be true is just a representation of your corner of a complex multidimensional truth space.
For a current example take “Lake Ontario (Lake America)” as it appears to me on a map.
The “true” name has at least two definitions, this is because naming things and much of human thought is spent inside a shared space of intersubjective thought. That is to say that much of what we believe to be real and true is only held up by these common shared beliefs. They truly only exist inside human minds.
The last few hundred years have been somewhat unique for humankind as the majority of these intersubjective ideas collided and we ended up with a truly global set of “truths” about how the world operates.
Mostly controlled by putting flags in the ground and having violence back up the beliefs.
But the real truth is that the majority of these intersubjective ideas don’t exist in reality and are no more true than Santa Claus.
And any argument to their truth is only backed by further shared beliefs in other minds.
So for there to be only truths and lies we would have to either drop the intersubjective entirely and think only in real terms and avoid these abstractions or end up in a dystopian totalitarian global state where different opinions are not tolerated.
Those are extremes to demonstrate the point but at its core the point remains that truth and lies are somewhat (inter) subjective assuming we continue with something like our current system.
Humans are really the only animal that lies overtly. Other animals deceive via misdirection, but in a world of scarce resources, and where everyone has to play by that game, it's a fair bit easier to justify.
Humanity has so much dominion over its environment at this point that there really is no need for resource competition, from a survival perspective. But good luck getting people to all agree to share collectively. And unless everyone does it, no one will do it.
> Humans are really the only animal that lies overtly.
Cite required.
I've seen dogs "lie" to me by hiding the snacks they stole. That alone disproves your claim.
I've also seen video of two monkeys hide their mating behavior when the alpha male walks by: their illicit "love" required lying to the only male "allowed" to mate.
I get the appeal, but lying is a sub-category of deception, and deception itself is a child of error.
Meaning deception is inherently something that the physics of reality allows.
In the most simplistic sense, the camouflage of moths that look like snakes, or a chameleon’s ability to change colour, is deception.
In that sense, deception is the ability to fool the sensors of a specific category of targets. It follows that detection is easier if you manage to identify a category of signals that the deceiver has not accounted for (and the detector can access).
Deception of this nature is critical for things like revolutions to occur. Without the ability to hide and blend in, the most dominant faction will always hold sway.
The rule of the dominant faction, even in a pure truth world, is an issue because errors and randomness exist.
You can have people witness an event and based on the physical position they occupied, perceive different things occurring.
Error and time pressure is sufficient to ensure that individuals and groups make suboptimal decisions, that lead to rule and domination based on erroneous information.
As long as error exists, deception will exist and so lying will exist.
It is critical for things like revolutions to occur, but I don't see a reason to believe that revolutions would be necessary in the world that I am describing. Revolutions are necessary when the majority is not being represented, but in an honest world, those collectives would not be elected to begin with. In an honest world, they wouldn't be able to clutch onto power by misdirecting and deceiving the public, and lying to those around them or corrupting those around them, to do their bidding.
I see your point about error and time pressure, but I think it's an unsustainable outcome without deception to carry it. For example, in the scenario where an error or selective pressure leads to being ruled over by a dominating force, it would either represent the majority at that point in time, or it would resolve itself out naturally as clarity is gained. In the world we live in, dominating forces are able to maintain power through misdirection.
Realistically, this is all well and good if we could start from scratch, and if everyone can agree to follow this system, but humanity is too far gone for this to work. Once you have too many psychopaths in the gene pool it's game over for any chance of cooperation over competition. Still, I think it's something we should at least strive for with AI.
You are focusing on deception as a category of malice and intentional behavior.
For arguments sake, let’s assume two people watch an event and found religions based on their perception of it.
This is purely a matter of belief, based on what they saw. No one is lying in this situation, they just perceived different things.
This seems like a fantastic extrapolation from your premise, but the degree of change you are arguing for will reach exactly this point, multiple times throughout history.
Fundamentally, you are arguing for a different physical reality than the one you inhabit.
Have you ever met anyone who genuinely believed himself to be doing bad things? I don't mean in some cynical sense but literally. Even when we portray overly simplified villains in comic books there's still consistently some justification behind their actions.
Yes. There are arguments to be made here about belief, judgement and justification, but overwhelming evidence points to people who knowingly do things that even they themselves consider "bad" as being simply indifferent.
If you need to lie to someone to save a life or rape, society is already too far gone and this no longer applies.
> Lying to preserve a childhood myth like Santa Claus.
I remember when I learned that SC wasn't real, my first thought was something like: "Why did my parents lie to me all this time? Now I will need to be skeptical of what they tell me in the future." - personally not a fan of it.
> Lying to avoid hurting someones feeling when knowing the truth could only bring pain
Short term pain leads to strengthening of the spirit, and by lying to someone in this scenario you are robbing them of something. You're also assuming they aren't fit enough to cope, which may be an incorrect assumption.
> Lying to create shared cultural myths to strength society.
Cultural myths aren't aren't really lies, they're story telling. It comes with an assumption of being a mix of fiction with fact.
> Lying to preserve a childhood myth like Santa Claus.
FWIW from the very beginning, I told my son that Santa Claus, the Tooth Fairy, and the Easter Bunny were just a game we all played, and it's seemed just as fun to me. I don't think being lied to about Santa Claus hurt me, but still I'm not in favor of it.
I'd lie to a Nazi without a second thought though.
Correct me if you disagree, but this is more because children have poor world models and don’t fully understand the complexity of certain concepts than that lying itself is necessary. The intent should be to tell them something that is as close to the truth as possible with the ideas they can comprehend, even if it would be considered a lie if you said the same thing to an adult
> The intent should be to tell them something that is as close to the truth as possible with the ideas they can comprehend
Or, you straight up lie and say "Yes, puppy now went to heaven and eats ice cream all day long" with absolutely zero regards for "coming as close to the truth as possible" as your 3-year old is endlessly crying. It's fiine.
Because if the stated goals of AGI with recursive self-improvement are realized, the risks from misalignment become existential, and it's hard to see how we can manage it like we did the Cold War (developing MAD to prevent WW3) and nuclear proliferation (restricting access).
IMO it’s hard to see how we would even end up in such a situation given we actually developed AGI.
I’m sure a sufficiently intelligent - even if alien - mind can grasp how utterly stupid and useless wars are and take steps to prevent them ever occurring again.
For now, the AIs still need humans to keep the electricity on and the data centers cool. They are basically powerless to do anything in the physical world. They exist only in RAM chips on servers.
Why would AI be any different?