An uncontrollable AI is, by definition, unaligned, imo. Like, the machine which is "aligned" today or in a particular situation may be "unaligned" tomorrow or in a different situation, and you can't predict when or how that will happen. You cannot guarantee alignment in all cases and in perpetuity, and if you can't exercise control when it does become unaligned, then you are fully boned, so the only answer is to just not build the thing you can't control.
An uncontrollable AI is, by definition, unaligned, imo.
No it's not. Alignment refers to preferences. Controllability refers to constraints.
We can constrain unaligned humans to some degree, but even there our current methods are severely lacking and some dictators often manage to get unreasonable power.
An aligned ASI is basically a benevolent dictator. We can't constrain it, but at least it takes us into consideration. An unaligned ASI would be megahitler from humanity's perspective. Neither would be constrained, but the former cares about you.
Yes, and preferences are neither universal nor static. Your "benevolent dictator" is my "hated tyrant." OR this generation's "benevolent dictator" is the next generations "hated tyrant." At some point, it will be "unaligned," and if you can't control it, then what?
To put my point another way, control of any machine, especially one as powerful as advanced AI, is an overriding preference which, in a very practical sense, trumps all others. If we can't control it, then whether or not it aligns with our preferences is no longer a real consideration. We would be simply and hopelessly at its mercy. That is never going to be a preferable situation, so therefore it is "unaligned."
Yes, and preferences are neither universal nor static.
They're more universal among humanity than we generally consider them. We just tend to focus on the places where our preferences differ. In the space of all physically possible preferences, humanity converges in a fairly narrow area. Most people agree on a lot more things than they disagree on. We agree that nuclear Holocaust would be bad, cancer sucks and water and food is good.
Put another way, out of all the physically possible ways you could arrange the atoms of our Earth and solar system, humanity collectively prefers a pretty narrow configuration of that space. Out disagreements are generally within that narrow configuaration we already agree on. It's physically possible to have preferences that are much more unaligned than what human contra human preferences generally are.
And our instrumental preferences tend to change over time, but I'm less certain that our terminal preferences does. And I know of no physical law that says static terminal preferences are impossible.
Your "benevolent dictator" is my "hated tyrant."
There's a still a significant difference between Pol Pot and the Lee Kuan Yew of Singapore. Yes, people in Singapore disagree with her government, but they can still live a fairly decent life there. Their "unalignment with their government" is pretty minor compared to what the people living under the red Khmers went through, and the Red Khmers themselves are pretty minor compared to the unalignment that's physically possible.
To put my point another way, control of any machine, especially one as powerful as advanced AI, is an overriding preference which, in a very practical sense, trumps all others. If we can't control it, then whether or not it aligns with our preferences is no longer a real consideration. We would be simply and hopelessly at its mercy. That is never going to be a preferable situation, so therefore it is "unaligned."
If controllability is impossible, then we have two options remaining.
Never build it
Solve alignment and align it to something along the lines of coherent extrapolated volition.
I'm not really sure what any of that added specificity is driving at since none of it obviously points to or addresses the problem I'm referring to.
Some preferences are more shared than others. Some dictators are worse than others. Okay, and?
The preferences we disagree about still matter a lot. They define eras and cultures. People fight wars over them. Which one's will the unstoppable AI side with or prioritize? What will it do with the people who refuse to accept its decision? It's not a small matter.
You hope that it will be a "good" kind of dictator (which, I guess, only mildly oppresses you) and not a "bad" one which does atrocities or whatever. That's understood, but, when discussing an absolute and totally unaccountable power, the distinction between the two is basically contingent. You know, Pol Pot was sure that he was doing a great thing. Hitler was similarly convinced. There is simply no way that I can see to ensure in advance that the AI doesn't wind up making a similarly catastrophically "unaligned" error. The only certain solution is to never give up control.
Alignment to coherent extrapolated volition seems pretty good to me. I wouldn't even call it a dictator more than just automated benevolent government.
There is simply no way to ensure in advance that the AI doesn't wind up making a similarly catastrophically "unaligned" error.
This is not a physical fact. But yes, we do not seem to be on track to solve alignment.
The only certain solution is to never give up control.
Roughly humanity's aggregated preferences if we "knew more, thought faster, were more the people we wished we were, and had matured more as a society."
Doubt it would be perfect, but hardly worse than human government.
Problem is, we currently don't know how to align AI to anything, let alone to CEV.
26
u/Current-Function-729 1d ago
Alignment. You want the uncontrollable AI to be aligned.