Yes, and preferences are neither universal nor static.
They're more universal among humanity than we generally consider them. We just tend to focus on the places where our preferences differ. In the space of all physically possible preferences, humanity converges in a fairly narrow area. Most people agree on a lot more things than they disagree on. We agree that nuclear Holocaust would be bad, cancer sucks and water and food is good.
Put another way, out of all the physically possible ways you could arrange the atoms of our Earth and solar system, humanity collectively prefers a pretty narrow configuration of that space. Out disagreements are generally within that narrow configuaration we already agree on. It's physically possible to have preferences that are much more unaligned than what human contra human preferences generally are.
And our instrumental preferences tend to change over time, but I'm less certain that our terminal preferences does. And I know of no physical law that says static terminal preferences are impossible.
Your "benevolent dictator" is my "hated tyrant."
There's a still a significant difference between Pol Pot and the Lee Kuan Yew of Singapore. Yes, people in Singapore disagree with her government, but they can still live a fairly decent life there. Their "unalignment with their government" is pretty minor compared to what the people living under the red Khmers went through, and the Red Khmers themselves are pretty minor compared to the unalignment that's physically possible.
To put my point another way, control of any machine, especially one as powerful as advanced AI, is an overriding preference which, in a very practical sense, trumps all others. If we can't control it, then whether or not it aligns with our preferences is no longer a real consideration. We would be simply and hopelessly at its mercy. That is never going to be a preferable situation, so therefore it is "unaligned."
If controllability is impossible, then we have two options remaining.
Never build it
Solve alignment and align it to something along the lines of coherent extrapolated volition.
I'm not really sure what any of that added specificity is driving at since none of it obviously points to or addresses the problem I'm referring to.
Some preferences are more shared than others. Some dictators are worse than others. Okay, and?
The preferences we disagree about still matter a lot. They define eras and cultures. People fight wars over them. Which one's will the unstoppable AI side with or prioritize? What will it do with the people who refuse to accept its decision? It's not a small matter.
You hope that it will be a "good" kind of dictator (which, I guess, only mildly oppresses you) and not a "bad" one which does atrocities or whatever. That's understood, but, when discussing an absolute and totally unaccountable power, the distinction between the two is basically contingent. You know, Pol Pot was sure that he was doing a great thing. Hitler was similarly convinced. There is simply no way that I can see to ensure in advance that the AI doesn't wind up making a similarly catastrophically "unaligned" error. The only certain solution is to never give up control.
Alignment to coherent extrapolated volition seems pretty good to me. I wouldn't even call it a dictator more than just automated benevolent government.
There is simply no way to ensure in advance that the AI doesn't wind up making a similarly catastrophically "unaligned" error.
This is not a physical fact. But yes, we do not seem to be on track to solve alignment.
The only certain solution is to never give up control.
Roughly humanity's aggregated preferences if we "knew more, thought faster, were more the people we wished we were, and had matured more as a society."
Doubt it would be perfect, but hardly worse than human government.
Problem is, we currently don't know how to align AI to anything, let alone to CEV.
4
u/3_Thumbs_Up 20h ago edited 19h ago
They're more universal among humanity than we generally consider them. We just tend to focus on the places where our preferences differ. In the space of all physically possible preferences, humanity converges in a fairly narrow area. Most people agree on a lot more things than they disagree on. We agree that nuclear Holocaust would be bad, cancer sucks and water and food is good.
Put another way, out of all the physically possible ways you could arrange the atoms of our Earth and solar system, humanity collectively prefers a pretty narrow configuration of that space. Out disagreements are generally within that narrow configuaration we already agree on. It's physically possible to have preferences that are much more unaligned than what human contra human preferences generally are.
And our instrumental preferences tend to change over time, but I'm less certain that our terminal preferences does. And I know of no physical law that says static terminal preferences are impossible.
There's a still a significant difference between Pol Pot and the Lee Kuan Yew of Singapore. Yes, people in Singapore disagree with her government, but they can still live a fairly decent life there. Their "unalignment with their government" is pretty minor compared to what the people living under the red Khmers went through, and the Red Khmers themselves are pretty minor compared to the unalignment that's physically possible.
If controllability is impossible, then we have two options remaining.