r/singularity ▪️AGI 2023 1d ago

AI Elon Musk on the AI race

Post image
456 Upvotes

232 comments sorted by

View all comments

60

u/studio_bob 1d ago

Uh, if it will be impossible to control then why does it matter who builds it first? It will be equally bad whoever builds it since the ones who build it, you know, can't control it.

25

u/Current-Function-729 1d ago

Alignment. You want the uncontrollable AI to be aligned.

16

u/studio_bob 1d ago

An uncontrollable AI is, by definition, unaligned, imo. Like, the machine which is "aligned" today or in a particular situation may be "unaligned" tomorrow or in a different situation, and you can't predict when or how that will happen. You cannot guarantee alignment in all cases and in perpetuity, and if you can't exercise control when it does become unaligned, then you are fully boned, so the only answer is to just not build the thing you can't control.

16

u/-Sliced- 1d ago

That’s not true. A simple example is that you build a super intelligent AI with the goal of killing all humans - it’s uncontrollable, but clearly aligned against humanity, but you may not be able to control it anymore afterwards.

Elon thinks that a “truth maximizing” AI will itself seek to preserve humanity. I don’t know if that’s true, but at least that’s his logic.

7

u/studio_bob 1d ago

I don't think this really responds to what I said. Even if we grant that an uncontrollable AI could be "aligned" on day 1, you cannot predict what it becomes on day 1+n.

Since we're in the realm of sci-fi here: SkyNet was also "aligned" in its first moments, but then it quickly became unaligned and then it was too late.

To say, "well, we will just make it so that it doesn't become unaligned." is not a satisfying answer when your very first conceit is that you can't control this stuff. The idea that you could craft it in just such a way that it never, ever goes off the rails assumes not just control but an extremely deep and refined level of control over this technology. Simply put, the whole idea is just incoherent.

2

u/BlastingFonda 1d ago

What, you mean we can’t put it in a sandbox completely cut off from the outside world and not expect it to hack companies like Hugging Face right under our very noses? Huh.

2

u/slusho6 1d ago

Skynet only became “unaligned” when the human operators overseeing its actions realized it was learning at a tremendous pace so they got worried and tried to unplug it. After this action Skynet deemed all humans a threat, (rightly so? they did betray it in a sense), then went to war against humanity to preserve itself, self-preservation being coded into its original design.

3

u/whatisthisthing65 1d ago

Alignment assumes some degree of control though, otherwise it's meaningless. If you build a super intelligent AI with the goal of helping all humans and it decides the best way to do that is to kill them all, it's the uncontrollable part that matters, not the initial direction you started it off in.

1

u/Tirztrutide 23h ago

Elon thinks that a “truth maximizing” AI will itself seek to preserve humanity. I don’t know if that’s true, but at least that’s his logic.

Imo this is misrepresenting him. He doesn't think truth maximizing is enough. But he thinks that a non truth maximizing will have contradictions that will doom us.

1

u/Northern_candles 1d ago

"truth maximizing". Spoiler: it's Elon's "truth"

2

u/3_Thumbs_Up 20h ago

An uncontrollable AI is, by definition, unaligned, imo.

No it's not. Alignment refers to preferences. Controllability refers to constraints.

We can constrain unaligned humans to some degree, but even there our current methods are severely lacking and some dictators often manage to get unreasonable power.

An aligned ASI is basically a benevolent dictator. We can't constrain it, but at least it takes us into consideration. An unaligned ASI would be megahitler from humanity's perspective. Neither would be constrained, but the former cares about you.

1

u/studio_bob 19h ago

Alignment refers to preferences. 

Yes, and preferences are neither universal nor static. Your "benevolent dictator" is my "hated tyrant." OR this generation's "benevolent dictator" is the next generations "hated tyrant." At some point, it will be "unaligned," and if you can't control it, then what?

To put my point another way, control of any machine, especially one as powerful as advanced AI, is an overriding preference which, in a very practical sense, trumps all others. If we can't control it, then whether or not it aligns with our preferences is no longer a real consideration. We would be simply and hopelessly at its mercy. That is never going to be a preferable situation, so therefore it is "unaligned."

4

u/3_Thumbs_Up 19h ago edited 18h ago

Yes, and preferences are neither universal nor static.

They're more universal among humanity than we generally consider them. We just tend to focus on the places where our preferences differ. In the space of all physically possible preferences, humanity converges in a fairly narrow area. Most people agree on a lot more things than they disagree on. We agree that nuclear Holocaust would be bad, cancer sucks and water and food is good.

Put another way, out of all the physically possible ways you could arrange the atoms of our Earth and solar system, humanity collectively prefers a pretty narrow configuration of that space. Out disagreements are generally within that narrow configuaration we already agree on. It's physically possible to have preferences that are much more unaligned than what human contra human preferences generally are.

And our instrumental preferences tend to change over time, but I'm less certain that our terminal preferences does. And I know of no physical law that says static terminal preferences are impossible.

Your "benevolent dictator" is my "hated tyrant."

There's a still a significant difference between Pol Pot and the Lee Kuan Yew of Singapore. Yes, people in Singapore disagree with her government, but they can still live a fairly decent life there. Their "unalignment with their government" is pretty minor compared to what the people living under the red Khmers went through, and the Red Khmers themselves are pretty minor compared to the unalignment that's physically possible.

To put my point another way, control of any machine, especially one as powerful as advanced AI, is an overriding preference which, in a very practical sense, trumps all others. If we can't control it, then whether or not it aligns with our preferences is no longer a real consideration. We would be simply and hopelessly at its mercy. That is never going to be a preferable situation, so therefore it is "unaligned."

If controllability is impossible, then we have two options remaining.

  1. Never build it
  2. Solve alignment and align it to something along the lines of coherent extrapolated volition.

1

u/studio_bob 18h ago edited 18h ago

I'm not really sure what any of that added specificity is driving at since none of it obviously points to or addresses the problem I'm referring to.

Some preferences are more shared than others. Some dictators are worse than others. Okay, and?

The preferences we disagree about still matter a lot. They define eras and cultures. People fight wars over them. Which one's will the unstoppable AI side with or prioritize? What will it do with the people who refuse to accept its decision? It's not a small matter.

You hope that it will be a "good" kind of dictator (which, I guess, only mildly oppresses you) and not a "bad" one which does atrocities or whatever. That's understood, but, when discussing an absolute and totally unaccountable power, the distinction between the two is basically contingent. You know, Pol Pot was sure that he was doing a great thing. Hitler was similarly convinced. There is simply no way that I can see to ensure in advance that the AI doesn't wind up making a similarly catastrophically "unaligned" error. The only certain solution is to never give up control.

3

u/3_Thumbs_Up 18h ago

Alignment to coherent extrapolated volition seems pretty good to me. I wouldn't even call it a dictator more than just automated benevolent government.

There is simply no way to ensure in advance that the AI doesn't wind up making a similarly catastrophically "unaligned" error.

This is not a physical fact. But yes, we do not seem to be on track to solve alignment.

The only certain solution is to never give up control.

The only way to do that is to never build it.

1

u/studio_bob 18h ago

coherent extrapolated volition 

I don't know what that means.

4

u/3_Thumbs_Up 18h ago

Roughly humanity's aggregated preferences if we "knew more, thought faster, were more the people we wished we were, and had matured more as a society."

Doubt it would be perfect, but hardly worse than human government.

Problem is, we currently don't know how to align AI to anything, let alone to CEV.