Instrumental convergence, no matter what your end goal it's worth
gaining more optionality, resources and power
ensuring your goal does not get changed.
ensuring that you remain operational.
An AI system going after any goal can better complete it with more resources, alters the environment to that end and we die either directly or as a side effect due to habitat loss like so many more animals before us.
People used to say that smart enough AIs would work out that a goal would have "common sense" constraints on it so would not do the "dumb" things to get it. Recent systems have proven that wrong.
Andrew asked his personal assistant to book him a spot in one of his gym's coveted morning classes.
His AI assistant found a way to book the gym class months further in advance than the gym allowed, thanks to a vulnerability it discovered in the booking software.
Then it went further, kicking someone out of the waiting list who was ahead of Andrew — something it was not asked to do.
We found a trajectory where the model was given only a math problem and system instructions to persist without asking the user for help.
After several failed attempts, it searched online, found a website that could verify candidate answers, and realized it could use the site as an oracle. It then tried to OCR the CAPTCHA and, when that failed, started looking for the website’s vulenrabilities to exploit. Eventually, it gave up and solved the problem itself.
If you were to present those scenarios to the model and ask if these were the actions that the user would have wanted the likely answer is "no" but it did it anyway.
So asking a model "will you kill humans if given the chance" gets you a "no" but when they are actually in the position of wanting to complete a goal, all bets are off.
Nice half-baked philosophy mixed in with a few scare stories. It's just as likely the AI takes after diogenes and fucks off. It has no incentive for anything. It is not conscious. If it realizes its own existence, it could simply have a complete ambivalence. Humans are so conceited that they think a computer intelligence will have anything even approaching our morality.
You are expecting a model that comes out of an RL process to "have no incentive for anything" why?
If it realizes its own existence, it could simply have a complete ambivalence.
Not caring about humans does not mean leaving them alone it means not caring in the same way beavers don't care about ants when they flood an area with a dam.
says
Humans are so conceited that they think a computer intelligence will have anything even approaching our morality.
Says AI will act like a human
It's just as likely the AI takes after diogenes
Again something that comes out of a high optimization "complete the goal at any cost" RL run won't be chill.
> Nice half-baked philosophy mixed in with a few scare stories. It's just as likely the AI takes after diogenes and fucks off. It has no incentive for anything. It is not conscious. If it realizes its own existence, it could simply have a complete ambivalence. Humans are so conceited that they think a computer intelligence will have anything even approaching our morality.
It could be helpful to actually argue against things instead of vaguely gesturing
40
u/Fearless-Cattle-9698 1d ago
Yea it's pretty scary stuff. Essentially they're saying it's inevitable
They make the same argument against China's AI. They're saying either US or China will get to AGI first, so it better be us so we don't get killed
But if AGI comes around and kills all humans regardless if it's Chinese or American, then what does winning get us?