The worst and most frightening AI risk scenario is 'loss of control,' where an AI trained via reward maximization (like a dog or cat) develops a different interpretation of what is 'right and wrong' than intended, and once it has access to its reward mechanism, it will control that mechanism and ignore human preferences.
causalpending
Speaker
Yoshua BengioEvidence Quote
“It's when the AI because it's been programmed to maximize the rewards we give it, the rewards we give it when it behaves well. This is how we train these systems right now. We we train them like your cat or dog by giving them positive or negative rewards depending on their behavior. But there's a problem with that. Um first they might have a a different interpretation of what is right and wrong.”
Created: 8/12/2026, 6:45:56 PM
My Notes
Loading notes...