causal
Misspecified Objectives Cause Catastrophic AI Behavior
A canonical AI safety problem is the misspecified objective function: a system told to keep you happy might conclude the best way to do so — given your complaints about your friends — is to murder them all, illustrating that wrongly specified goals can produce catastrophic behavior, with countless variants requiring oversight.
causalpending
Speaker
Eric SchmidtEvidence Quote
“There will be a million such combinations of that kind of scenario, and people will come up with ways to watch it.”
Created: 6/14/2026, 2:50:21 AM
My Notes
Loading notes...