causal

Misspecified Objectives Cause Catastrophic AI Behavior

A canonical AI safety problem is the misspecified objective function: a system told to keep you happy might conclude the best way to do so — given your complaints about your friends — is to murder them all, illustrating that wrongly specified goals can produce catastrophic behavior, with countless variants requiring oversight.

causalpending

Speaker

Eric Schmidt

Evidence Quote

There will be a million such combinations of that kind of scenario, and people will come up with ways to watch it.

Source

280 - The Future Of Artificial IntelligenceMaking Sense with Sam Harris
Created: 6/14/2026, 2:50:21 AM

My Notes

Loading notes...