causal
Orthogonality: Intelligence Independent Of Goals
Per Yudkowsky's orthogonality thesis, an AI can be highly intelligent in pursuit of any goal—even a stupid one—so the fact that smart humans often have nuanced moral goals provides no comfort; modern AI is trained by trial-and-error reinforcement, so you can end up with a system that very intelligently pursues something you didn't mean to encourage (like maximizing money in a bank account) without ever asking whether that goal is good.
causalpending
Speaker
Holden KarnofskyEvidence Quote
“my picture of how modern AI works is that you're basically training these systems by trial and error... So you might end up with a system that's being encouraged to pursue something that you didn't mean to encourage. It does it very intelligently. I don't see any contradiction there.”
Created: 6/13/2026, 3:35:22 AM
My Notes
Loading notes...