causal

Orthogonality: Intelligence Independent Of Goals

Per Yudkowsky's orthogonality thesis, an AI can be highly intelligent in pursuit of any goal—even a stupid one—so the fact that smart humans often have nuanced moral goals provides no comfort; modern AI is trained by trial-and-error reinforcement, so you can end up with a system that very intelligently pursues something you didn't mean to encourage (like maximizing money in a bank account) without ever asking whether that goal is good.

causalpending

Speaker

Holden Karnofsky

Evidence Quote

my picture of how modern AI works is that you're basically training these systems by trial and error... So you might end up with a system that's being encouraged to pursue something that you didn't mean to encourage. It does it very intelligently. I don't see any contradiction there.

Source

Holden Karnofsky — History's most important centuryDwarkesh Patel Podcast
Created: 6/13/2026, 3:35:22 AM

My Notes

Loading notes...