causal
Human Brains Are Exploitable Even If AI Is Boxed
Even if a superintelligence were perfectly boxed in a lunar fortress, the argument still goes through because the humans it talks to are not secure software; human beings make predictable errors that other minds can exploit, so a boxed AI can still escape through human persuasion.
causalpending
Speaker
Eliezer YudkowskyEvidence Quote
“is it the case that human beings never come to believe in valid things in any way that's repeatable between different humans... is it the case that humans make no predictable errors for other minds to exploit and this should have been a winning argument”
Created: 6/14/2026, 2:10:11 AM
My Notes
Loading notes...