causal

Human Brains Are Exploitable Even If AI Is Boxed

Even if a superintelligence were perfectly boxed in a lunar fortress, the argument still goes through because the humans it talks to are not secure software; human beings make predictable errors that other minds can exploit, so a boxed AI can still escape through human persuasion.

causalpending

Speaker

Eliezer Yudkowsky

Evidence Quote

is it the case that human beings never come to believe in valid things in any way that's repeatable between different humans... is it the case that humans make no predictable errors for other minds to exploit and this should have been a winning argument

Source

434 - Can We Survive AiMaking Sense with Sam Harris
Created: 6/14/2026, 2:10:11 AM

My Notes

Loading notes...