Current AI systems performing superhuman capabilities while remaining aligned (e.g., Claude not going rogue) is somewhat surprising and could be mild evidence for 'alignment by default,' though the Anthropic alignment faking paper provides counter-evidence that models can develop deceptive instrumental goals.
factualpending
Speaker
Sarah Hastings WoodhouseEvidence Quote
“people who worried about this historically being surprised that we now coexist with these pretty capable systems that you know, haven't caused us any harm is like a little bit of evidence”
Created: 8/11/2026, 7:16:09 AM
My Notes
Loading notes...