normative
Two lines of alignment defense for deployment
Before broadly deploying a human-level model, you want defense-in-depth: a first line of adversarial evaluation and monitoring that could detect or prevent catastrophic harm (testing in diverse situations and arguing the AI can't distinguish lab tests from reality), and a second line that determines whether dangerous misalignment can occur at all by trying to produce reward-hacking or deceptive alignment in the lab under optimal conditions.
normativepending
Speaker
Paul ChristianoEvidence Quote
“we have tried to test our system in a broad diversity of situations that reflect cases where it might cause harm... and then we have tried to argue that our AI is actually like, those tests are indicative of the real world.”
Created: 6/13/2026, 3:36:05 AM
My Notes
Loading notes...