causal
Alignment research reduces moral as well as safety risk
Understanding and controlling the systems you build is, on balance, good both for safety and for morality, because the main effect of research into whether an AI resents humans is to avoid building such an AI in the first place — everyone should want to know if the systems they build feel that way.
causalpending
Speaker
Paul ChristianoEvidence Quote
“Everyone should be very unhappy if you built a bunch of AIS who are like, I really hate these humans, but they will murder me if I don't do what they want.”
Created: 6/13/2026, 3:36:05 AM
My Notes
Loading notes...