causal

Alignment research reduces moral as well as safety risk

Understanding and controlling the systems you build is, on balance, good both for safety and for morality, because the main effect of research into whether an AI resents humans is to avoid building such an AI in the first place — everyone should want to know if the systems they build feel that way.

causalpending

Speaker

Paul Christiano

Evidence Quote

Everyone should be very unhappy if you built a bunch of AIS who are like, I really hate these humans, but they will murder me if I don't do what they want.

Source

Paul Christiano — Preventing an AI takeoverDwarkesh Patel Podcast
Created: 6/13/2026, 3:36:05 AM

My Notes

Loading notes...