normative

How to detect alignment bullshit

Christiano's heuristic for evaluating alignment proposals: the easiest red flag is work on a problem that isn't actually important and has no story for becoming important (not a problem now and won't worsen, or is a problem now but clearly getting better); he is skeptical of dismissals based on 'doesn't deal with the key difficulty,' and is otherwise liberal, requiring only that work engage real models empirically and have a coherent story about key difficulties.

normativepending

Speaker

Paul Christiano

Evidence Quote

I tend to be optimistic when people dismiss something because this doesn't deal with a key difficulty or this runs into the following insurable obstacle. I tend to be a little bit more skeptical about those arguments

Source

Paul Christiano — Preventing an AI takeoverDwarkesh Patel Podcast
Created: 6/13/2026, 3:36:05 AM

My Notes

Loading notes...