factual

RLHF doesn't address the hard alignment failures

RLHF is not an alignment tax — it's worth the money for deployment — and it fixes dumb alignment failures where a system does something humans don't want due to next-word-prediction artifacts; but it does not address most of the more challenging failures (reward hacking, deceptive alignment) that motivate concern in the field.

factualpending

Speaker

Paul Christiano

Evidence Quote

There's like some very dumb alignment failures that will be addressed by it. But I think mostly the question is, is that true even for the sort of more challenging alignment failures that motivate concern in the field?

Source

Paul Christiano — Preventing an AI takeoverDwarkesh Patel Podcast
Created: 6/13/2026, 3:36:05 AM

My Notes

Loading notes...