factual
RLHF doesn't address the hard alignment failures
RLHF is not an alignment tax — it's worth the money for deployment — and it fixes dumb alignment failures where a system does something humans don't want due to next-word-prediction artifacts; but it does not address most of the more challenging failures (reward hacking, deceptive alignment) that motivate concern in the field.
factualpending
Speaker
Paul ChristianoEvidence Quote
“There's like some very dumb alignment failures that will be addressed by it. But I think mostly the question is, is that true even for the sort of more challenging alignment failures that motivate concern in the field?”
Created: 6/13/2026, 3:36:05 AM
My Notes
Loading notes...