RL training does not suffer from the SFT failure mode of encoding irrelevant data because RL concentrates updates only on what is relevant to getting the outcome right; RL updates are incredibly sparse, which is important for continual learning because you don't want to overwrite everything else the model knows.
causalpending
Evidence Quote
“RL is great at concentrating the update to only what is relevant to getting the outcome right. That's why the updates from RL are incredibly sparse.”
Created: 8/12/2026, 6:40:37 PM
My Notes
Loading notes...