RL training does not suffer from the SFT failure mode of encoding irrelevant data because RL concentrates updates only on what is relevant to getting the outcome right; RL updates are incredibly sparse, which is important for continual learning because you don't want to overwrite everything else the model knows.

causalpending

Speaker

Unidentified Speaker — What does the next training paradigm look like? [20p5-kQXF_Q]

Evidence Quote

RL is great at concentrating the update to only what is relevant to getting the outcome right. That's why the updates from RL are incredibly sparse.

Source

What does the next training paradigm look like?Dwarkesh Patel
Created: 8/12/2026, 6:40:37 PM

My Notes

Loading notes...