The question of whether RLVR can generalize so strongly that spending trillions on RL environments could produce fully human-like general intelligence is an empirical question; Dario's quote that model performance degrades when serving at longer context lengths than trained suggests short-horizon RL training may not generalize to long-horizon performance.

causalpending

Speaker

Unidentified Speaker — What does the next training paradigm look like? [20p5-kQXF_Q]

Evidence Quote

Now, maybe I'm reading too much into this, but it seems like he's saying that short-horizon RL training doesn't necessarily generalize to long-horizon RL performance.

Source

What does the next training paradigm look like?Dwarkesh Patel
Created: 8/12/2026, 6:40:37 PM

My Notes

Loading notes...