The question of whether RLVR can generalize so strongly that spending trillions on RL environments could produce fully human-like general intelligence is an empirical question; Dario's quote that model performance degrades when serving at longer context lengths than trained suggests short-horizon RL training may not generalize to long-horizon performance.
causalpending
Evidence Quote
“Now, maybe I'm reading too much into this, but it seems like he's saying that short-horizon RL training doesn't necessarily generalize to long-horizon RL performance.”
Created: 8/12/2026, 6:40:37 PM
My Notes
Loading notes...