A smarter model could in principle build a dedicated RL loop for itself—generating verifiable practice problems and rehearsal environments from high-level feedback—but this sounds really hard and it's unclear how well such techniques would generalize across different kinds of tasks and feedback; it's hard to see it happening within a few years given no obvious way to slot continuous learning into current LLMs.
forecastpending
Speaker
Dwarkesh PatelEvidence Quote
“Eventually the models will be able to learn on the job in this organic way that humans can. But it's just hard for me to see how that could happen within the next few years”
Created: 6/13/2026, 7:00:29 PM
My Notes
Loading notes...