Unidentified Speaker — What does the next training paradigm look like? [20p5-kQXF_Q]
Extraction couldn't determine who this speaker is in What does the next training paradigm look like?. If you recognize them, use "Identify this speaker" above to merge their claims onto the real person.
About
Speaker whose identity could not be determined from "What does the next training paradigm look like?".
Cast within
No topic-region cast yet — this appears once Unidentified Speaker — What does the next training paradigm look like? [20p5-kQXF_Q]'s compiled claims are aligned into a topic region's argument tree.
Claims by Unidentified Speaker — What does the next training paradigm look like? [20p5-kQXF_Q] (20 of 21)
Just as pretraining created a base intelligence smart enough to become a competent agent with enough RLVR on top, RLVR has created an agent competent enough to actually be broadly deployed in the world, and from this broad deployment the agent can learn on the job once the training recipe for continual learning arrives.
All major AI labs are making a research bet that training AIs on millions of verifiable tasks across thousands of diverse RL environments will produce AGI, because such training creates general problem-solving agents capable of making progress on open-ended tasks for weeks despite errors, ambiguity, and mistakes.
Computer use has made much slower progress than other verifiable domains like coding and math, despite being clearly verifiable (e.g., 'did the package get delivered?', 'is the venue booked?', 'have taxes been submitted?'), which is unusual and reveals something important about AI progress constraints.
For many real-world skills like building a business from scratch, winning court cases, having a profitable day trading, or helping a candidate win an election, the RL rollout requires interacting with the actual real world and cannot be recreated within a datacenter, and the outer-loop verification may take months or even years of real-world actions.
The most valuable bits of information that a model could learn from are revealed only during deployment, including what organizations are actually using the model for, what kinds of mistakes the model makes in the real world, and domain-specific tacit knowledge; yet deployed AIs are not able to use this information to improve themselves.
The current approach is analogous to having a genius graduate student who has never been allowed to take a real internship, given only classroom case studies in the form of RL training on environments, even though the student could be deployed in the economy and privy to so much domain- and organization-specific tacit knowledge.
Continual learning requires going back to update the weights; AIs cannot simply keep building up a larger and larger KV cache as they learn from more users, because that is not scalable and is not how humans learn—humans have no clean separation between parameters and activations, and the brain doesn't expand as we learn.
RL training does not suffer from the SFT failure mode of encoding irrelevant data because RL concentrates updates only on what is relevant to getting the outcome right; RL updates are incredibly sparse, which is important for continual learning because you don't want to overwrite everything else the model knows.
Another more speculative idea for addressing sample efficiency is 'dreaming'—where AIs build a good simulation of reality to rehearse new skills, try alternative strategies, and reinforce what works, allowing AIs to experience orders of magnitude more simulated samples in the same wall-clock time.
A couple years after DeepMind released AlphaZero, researchers trained EfficientZero to be very data-efficient, where if given two hours to play an unseen Atari game, the model would beat a novice human, but this doesn't necessarily mean the model was more sample-efficient than humans—it depends on how you measure, because EfficientZero plays dozens of simulated games in its head for each real-world game step.
After a week of work at a deployment, the AI receives a thumbs up or thumbs down (work review), and the base model distills everything learned during the session via OPSD, dreaming, or combinations thereof, allowing the AI to get better at domains adjacent to what it was trained for, starting a cycle of expanding capabilities.
The question of whether RLVR can generalize so strongly that spending trillions on RL environments could produce fully human-like general intelligence is an empirical question; Dario's quote that model performance degrades when serving at longer context lengths than trained suggests short-horizon RL training may not generalize to long-horizon performance.
Every time a user interacts with an AI, the AI will be smarter not only because it's learning from that user's previous sessions, but also because it's learning from all other interactions across the entire user base—this is very scary, exciting, and different from how AI currently improves.
My Notes
Loading notes...