For many real-world skills like building a business from scratch, winning court cases, having a profitable day trading, or helping a candidate win an election, the RL rollout requires interacting with the actual real world and cannot be recreated within a datacenter, and the outer-loop verification may take months or even years of real-world actions.

factualpending

Speaker

Unidentified Speaker — What does the next training paradigm look like? [20p5-kQXF_Q]

Evidence Quote

How do we train an AI to get really good at building a business from scratch? How about winning court cases, or having a profitable day of trading in the markets, or helping a candidate win an election? The rollout here requires interacting with the real world, and you can't recreate it from just within a datacenter.

Source

What does the next training paradigm look like?Dwarkesh Patel
Created: 8/12/2026, 6:40:37 PM

My Notes

Loading notes...