A couple years after DeepMind released AlphaZero, researchers trained EfficientZero to be very data-efficient, where if given two hours to play an unseen Atari game, the model would beat a novice human, but this doesn't necessarily mean the model was more sample-efficient than humans—it depends on how you measure, because EfficientZero plays dozens of simulated games in its head for each real-world game step.

causalpending

Speaker

Unidentified Speaker — What does the next training paradigm look like? [20p5-kQXF_Q]

Evidence Quote

Because for each step in the real game, EfficientZero is playing dozens of simulated games in its head.

Source

What does the next training paradigm look like?Dwarkesh Patel
Created: 8/12/2026, 6:40:37 PM

My Notes

Loading notes...