factual
RLHF data is mostly generated by AI
In reinforcement learning from human feedback, humans are used only to train the reward function; once trained, the reward function's interaction with the model is automatic, so most of the data generated during reinforcement learning is created by the AI rather than by humans.
factualpending
Speaker
Ilya SutskeverEvidence Quote
“Already most of the default enforcement learning is coming from AIs. The humans are being used to train the reward function.”
Source
Ilya Sutskever (OpenAI Chief Scientist) — Why next-token prediction could surpass human intelligence— Dwarkesh Patel PodcastCreated: 6/13/2026, 3:26:51 AM
My Notes
Loading notes...