factual

RLHF data is mostly generated by AI

In reinforcement learning from human feedback, humans are used only to train the reward function; once trained, the reward function's interaction with the model is automatic, so most of the data generated during reinforcement learning is created by the AI rather than by humans.

factualpending

Speaker

Ilya Sutskever

Evidence Quote

Already most of the default enforcement learning is coming from AIs. The humans are being used to train the reward function.

Source

Ilya Sutskever (OpenAI Chief Scientist) — Why next-token prediction could surpass human intelligenceDwarkesh Patel Podcast
Created: 6/13/2026, 3:26:51 AM

My Notes

Loading notes...