causal

RLHF works via subtle aggregated human preferences

RLHF works so well because human preference data contains a huge amount of subtle information — different people pick up on small things (like correct semicolon usage) that an observer wouldn't even notice — and the model learns across all domains and contexts what humans want, similar to how deep learning beats hand-coded edge detection.

causalpending

Speaker

Amanda Askell

Evidence Quote

each of these single data points, and this model just has so many of those, it has to try and figure out what is it that humans want in this really complex across all domains.

Source

Dario AmodeiLex Fridman Podcast
Created: 6/13/2026, 3:36:36 AM

My Notes

Loading notes...