causal

Why bigger networks learn more: one-over-f noise analogy

Bigger networks are better because language (like physical processes) has a long-tail, one-over-f distribution of patterns at many scales; small networks capture only common simple patterns, while larger networks progressively capture rarer, more complex patterns up the hierarchy from grammar to sentences to paragraphs to themes.

causalpending

Speaker

Dario Amodei

Evidence Quote

if that long tail of other patterns is really smooth like it is with the one over F noise in physical processes like resistors, then you can imagine as you make the network larger, it’s kind of capturing more and more of that distribution.

Source

Dario AmodeiLex Fridman Podcast
Created: 6/13/2026, 3:36:36 AM

My Notes

Loading notes...