Modern transformer-based LLMs work in essentially the same way as Hinton's 1985 tiny language model—turning words into feature activations, having features interact to predict the features of the next word, and backpropagating prediction error—just with many more input words, many more layers, and far more complicated feature interactions (including query-key matching and disambiguation across layers).
factualpending
Speaker
Geoffrey HintonEvidence Quote
“the large language models which I like to think are descendants of my tiny language model”
Source
Will AI outsmart human intelligence? - with 'Godfather of AI' Geoffrey Hinton— The Royal InstitutionCreated: 6/18/2026, 2:17:54 PM
My Notes
Loading notes...