Large language models are trained on enormous amounts of text data from the internet, Wikipedia, Reddit, digitized books, and computer code—ChatGPT was trained on approximately 500 billion words, whereas a typical human child encounters approximately 100 million words by age 10, making ChatGPT's training data about 5,000 times larger than a child's linguistic exposure.
factualpending
Speaker
Melanie MitchellEvidence Quote
“chat TPT was trained on the order of 50 500 billion words so just to put that in context a typical human child encounters either by hearing or reading roughly 100 million words by age 10 that's 5,000 times less than chat gbt”
Created: 8/12/2026, 6:19:21 PM
My Notes
Loading notes...