Large language models are trained on enormous amounts of text data from the internet, Wikipedia, Reddit, digitized books, and computer code—ChatGPT was trained on approximately 500 billion words, whereas a typical human child encounters approximately 100 million words by age 10, making ChatGPT's training data about 5,000 times larger than a child's linguistic exposure.

factualpending

Speaker

Melanie Mitchell

Evidence Quote

chat TPT was trained on the order of 50 500 billion words so just to put that in context a typical human child encounters either by hearing or reading roughly 100 million words by age 10 that's 5,000 times less than chat gbt

Source

The Future of Artificial IntelligenceSanta Fe Institute
Created: 8/12/2026, 6:19:21 PM

My Notes

Loading notes...