Modern large language models operate on tokens (letters, words, or punctuation) rather than just an alphabet, predicting the odds of the next token given a string of prior tokens—but unlike simple Markov chains they use 'attention' to decide which prior tokens matter, allowing context like 'blood' and 'mitochondria' to disambiguate that 'cell' means biology rather than a prison.

definitionpending

Speaker

Derek Muller

Evidence Quote

unlike simple Markov chains, they also use something called attention, which tells the model what to pay attention to

Source

The Strange Math That Predicts (Almost) AnythingVeritasium
Created: 6/18/2026, 1:59:31 PM

My Notes

Loading notes...