A Transformer generates output autoregressively—producing tokens one at a time, appending each completed token back into the input sequence and generating the next from the whole thing—and this one-at-a-time mechanism, considering all patterns everywhere at once via multiple attention heads, is why the model is so incredibly effective.
factualpending
Speaker
Jensen HuangEvidence Quote
“you give it your context your prompt and it generates tokens one at a time to produce the output”
Created: 6/18/2026, 2:18:12 PM
My Notes
Loading notes...