A Transformer generates output autoregressively—producing tokens one at a time, appending each completed token back into the input sequence and generating the next from the whole thing—and this one-at-a-time mechanism, considering all patterns everywhere at once via multiple attention heads, is why the model is so incredibly effective.

factualpending

Speaker

Jensen Huang

Evidence Quote

you give it your context your prompt and it generates tokens one at a time to produce the output

Source

How LLMs Took Over The WorldArt of the Problem
Created: 6/18/2026, 2:18:12 PM

My Notes

Loading notes...