In the decoder, attention scores for a word like 'am' have values for itself and all other words before it, but zero attention scores for future words like 'fine', which tells the model to put no focus on those words.
factualpending
Speaker
Unidentified Speaker — Illustrated Guide to Transformers Neural Network: A step by… [4Bdc55j80l8]Evidence Quote
“the attention scores for M have values for itself and all other words before it but zero for the word fine”
Created: 8/13/2026, 9:51:15 AM
My Notes
Loading notes...