In the decoder, attention scores for a word like 'am' have values for itself and all other words before it, but zero attention scores for future words like 'fine', which tells the model to put no focus on those words.

factualpending

Speaker

Unidentified Speaker — Illustrated Guide to Transformers Neural Network: A step by… [4Bdc55j80l8]

Evidence Quote

the attention scores for M have values for itself and all other words before it but zero for the word fine

Source

Illustrated Guide to Transformers Neural Network: A step by step explanationThe AI Hacker
Created: 8/13/2026, 9:51:15 AM

My Notes

Loading notes...