The softmax function is applied to scaled attention scores to produce attention weights with probability values between 0 and 1, where higher scores are heightened and lower scores are depressed, allowing the model to be more confident about which words to attend to.

factualpending

Speaker

Unidentified Speaker — Illustrated Guide to Transformers Neural Network: A step by… [4Bdc55j80l8]

Evidence Quote

next you take the softmax the scaled score to get the attention weights which gives you probability values between 0 & 1 by doing the softmax the higher scores get heightened and the lower scores are depressed

Source

Illustrated Guide to Transformers Neural Network: A step by step explanationThe AI Hacker
Created: 8/13/2026, 9:51:15 AM

My Notes

Loading notes...