Distillation is much more effective when the teacher provides probabilities for all output classes rather than just the correct answer. When a teacher gives probabilities for 1024 output categories (which sum to 1), training a student to match those probabilities provides far more information per training example than just telling the student the right class label (which only provides 10 bits of information).
factualpending
Speaker
Geoffrey HintonEvidence Quote
“if you just tell the agent the right answer you're only giving them 10 bits of information... the teacher will give probabilities for all the outputs and that's 1023 real numbers... each training example is far more valuable”
Created: 8/12/2026, 5:56:09 PM
My Notes
Loading notes...