definition
Mechanistic interpretability targets algorithms in weights
Mechanistic interpretability is distinguished from approaches like saliency maps by aiming to reverse-engineer the actual algorithms running in a model: treating the weights like a compiled binary and activations like the running program, the goal is to figure out how the weights correspond to algorithms.
definitionpending
Speaker
Chris OlahEvidence Quote
“we’d like to reverse engineer those weights and figure out what algorithms are running.”
Created: 6/13/2026, 3:36:36 AM
My Notes
Loading notes...