definition

Mechanistic interpretability targets algorithms in weights

Mechanistic interpretability is distinguished from approaches like saliency maps by aiming to reverse-engineer the actual algorithms running in a model: treating the weights like a compiled binary and activations like the running program, the goal is to figure out how the weights correspond to algorithms.

definitionpending

Speaker

Chris Olah

Evidence Quote

we’d like to reverse engineer those weights and figure out what algorithms are running.

Source

Dario AmodeiLex Fridman Podcast
Created: 6/13/2026, 3:36:36 AM

My Notes

Loading notes...