Mechanistic interpretability—the field trying to understand what goes on inside neural networks—has found that the internal structures are 'a mess' without clear connection to intuitive categories we expected, further confirming that human cognition is not built on intelligible symbolic structures.
factualpending
Speaker
Murray ShanahanEvidence Quote
“you find that you you know you train these things then there's a whole field of mechanistic interpretability that's trying to understand what goes on inside well you know that's a mess as well”
Created: 8/11/2026, 6:51:36 AM
My Notes
Loading notes...