factual
Dark matter of unobserved features in networks
Sparse autoencoders are like telescopes revealing features, but there is strong evidence we still see only a small fraction of them — a kind of 'dark matter' in the neural network universe that may never be observable or computationally tractable to observe, which has implications for safety if a significant fraction of networks remains inaccessible.
factualpending
Speaker
Chris OlahEvidence Quote
“what that means for safety if we can’t observe it, if some significant fraction of neural networks are not accessible to us.”
Created: 6/13/2026, 3:36:36 AM
My Notes
Loading notes...