causal
Automated interpretability trust concern
Using neural networks to audit neural networks raises a 'reflections on trusting trust' problem: if you rely on powerful AI to verify that your AI systems are safe, you must worry whether the auditing model could be screwing with you — a concern Olah considers minor now but potentially serious in the long run.
causalpending
Speaker
Chris OlahEvidence Quote
“I do wonder in the long run, if we have to use really powerful AI systems to go and audit our AI systems, is that actually something we can trust?”
Created: 6/13/2026, 3:36:36 AM
My Notes
Loading notes...