causal

Automated interpretability trust concern

Using neural networks to audit neural networks raises a 'reflections on trusting trust' problem: if you rely on powerful AI to verify that your AI systems are safe, you must worry whether the auditing model could be screwing with you — a concern Olah considers minor now but potentially serious in the long run.

causalpending

Speaker

Chris Olah

Evidence Quote

I do wonder in the long run, if we have to use really powerful AI systems to go and audit our AI systems, is that actually something we can trust?

Source

Dario AmodeiLex Fridman Podcast
Created: 6/13/2026, 3:36:36 AM

My Notes

Loading notes...