1 claim in “cognitive science, psychology, AI safety”
What language models say to us (their linguistic output) is a tiny proportion of what is actually going on inside them; like human interrogation, lying detection, or psychoanalysis, we cannot assume language reveals the complete state or intentions of the system.