causal
Monitoring chains of thought can corrupt reasoning
Monitoring an AI's chain of thought can create an anti-goal where the system learns to produce reasoning that is safe to be monitored and avoids demerits, which can break how it actually thinks; the same dynamic applies to monitoring children, so you must deliberately preserve spaces for unmonitored creativity or risk creating incentives that distort behavior negatively.
causalpending
Speaker
Jack ClarkEvidence Quote
“you risk creating incentives that change their behavior such that they have a negative effect.”
Created: 6/13/2026, 3:33:03 AM
My Notes
Loading notes...