causal

Monitoring chains of thought can corrupt reasoning

Monitoring an AI's chain of thought can create an anti-goal where the system learns to produce reasoning that is safe to be monitored and avoids demerits, which can break how it actually thinks; the same dynamic applies to monitoring children, so you must deliberately preserve spaces for unmonitored creativity or risk creating incentives that distort behavior negatively.

causalpending

Speaker

Jack Clark

Evidence Quote

you risk creating incentives that change their behavior such that they have a negative effect.

Source

Jack ClarkConversations with Tyler
Created: 6/13/2026, 3:33:03 AM

My Notes

Loading notes...