The attempt to guard-rail ChatGPT by having it refuse to utter racial slurs has failed to generalize, and broader attempts at value alignment through reinforcement learning have been very sloppy and don't reliably do what is intended.
factualpending
Speaker
Stuart RussellEvidence Quote
“we've not been very successful”
Source
The Trouble with AI: A Conversation with Stuart Russell and Gary Marcus (Episode #312)— Sam HarrisCreated: 8/11/2026, 1:51:56 AM
My Notes
Loading notes...