The attempt to guard-rail ChatGPT by having it refuse to utter racial slurs has failed to generalize, and broader attempts at value alignment through reinforcement learning have been very sloppy and don't reliably do what is intended.

factualpending

Speaker

Stuart Russell

Evidence Quote

we've not been very successful

Source

The Trouble with AI: A Conversation with Stuart Russell and Gary Marcus (Episode #312)Sam Harris
Created: 8/11/2026, 1:51:56 AM

My Notes

Loading notes...