The alignment problem may be partially solved 'by default' through properties of large language models that naturally develop common sense understanding of goals from next-token prediction training, making them less prone to the Malign Omniscience failure modes that earlier alignment researchers feared.

factualpending

Speaker

Scott Alexander

Evidence Quote

I think one of the reasons I'm more hopeful than I used to be is that LLMs are great compared to the kind of reinforcement learning self-play agents that they expected.

Source

AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel KokotajloDwarkesh Patel
Created: 8/11/2026, 7:03:29 AM

My Notes

Loading notes...