Current AI safety plans (Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, DeepMind's policies) are vague commitments rather than concrete plans, as they do not specify which evaluations will be run, at what frequency, with which compute thresholds, or what success criteria trigger pause decisions.
factualpending
Speaker
Sarah Hastings WoodhouseEvidence Quote
“none of them specify which evaluations they're going to run...don't actually say these are the ones we're going to run”
Created: 8/11/2026, 7:16:09 AM
My Notes
Loading notes...