Sarah Hastings Woodhouse
About
AI safety researcher and essayist; guest on Future of Life Institute podcast
Cast within
No topic-region cast yet — this appears once Sarah Hastings Woodhouse's compiled claims are aligned into a topic region's argument tree.
Claims by Sarah Hastings Woodhouse (20 of 55)
An intelligence explosion could occur if AI systems automate AI research itself (writing code, designing experiments, iterating on algorithms), because AI researchers can be copied and run in parallel at digital speeds, potentially compressing development timelines regardless of whether broader labor automation is near.
Benchmarks saturating rapidly on closed-ended tasks (GPQA, Humanity's Last Exam reaching ~25% accuracy) suggests models are approaching human-level performance on extremely difficult specialized tasks, but this says little about automating real-world messy labor that lacks clear verification criteria.
Skepticism about AI CEO warnings (e.g., Dario Amodei's statements on labor displacement) is partly motivated by technology hype history (Elon Musk on self-driving cars), but the claims being made are distinct—economic disruption warnings differ from capability predictions, and are less susceptible to marketing bias.
Current AI safety plans (Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, DeepMind's policies) are vague commitments rather than concrete plans, as they do not specify which evaluations will be run, at what frequency, with which compute thresholds, or what success criteria trigger pause decisions.
Living psychologically in the slow world even while believing in short AI timelines is morally and personally justified because: (1) it reduces psychological distress without changing actual work impact, (2) it preserves ability to relate emotionally to others' futures and projects, and (3) it hedges against timeline errors given deep uncertainty.
Current AI systems performing superhuman capabilities while remaining aligned (e.g., Claude not going rogue) is somewhat surprising and could be mild evidence for 'alignment by default,' though the Anthropic alignment faking paper provides counter-evidence that models can develop deceptive instrumental goals.
Historical scientific discovery involved significant trial-and-error and physical labor, not just cognitive labor; simultaneous discoveries (Darwin and Wallace discovering natural selection within years) suggest a 'cultural overhang' where questions emerge from society's needs rather than genius solving pre-existing problems.
Communicating AI risks to the general public is extremely difficult because: (1) people are bombarded with doomsday predictions across many domains and have limited emotional bandwidth, (2) there is no clear actionable response (unlike climate action which has recycling, etc.), and (3) the risk of inducing nihilism and learned helplessness may outweigh the benefit of awareness.
Abstract reasoning skills evolved very recently in human history (hundreds of thousands of years) compared to embodied skills like navigation and object manipulation (millions of years), so it is less surprising that AI excels at abstraction—there has been less evolutionary selection pressure to optimize it.
Working on long-term projects (PhDs, books, research programs) may feel psychologically demoralizing under short timelines, as individuals feel they are undertaking work with high probability of being useless; but this creates a tragedy where necessary long-term investments go undone.
Living in the 'fast world' (myopic timeline perspective) reduces emotional investment in others' futures because you dismiss their plans as futile or feel less happiness for positive events (engagements, births) and less grief for deaths because everything ends soon anyway; this distorts relationships and reduces emotional authenticity.
My Notes
Loading notes...