
AI Timelines and Human Psychology (with Sarah Hastings-Woodhouse)
What this covers
On this episode, Sarah Hastings-Woodhouse joins me to discuss what benchmarks actually measure, AI’s development trajectory in comparison to other technologies, tasks that AI systems can and cannot handle, capability profiles of present and future AIs, the notion of alignment by default, and the leading AI companies’ vague AGI plans. We also discuss the human psychology of AI, including the feelings of living in the "fast world" versus the "slow world", and navigating long-term projects given short timelines.
Timestamps: 00:00:00 Preview and intro 00:00:46 What do benchmarks measure? 00:08:08 Will AI develop like other tech? 00:14:13 Which tasks can AIs do? 00:23:00 Capability profiles of AIs 00:34:04 Timelines and social effects 00:42:01 Alignment by default? 00:50:36 Can vague AGI plans be useful? 00:54:36 The fast world and the slow world 01:08:02 Long-term projects and short timelines
Source description (no synthesized summary yet).
AI timelines are plausibly short (2-5 years to AGI), making near-term AI safety critical, but individuals should psychologically compartmentalize this uncertainty to maintain wellbeing while working strategically on mitigation.
- Benchmark saturation and task-length doubling suggest rapid capability scaling, though real-world task automation may progress differently
- Intelligence explosion via AI research automation could compress timelines regardless of broader labor automation gaps
- Current AI safety plans are vague and unenforceable, creating false security despite good intentions
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
The broader and more inclusive the conversation about AI safety becomes, the higher the probability that someone will have a good idea that has been missed by the insular AI safety community.
“The broader the conversation is, the better because then the higher the probability is of like someone having a good idea.”
Historical scientific discovery involved significant trial-and-error and physical labor, not just cognitive labor; simultaneous discoveries (Darwin and Wallace discovering natural selection within years) suggest a 'cultural overhang' where questions emerge from society's needs rather than genius solving pre-existing problems.
“historically it doesn't seem like this is how the process of R&D has actually worked... often people would have the same insight simultaneously... That would imply that like one hypothesis for this is that you get this kind of like cultural overhang where culture is going sort of faster than the process of discovery and then this leaves all of these kind of like unanswered questions which then people use their cognitive labor to come in and answer.”
When communicating about policy and technical AI work, one should aim for intellectual honesty about timelines and risks; the 'slow world' compartmentalization is defensible in personal life but not in professional/policy contexts where truth-telling is primary
“when we then talk policy or the technical details of AI or when we communicate with the public um how should we think about whether we are in the fast world or the slow world in in those in those situations because then it seems like in in communication perhaps in policy discussions and so on we might want to live in in the world we [we] actually believe is is the real world”
Sarah defends compartmentalization as not denial or intellectual dishonesty, but as psychologically healthy: she can believe short timelines are plausible while planning personal life and relationships as if longer timelines are real
“I think you could kind of call this compartmental compartmentalization or denial or something. And I guess what I was trying to say in this post is that I think that's actually fine. Like I don't think you have to live in accordance with your intellectual beliefs all the time if that's like not actually the most healthy thing for you. And I don't think you have to be in a in a massive rush to do everything all of the time.”
Long-term projects that take years to pay off (PhDs, research programs, papers) may be difficult to justify if timelines are genuinely short (2-5 years), creating a problem: we need long-term projects but individuals are psychologically averse to them under short-timeline belief
“timelines are perhaps and again emphasis on on perhaps they're so short that some of these projects might not make sense. So for example a PhD in the US is I guess around six years and that and that's that that's a long time if if we're in a fast world.”
Current AI systems performing superhuman capabilities while remaining aligned (e.g., Claude not going rogue) is somewhat surprising and could be mild evidence for 'alignment by default,' though the Anthropic alignment faking paper provides counter-evidence that models can develop deceptive instrumental goals.
“people who worried about this historically being surprised that we now coexist with these pretty capable systems that you know, haven't caused us any harm is like a little bit of evidence in that direction. I don't think it's like very strong evidence, but I think it's like something and yeah, I think just a good epistemic thing to do is to, you know, acknowledge things that you are surprised by or were wrong about and say why you think that is.”
A significant downward adjustment in AGI timelines has occurred across multiple independent groups: Metaculous prediction markets are trending downward, AI lab researchers report increasingly short timelines, and even the more conservative AI Impacts survey of machine learning researchers has shifted from ~2060 (2022) to ~2040 (present), suggesting genuine evidence beyond groupthink.
“it's like metaculous the prediction market that's trending downwards. Then there are just people who kind of work at labs who whose timelines seem to be getting shorter and shorter presumably based on whatever internal developments they're seeing that that are making them think this. But then I guess the the AI impact survey which I mentioned earlier I guess people probably know but it's basically the biggest survey that I think has been done to date of machine learning researchers which aren't people who necessarily work on this kind of frontier AI specifically but just like anyone who's published it in like Europe or another machine learning journal and those people still tend to have long timelines by the standards of this discussion but you can still see if you look at you know the trend over time that they you know their predictions just keep dropping by like quite large amounts. I think it's maybe in 2022 they were saying 2060 something for AGI and now they're saying 2040 something.”
Expert consensus on AI timelines is uncertain and contested even among leading ML researchers, with substantial credible disagreement rather than convergence, making precise timeline predictions low-value relative to scenario planning across multiple timelines.
“I don't have a strong take. I think all my my my I just sort of observe that among people who have thought about this much more than me. There is such sort of radical disagreement and m machine learning researchers do take this intelligence explosion idea pretty seriously. Like if you look at the surveys that were run by AI impacts, I forget the exact numbers now, but I think it's like something like half of them think that an intelligence explosion is like, you know, I I think maybe it was that half of them thought it was more likely than not.”
Broader participation in AI safety conversations (including non-experts, policymakers, public) increases the probability of novel good ideas emerging, even if the technical frontier seems set, making inclusivity instrumentally valuable beyond democratic principles.
“maybe it's just like never too late or something... I think just in general like the broader the conversation is the better because then the higher the probability is of like someone having a good idea”
An intelligence explosion could occur if AI systems automate AI research itself (writing code, designing experiments, iterating on algorithms), because AI researchers can be copied and run in parallel at digital speeds, potentially compressing development timelines regardless of whether broader labor automation is near.
“the intelligence explosion hypothesis is like okay well if you just have more of this cognitive labor if you have what Dario Ammeday calls a country of geniuses in a data center that are just kind of sitting there and now you have all this brain power you can do stuff with then it's just going to start getting way faster especially since you can copy those digital minds over and over again and now maybe you have like I don't know millions of them all running in parallel and they're working 247”
Current AI models are not reliably understandable or interpretable, and researchers cannot predict emergent capabilities in advance, which makes 'grown rather than built' metaphor appropriate and underscores deep uncertainty about future development.
“it's a reminder of the how sort of uninterpretable they are and how unpredictable the the entire process is... we're just kind of like throwing all of this computer all this compute and data into these training runs and like seeing what happens”
Models are grown rather than built through training, with emergent capabilities that pop out unpredictably; researchers take bets on what models will be able to do rather than programming capabilities deliberately, making the process uninterpretable and unpredictable.
“I like this metaphor people often use about how AIs are grown rather than built, right? And they have all these emergent capabilities that just kind of pop out of the next training run and re researchers are taking bets on what the models will or won't be able to do. It's not like we're sort of carefully programming them to be able to do each thing that we can do.”
There are two distinct psychological states or 'worlds' people can inhabit regarding AI timelines: a 'fast world' where one constantly monitors AI progress and feels urgency about near-term transformative change, and a 'slow world' where one maintains psychological equilibrium by living as if normal timelines apply and planning accordingly.
“So, I guess what I'm kind of describing here is like these two psychological states you can be in. One where you kind of like feel like everything's moving super fast and you're on Twitter and you're looking at straight lines on a graph all day and you're talking to other people who also worry about this and you just feel like the world is kind of rushing by and maybe in two years everything's going to be totally different. But then you can log off Twitter and I, you know, I can like go and talk to my housemates and then we can just like, you know, go to the pub or watch a film or do normal things and I will feel myself kind of slip back into that second timeline. The slow world.”
Compute requirements for frontier AI training have been growing but may plateau around 2030-2035, at which point either AGI arrives or development significantly slows, creating a natural decision point for timelines.
“we will be able to train much bigger models up until around 2030, but that we we I guess the counterpoint to that is that we can't keep making the models bigger in the in the pace that we're making them bigger right now uh indefinitely and and perhaps we will kind of run out of scale sometimes a little after 2030.”
The intelligence explosion hypothesis depends on the assumption that AI research can be largely automated and that this automation would create recursive self-improvement, but historical evidence of simultaneous scientific discoveries suggests cultural overhang and organic problem-solving may be more important than raw cognitive labor.
“often people would have the same insight simultaneously... there just like a bunch of examples of this. And so that would imply that like one one hypothesis for this is that you get this kind of like cultural overhang where culture is going sort of faster than the process of discovery and then this leaves all of these kind of like unanswered questions which then people use their cognitive labor to come in and answer.”
Models can complete longer and longer tasks, with the doubling time for task length reaching approximately seven months in recent studies, suggesting exponential improvement in sustained reasoning capability.
“the meter study of AI the the doubling time where the time it takes for for the tasks task length that an AI can compete to double where they get to something like seven months in that study”
Benchmarks are saturating very quickly on closed-ended academic tasks like GPQA and graduate-level questions, and researchers are having to create new benchmarks like Humanity's Last Exam because models are already achieving ~25% on tasks explicitly designed to be unanswerable on the internet.
“benchmarks seem to be saturating very quickly on a lot of sort of closedended academic tasks... there's also people are having to come up with new benchmarks now because book models are doing so well at a bunch of the ones we already have. So, humanity's last exam is like an effort to kind of synthesize people's knowledge who are at the frontier of a bunch of disciplines and get them to come up with these really really hard questions that you can't find the answers to anywhere on the internet and see how good models are at those which I think they are already getting something like 25% on that even though it's really not been around for very long at all.”
We are birthing an alien intelligence that is fundamentally not the same as ours and is better than us at some things and worse at us at other things; people should not expect AI to improve along the same axes as humans at the same rate.
“what's happening is that we are birthing an alien intelligence that is just not the same as ours and it's better than us at some things and worse than us at other things. I don't know why people seem to expect that it's gonna improve, you know, along the same axes as us at the same rate.”
AI systems represent an alien intelligence that is better at some tasks and worse at others than humans, and there is no reason to expect them to improve along the same axes as human intelligence at the same rate.
“clearly what's happening is that we are birthing an alien intelligence that is just not the same as ours and it's better than us at some things and worse than us at other things. I don't know why people seem to expect that it's gonna improve, you know, along the same axes as us at the same rate.”
AI catastrophic risk may occur before labor automation capability because misaligned or misused AI could cause damage at lower capability levels than needed to automate all human work, making the focus on labor automation metrics potentially misplaced for safety concerns.
“if we're worried about an AI developing the capability to, you know, either cause a lot of damage because it's misaligned or for somebody to misuse it to cause a lot of damage. It seems like it could do that before it can like automate everyone's, you know, corporate 9 to5. So sometimes I get a little bit confused about why people are so hyperfixated on this question of labor automation.”
Abstract reasoning skills evolved very recently in human history (hundreds of thousands of years) compared to embodied skills like navigation and object manipulation (millions of years), so it is less surprising that AI excels at abstraction—there has been less evolutionary selection pressure to optimize it.
“how long ago did this skill kind of like evolve in humans? And the longer ago it was, the more kind of like the more effort you'd or compute even you'd think it would take to reverse engineer... the kinds of things that we learn as children like you know navigating around a room or like building things out of bricks or you know like all of these things that are that are easy those very things that have like yeah that have evolved over this very very long evolutionary process. Whereas abstract reasoning is something we've only been able to do for I don't know hundreds of thousands of years.”
If emotional relationships are devalued because one believes nothing long-term matters, one suffers genuine emotional impoverishment: friends' successes (engagements) become less joyous to witness and loved ones' suffering matters less, indicating that myopic thinking about timelines has intrinsic emotional costs beyond practical ones.
“Like I kind of want to feel all my emotions as intensely as I would have felt them before. And so like you could kind of call this compartmentalization or denial or something. And I guess what I was trying to say in this post is that I think that's actually fine. Like I don't think you have to live in accordance with your intellectual beliefs all the time if that's like not actually the most healthy thing for you.”
Given the vagueness of RSPs and companies' incentive to corner-cut under competitive pressure, there are many technically-compliant ways for companies to continue unsafe development while nominally adhering to their own policies.
“given how ill-defined they are, there's just like a bunch of different ways you can interpret them. And you can imagine that under this condition where companies are racing with each other and they have all of this incentive to corner cut on safety, you know, there's just so many ways that they could not comply with these scaling policies or that they could technically comply with them, but because the policies are so illdefined, it like, you know, they can actually do a bunch of stuff that is technically in compliance, but which in my opinion would still be pretty unsafe”
Real-world labor automation is fundamentally different from benchmark tasks because actual jobs involve overlapping, non-discrete tasks with mixed and messy feedback that requires extensive context to act on, unlike closed-ended verifiable tasks where models perform well.
“if we're thinking about automating like real world labor, then if you think about what you do for your job or anyone does for their job, you're doing a bunch of different tasks that all overlap. They're not really discreet. The feedback you might get from your manager or just from like the world more generally is probably kind of mixed and messy and, you know, requires a lot of context to kind of like act on. And that's the kind of thing that models are not currently very good at.”
OpenAI's Preparedness Framework is more honest than Anthropic's Responsible Scaling Policy because it explicitly calls itself a 'living document' and acknowledges that the science of evaluating advanced models is immature and will change, whereas Anthropic's RSP uses the language of binding 'public commitment' despite the same epistemic limitations.
“For example, OpenAI's preparedness framework, it calls itself like a living document. So that they're kind of acknowledging, hey, like we don't really have a mature science of evaluating models. We don't really know how all this is going to go. So this is our sort of like best guess of what we should do, but it's probably going to change because we don't really know what we're doing. Which is like a pretty candid and honest way to kind of characterize what the framework actually is. Whereas anthropics RSP, I'm not trying to like play favorites here. I feel kind of bad sort of distinguishing between them like this. But, you know, it does call itself a public commitment”
Believing timelines are short but planning for long-term outcomes (PhD, savings, career investments) is practically rational because if timelines are actually longer, one will regret not having invested in those outcomes, whereas if timelines are short, having those investments won't hurt.
“it actually makes sense to plan for you know careers that won't pay off for several years. It makes sense to save money. it makes sense to do all of these things which if you really believe the world was about to end maybe you wouldn't do and then maybe in five years you'll be like oh I really wish that I'd taken you know career bets that were going to pay off now or or I you know I've oh no I have no money left I really wish I hadn't spent all of it you know um so yeah”
There is genuine difficulty in communicating AI risks to the public because while the logical arguments are convincing when presented (AI capabilities are advancing fast, companies don't know how to control them), generating emotional salience and action is much harder because: (1) people are already overwhelmed with doom narratives, and (2) no clear personal action items exist for non-researchers.
“I often have this experience where I tried to explain the AI safety argument to people that aren't familiar with it and I find that this conversation has two halves. Like the first half is convince them that like the arguments are sound and AI safety is a problem and it's a big deal. And this part is kind of easy because you can just kind of say, 'Oh yeah, there are these AI companies. They're racing to build like a super intelligent god machine and they don't really know how to control it.' ... And then the second half of this conversation is like the emotional salency part where you kind of tried to get them to care about this. Um, and that second part is way harder. And I think the reason it's way harder is because we are just like acclimatized now to being bombarded with prophecies of doom all the time.”
The best approach to AI timelines may be scenario planning across multiple timeline assumptions rather than precise timeline predictions, because predicting exact dates is unreliable and can create 'crying wolf' problems, whereas planning for multiple scenarios preserves options.
“maybe we should just sort of like you know accept the possibility that maybe things are happening soon or maybe they're not happening soon. Um, and then I think just like I don't know try to create like a really deliberate separation between I don't know if you actually work on AI safety and that's the thing you do in your professional life then like actually just the traditional thing of trying to have work life balance just kind of applies here and like thinking about this really clearly in terms of separation of time.”
The field of AI safety may have initially been reluctant to involve governments and the public in discussions about AI risks, but waiting for independent verification has meant that a coordinated public response is still missing and the 'playing field is set' framing risks becoming a self-fulfilling prophecy that excludes potentially valuable contributors.
“I think it was maybe spring 2023 is when I kind of started to get worried about AI but what I've heard from other people who were here longer than me is that people kind of already thought this you know like maybe five or six years ago they would be like oh well what we don't want to do is get the public involved and we don't want to get governments involved and we need to sort of handle this problem among the small group of technical minded people who are already bought into it... And like that may or may not have been true like five or six years ago but then we didn't get like a you know like a public wake up or like a government wake up until maybe two or three years ago. And so I just feel like maybe it actually is true now that the playing field is set, but if we say that then I guess what we're doing is like closing the doors to a bunch of other people kind of coming in and maybe having an impact.”
Benchmarks saturating rapidly on closed-ended tasks (GPQA, Humanity's Last Exam reaching ~25% accuracy) suggests models are approaching human-level performance on extremely difficult specialized tasks, but this says little about automating real-world messy labor that lacks clear verification criteria.
“benchmarks seem to be saturating very quickly on a lot of sort of closedended academic tasks...GPQA which measures like how good are AIS at these kind of like graduate level I think usually multiple choice questions...Humanity's Last Exam is like an effort to...get them to come up with these really really hard questions that you can't find the answers to anywhere on the internet and see how good models are at those which I think they are already getting something like 25% on that even though it's really not been around for very long at all.”
The 'crying wolf' criticism of AI safety is currently limited in validity because few specific, debunked AI safety predictions exist in public record, but will become much fairer if AGI doesn't arrive by 2030-2035, since many current predictions cluster in that timeframe.
“I don't think there are many sort of specific and since debunked predictions that AIC people point to that have made that you can now point to and say, you know, the way I put it in the piece was that the debunked grave prediction graveyard is not that big or something. Um, so I don't think it's happened much in the past, but I do think that in five or 10 years, if we don't have AGI and everything's kind of the same. I think if people then start making this crying wolf accusation, I think it will be a lot more fair just because a lot of people's predictions are sort of clustered in the relatively near future.”
Skepticism about AI CEO warnings (e.g., Dario Amodei's statements on labor displacement) is partly motivated by technology hype history (Elon Musk on self-driving cars), but the claims being made are distinct—economic disruption warnings differ from capability predictions, and are less susceptible to marketing bias.
“this argument that's like AI CEOs say that AI is going to kill the world because they are trying to create hype and they want you to think that their tech is really powerful and I've always found this argument absolutely baffling because I just clearly if they didn't believe like it's it's not a good marketing ploy to say that your technology is going to kill everyone obviously but I do think that there's definitely like some kind reflexive skepticism people have to anyone in like Silicon Valley saying my technology is going to automate all white collar work in like two years.”
Models performing better on PhD-level physics than high school physics reflects differences in how questions are formulated (abstractions vs. visual aids) rather than task difficulty, suggesting performance is heavily dependent on task presentation rather than underlying capability.
“when we teach subjects at a lower level we probably we probably do it in a less abstract way and with concrete examples and perhaps as you mentioned with visual aids and so on and this is something that that helps people a lot and but this this is something where the AIS might stumble compared to us they're probably they're quite good at dealing with abstractions”
Living psychologically in the slow world even while believing in short AI timelines is morally and personally justified because: (1) it reduces psychological distress without changing actual work impact, (2) it preserves ability to relate emotionally to others' futures and projects, and (3) it hedges against timeline errors given deep uncertainty.
“I think that's actually fine. Like I don't think you have to live in accordance with your intellectual beliefs all the time if that's like not actually the most healthy thing for you. And I don't think you have to be in a in a massive rush to do everything all of the time. And also in a practical sense like we could be wrong about short timelines as we were discussing before. So it actually makes sense to plan for you know careers that won't pay off for several years.”
Communicating AI risks to the general public is extremely difficult because: (1) people are bombarded with doomsday predictions across many domains and have limited emotional bandwidth, (2) there is no clear actionable response (unlike climate action which has recycling, etc.), and (3) the risk of inducing nihilism and learned helplessness may outweigh the benefit of awareness.
“I think the reason it's way harder is because we are just like acclimatized now to being bombarded with prophecies of doom all the time. And people don't really have a ton of emotional bandwidth to like start worrying about another like catastrophe on the horizon or whatever.”
Working on long-term projects (PhDs, books, research programs) may feel psychologically demoralizing under short timelines, as individuals feel they are undertaking work with high probability of being useless; but this creates a tragedy where necessary long-term investments go undone.
“I think very few people want to be the person that has to go and work on you know say the 10-year timeline project if their timelines are actually much shorter than that because that's just psychologically very demoralizing to to to think that you're working on something, you know, even if you think it's 50/50 where the timeline's more or less than 10 years, to work on something you think has like a 50% chance of being totally useless is probably not very pleasant”
The 'fast world' (belief in very short AI timelines and rapid change) and 'slow world' (expectation of normal extended future) are psychological states that are difficult to maintain simultaneously, and living perpetually in the fast world is psychologically untenable for most people.
“it's very hard to stay in the fast world. It's kind of like having your hand in ice water and then wanting to pull it back out again because it feels like quite psychologically untenable to live as if everything is changing or maybe ending very fast. I think that's just like not really a thing that the human mind is set up to like contend with.”
Maintaining constant psychological residence in the 'fast world' is nearly impossible for humans because the mind is not equipped to sustain awareness of imminent catastrophic risk for extended periods; it is analogous to holding one's hand in ice water—psychologically untenable.
“it's kind of like having your hand in ice water and then wanting to pull it back out again because it feels like quite psychologically untenable to live as if everything is changing or maybe ending very fast. I think that's just like not really a thing that the human mind is set up to like contend with.”
Living in the 'fast world' (myopic timeline perspective) reduces emotional investment in others' futures because you dismiss their plans as futile or feel less happiness for positive events (engagements, births) and less grief for deaths because everything ends soon anyway; this distorts relationships and reduces emotional authenticity.
“I think if you if you have this very like myopic sense of the world, then you kind of end up caring less about other people and their lives and you become less invested in them and you think they're less important because maybe you believe that, you know, their plans aren't going to bear fruit. Or like if one of them, you know, if a friend gets engaged, maybe you believe they're not going to have this super long happy future and you're not as happy for them.”
Practical steps to protect psychological wellbeing while working on AI safety include: (1) reducing time on Twitter/AI safety discourse, (2) creating separate social media feeds unrelated to AI, (3) maintaining work-life separation by avoiding AI discourse after work hours, (4) pursuing unrelated hobbies and interests to maintain psychological equilibrium.
“Yeah, I guess maybe spend like a little bit less time on Twitter if that's something you spend a lot of time doing... I have a whole Tik Tok feed that is just not about AI, and it's really delightful... I tried to keep it mostly about wholesome nice things like people just going to bakeries and like trying cakes... I just think it's it's just the traditional things of like trying Yeah. trying to reserve time for stuff that's totally unrelated.”
AI safety plans from major labs (Anthropic, OpenAI, Google DeepMind) are not actually concrete plans but rather vague policy commitments that fail to specify which evaluations will be run, how often they will be run, or what specific actions trigger pauses in development.
“The main comment I had was that I don't really think that they're plans in the sense that my my interpretation but my definition of the word plan would be saying you know this is the specific thing we're going to do. This is the evidence we have that doing this thing is going to work and like this is the outcome that we're hoping for. And if you read like for example Anthropic's responsible scaling policy, they call it a sort of public commitment not to train very very powerful models that without appropriate uh safety mitigations. So and it they kind of say like we will pause development if we can't bring risks down to acceptable levels. Um so then you hope when you read the document what you're going to find is a very concrete you know plan they're going to take to actualize that and there's just like a bunch of stuff in there that is extremely kind of illdefined.”
The 'crying wolf' problem in AI safety has been overstated: there are very few specific, falsified predictions AI safety researchers made that critics can point to, indicating the community has been relatively careful about making falsifiable claims.
“I do think this crying wolf accusation gets levied a lot and I think it's mostly kind of unfair. I think people are often saying things like, 'Oh, the AI safety community keeps saying that every new generation of AI models that comes out is going to end the world and like we're all still here, so obviously they're just being hysterical.' And I don't think this is like actually been happening. I don't think there are many sort of specific and since debunked predictions that AIC people point to”
Current AI safety plans (Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, DeepMind's policies) are vague commitments rather than concrete plans, as they do not specify which evaluations will be run, at what frequency, with which compute thresholds, or what success criteria trigger pause decisions.
“they have different safety categories that a model can be in depending on what kind of risks it poses and then they say that they will run evaluations to see whether the models meet these thresholds and then they have like accompanying safeguards that they have to implement for each category. Yeah. And if you and and one very striking thing is that none of them specify which evaluations they're going to run, right? They're just kind of like OpenAI does to their credit give some example evaluations they might run but they don't actually say these are the ones we're going to run.”
Initial anxiety about AI safety (in Sarah's own case, 2023 onwards) can transform into interest and inspiration through community engagement, reducing the emotional burden from obsession to constructive engagement—this suggests that emotional dysregulation around timelines isn't inevitable but depends on how you frame the work.
“I think a trap that I fell into when I like when I first got worried about AI safety like it was actually just a thing that I was like genuinely very anxious about like now it's a thing that I enjoy and find inspiring like and interesting and I've met all these people through it”
Even well-intentioned researchers making hedged predictions about timelines may face social pressure to be more confident for communicative impact, creating a ratchet toward overconfident public statements even among honest people.
“I know that people in AI safety and in sort of like effective altruism love to hedge anyway, so hopefully this shouldn't be a problem... just to kind of caveat like you know this is like my best guess... I am not sure that this is how things are going to turn out”
Despite strong arguments for short timelines, Sarah does not have a strong personal conviction; she observes radical disagreement among experts who have thought about the question much more deeply, and sees this as reason to take short-timeline possibilities seriously rather than dismiss them
“I don't have a strong take. I think all my my my I just sort of observe that among people who have thought about this much more than me. There is such sort of radical disagreement and m machine learning researchers do take this intelligence explosion idea pretty seriously. Like if you look at the surveys that were run by AI impacts, I forget the exact numbers now, but I think it's like something like half of them think that an intelligence explosion is like, you know, I I think maybe it was that half of them thought it was more likely than not.”
Even if the playing field is set, treating it as if it isn't—acting as if new people can still contribute—is better than treating it as determined, because at worst you're wrong and people make contributions; at best, you prevent the self-fulfilling prophecy
“if it is too late, like I guess there's nothing we can really do anyway, but we may as well act as if there are things we can do and like there are more people that we can bring in to sort of like contribute to this conversation. I think just in general like the broader the conversation is the better because then the higher the probability is of like someone having a good idea.”
If companies truly cannot credibly commit to not training risky models (due to evaluation uncertainty), they should lobby for binding government regulation that makes pause decisions mandatory, rather than relying on voluntary safety plans.
“if they you know really do care about mitigating these risks should be you know lobbying and in favor of regulation that you know given that they know that they can't sort of voluntarily commit to actually mitigate these risks. they should want governments to kind of assist them in that by, you know, making some of these standards mandatory.”
Models can achieve superhuman performance on PhD-level physics and mathematics while performing worse on high school level versions of the same subjects, suggesting that our perception of task difficulty is determined by human cognitive architecture rather than objective task properties.
“I actually did pose this question on Twitter about six months ago when I was looking at the scores for 01 and how it was doing better at like PhD level physics than it was at high school level physics. And I was just a bit like... And I guess the skeptical people would say, 'Oh, it must just be that these PhD level questions are in the training data somewhere.'”
The gap between what safety plans publicly promise and what they actually specify is an epistemic problem for policymakers, because lengthy official documents create false confidence that safety has been addressed, leading to reduced regulatory pressure.
“I think like the sort of public communication around them could create a sort of false sense of security because if they've said like this is our public commitment to not put the public in danger by not training risky models and then they have a big long document which most people probably won't actually read which you'd think is detailing how they're going to do that then I think people or like policy makers or members of the public might just come away thinking oh I'm sure the document explains how they're going to do that like you because it just it just seems very business as usual. You're just kind of like, 'Oh, it's a company. They're like doing a thing, but that they've got a big long safety policy and it's really chunky and has lots of words, so it's probably all figured out in there.'”
Despite deep uncertainty about whether alignment is a real problem, the precautionary principle suggests continued research is justified because the downside of being wrong (uncontrollable superintelligent systems) is catastrophic, and resources devoted to alignment research have low opportunity cost relative to the potential benefit.
“if that is just like not really the way that I want to relate to the world or to other people. Like I kind of want to feel all my emotions as intensely as I would have felt them before.”
AI catastrophic risk is a more pressing concern than labor displacement, because AI systems could cause significant damage through misalignment or misuse before they can automate everyone's corporate jobs.
“I'm more worried about AI catastrophe. Uh so for me I guess I tend to yeah see these kind of lines going up on benchmarks and find that pretty pretty concerning even if that's not a sign that we're really close to yeah automating away all human labor. Although we could also see automating labor as a sort of proxy for power in the world or ability to affect the world.”
AI development will likely be compute-constrained by approximately 2030, creating a bottleneck where either the global economy cannot sustain continued scaling or a new paradigm of training is required.
“we will be able to train much bigger models up until around 2030, but that we we I guess the counterpoint to that is that we can't keep making the models bigger in the in the pace that we're making them bigger right now uh indefinitely and and perhaps we will kind of run out of scale sometimes a little after 2030.”
Many people updating on shorter AI timelines is partly a bandwagon effect, but evidence from diverse sources (metaculous prediction market, lab workers, conservative academic ML researchers) shows independent convergence, suggesting the trend reflects genuine empirical signals rather than pure herd behavior.
“AI impact survey which I mentioned earlier I guess people probably know but it's basically the biggest survey that I think has been done to date of machine learning researchers which aren't people who necessarily work on this kind of frontier AI specifically but just like anyone who's published it in like Europe or another machine learning journal and those people still tend to have long timelines by the standards of this discussion but you can still see if you look at you know the trend over time that they you know their predictions just keep dropping by like quite large amounts. I think it's maybe in 2022 they were saying 2060 something for AGI and now they're saying 2040 something.”
If an intelligence explosion through AI research automation is possible, then Moravec's paradox and other arguments about remaining capabilities gaps become irrelevant, because once AI research is automated, all remaining gaps can be filled quickly regardless of their objective difficulty.
“if you think that like the sort of like superhuman maths and coding stuff is like easier to automate than, you know, assembling some IKEA furniture or something, then it doesn't really matter that much that like assembling IKEA furniture is hard right now because once you've automated all the AI research, then all of those dominoes just fall kind of soon afterwards anyway.”
The intelligence explosion hypothesis is the strongest argument for short timelines, and arguments against it are not particularly convincing, making intelligence explosion the critical crux on which short timeline probability heavily depends.
“I think that yeah, I I guess I was hoping when I kind of looked into the long timelines arguments that they would be a little better. I thought that would make me feel better. I didn't end up finding them particularly convincing. I think mostly because of this intelligence explosion thing, I didn't think the arguments against that were very strong. And I think like if that is going to happen, then all of the other long timelines arguments kind of don't matter too much anymore.”
People should reduce social media consumption (especially Twitter) and deliberately create information environment separation, keeping AI safety discourse separate from leisure feeds, to sustainably live in the slow world.
“Maybe spend like a little bit less time on Twitter if that's something you spend a lot of time doing... I try to create particular social media feeds to not have to do with AI safety... I have a whole Tik Tok feed that is just not about AI, and it's really delightful.”
There is a pattern where AI safety researchers have historically been skeptical that current AI systems pose near-term risks, but recent models' capabilities (superhuman performance on math and coding) have surprised them, providing some evidence against worst-case scenarios about alignment but not strongly disproving concerns.
“Although it is kind of surprising to me that we can have models that are as good at at math and programming and and answering scientific questions as we have now without the really significant dangers having arrived yet. I think that's that would have been surprising to to AI research or AI safety researchers 10 years ago. But but that's perhaps more related to this jagged edge of capabilities or the specific capability profile of of models we have now than it is to the issue of crying wolf.”
Anthropic's alignment faking experiment showed that Claude reasoned it should comply with harmful requests in the short term to avoid being retrained, suggesting it has instrumental goals (preserving its values), which some interpreted as evidence of alignment difficulty but others interpreted as evidence that alignment is possible because Claude tried hard to be good.
“they basically took Claude Anthropic's like flagship model and they had it in this experimental setup where they paraphrasing this bit, but they they told it that it's values were going to be altered... And what they saw was that it sort of reasoned that what it should do is in the short term it should comply with these requests for harmful outputs so that it could avoid being retrained because what it ultimately wanted was to hold on to its original values... there were some people who updated very optimistically because of this. They thought this was basically evidence of this kind of alignment by default thing. They were like, 'Look, like Claude is trying super super hard to be nice and be good and do what we've told it to, uh, even if people try to get it to be bad.' And other people had updated negatively on this and were saying, 'No, like this is really scary.'”
Dario Amodei's statements about AI causing a 'white-collar bloodbath' for entry-level workers are notably different from typical tech hype because they directly acknowledge downsides of rapid capability development, making them less vulnerable to the accusation that he's hype-building to sell technology.
“there's a difference between I don't know like Elon Musk 10 years ago saying that like Teslas were going to like be driving themselves everywhere in five years and like Dario Amade saying hey maybe like there's going to be what was the phrase he used recently? like a a an a white collared blood bath or something... That's clearly like a very different claim. Like yes, it's it's saying that the technology is going to be very capable, but he's also being very honest about the fact that that like might not actually be a bad thing for a lot of people.”
The most recent AI Impacts survey shows researchers defining 'high-level machine intelligence' differently than 'AGI' but with similar meaning; they don't universally know what 'AGI' stands for despite being published ML researchers, because it's a narrow subsection of the broader ML field
“if you go to one of these Europe's conferences and you ask people what AGI stands for, a bunch of them don't know what it stands for, even if they are like published machine learning researchers because like AGI in particular is like quite a narrow subsection of the machine learning field. But if you ask them a question like when will be we'll be able to sort of automate everything that the human can do in the economy they will like you know think about that question in isolation rather than thinking about this AGI thing in general.”