Holden Karnofsky — History's most important century
What this covers
Holden Karnofsky makes a case for why this century could be the most consequential in history, anchored in economic growth theory and the probability of transformative AI. The conversation with Dwarkesh Patel works through the argument from multiple angles: if AI systems capable of automating the "ideas" step of the people-ideas-resources feedback loop emerge this century—which Karnofsky estimates as more than 50% likely—it could restore the accelerating growth that stalled when human population growth decoupled from economic growth. That explosive expansion could culminate in a stable, post-human future. Because such a transition would be nearly irreversible, this becomes our last meaningful window to shape how it unfolds, if it unfolds at all.
The discussion spans how evidence for near-term transformative AI comes from multiple independent lines—economic history, biological scaling arguments, expert surveys, and deep learning progress—while also establishing that even ignoring AI entirely, our era stands out as extraordinarily early and dynamic by cosmic standards. Karnofsky addresses concrete concerns: lock-in risks under malign governance, the full automation bottleneck objection, and which tasks resist automation longest (physical dexterity and trust-based roles like teaching). On ethics, he develops future-proof ethics as a framework for acting under moral uncertainty, discusses Bostrom's moral parliament, and argues against ends-justify-the-means reasoning even for enormous stakes. He also defends the neglectedness heuristic—why working on emerging, under-recognized problems like AI alignment before institutional infrastructure solidifies can yield outsized long-run influence.
Karnofsky argues that if AI capable of automating all human contributions to science and technology is developed this century (which he considers more than 50% likely), it could trigger explosive growth leading to a stable, post-human future, making this potentially the most important century in history and our last chance to shape that trajectory.
- Economic growth theory predicts that AI filling the 'ideas' step of the people-ideas-resources feedback loop would restore accelerating, explosive growth that broke when people stopped having more children with more resources
- Multiple independent lines of evidence (economic history, biological anchors, expert surveys, deep learning progress) suggest transformative AI is plausibly near
- Even ignoring AI, commonly accepted facts show we live in an extraordinarily weird and early time, so the AI claim is only a moderate quantitative update rather than a complete revolution in thinking
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
When people believe they live in an especially important time, the best norm is neither to dismiss the thought (so genuinely important actors do the wrong thing) nor to take it so seriously they promote their own interests above everyone's; rather, people should take such beliefs reasonably seriously and act on them while adhering to common-sense ethical standards and avoiding 'ends justify the means' reasoning like lying or breaking the law.
“all these people should take their beliefs reasonably seriously and try to do the best thing according to their beliefs, but should also adhere to common sense standards of ethical conduct and not do too much "ends justify the means' reasoning."”
Moral progress is real—the term refers to changes in morality that are genuinely good (e.g., increased acceptance of homosexuality)—but it is not objective truth, not inevitable, and does not happen automatically just because time passes or intelligence increases; much of historical moral progress came from humans getting to know other humans they previously stereotyped, a mechanism a non-human AI would not share.
“I've used the term moral progress as just a term to refer to changes in morality that are good. I think there has been moral progress, but I don't think that means moral progress is something inevitable or something that happens every time you are intelligent.”
Faced with moral uncertainty (e.g., total-good views wanting a big world vs suffering-focused views wanting a small one vs integrity-focused views), Bostrom's moral parliament approach treats your competing moral views as multiple people living in your head who care about each other and negotiate a deal—so you favor actions that are really good according to one view and not too bad according to the rest, which moderates against extreme ends-justify-means behavior.
“The moral parliament idea is an idea that was laid out by Nick Bostrom in an Overcoming Bias post a decade ago that I like. I think about it as if I'm just multiple people.”
Even ignoring AI entirely, commonly accepted facts—the last couple hundred years showing far faster economic and technological growth than all prior history, our position in a tiny sliver of cosmic time, and the apparent absence of other galactic life—make our era one of the most extraordinary and earliest times ever, so claiming transformative AI is coming is only a moderate quantitative update rather than a wild leap.
“even if you ignore AI, there's a lot of things that are very strange about the time that our generation lives in”
If you want to do something new, world-changing, and revolutionary, the worst approach is to enter a well-established scientific field and try to revolutionize it; it's better to use common sense to ask an important question that no one works on because it isn't yet a recognized field with institutions—like AI alignment, which lacks academic departments—since that's where outsized impact is found.
“if people want to do something that's really new and world changing and dramatic and revolutionary, the worst way to do that is to go into some well-established scientific field and try to revolutionize that.”
Future-proof ethics is the project of adopting a moral system that would survive substantial moral progress—one where if you later became wiser and reflected more, you wouldn't look back on your current actions as horrible monstrous mistakes (as much of history does in hindsight)—built on principles like systemization, thin utilitarianism, and sentientism.
“a lot of people I know are trying to come up with a system of morality and a system of ethics that would survive a lot of moral progress.”
A CEO should know enough about a topic to manage its specialists effectively, but how much varies by how easily outputs can be judged: finance and audits can be judged by outputs (was it compliant, did we pass) without deep knowledge, whereas central functions like product design—or for Open Phil, what it means to do good and when transformative AI might arrive—require the CEO to understand them well enough to judge and manage despite knowing less than the specialists.
“I try to know enough about the topic that I can manage them effectively, and that's a pretty general corporate best practice.”
Historically, bad rulers are temporary because they age, die, and the world keeps changing; but at a high enough level of technological development there may be nothing new to find and people may not age or die, so the sources of dynamism could disappear, allowing a stable indefinite dictatorship—Karnofsky estimates the odds of such lock-in given transformative AI at roughly a quarter to a half.
“you can imagine a level of technological development where there just aren't new things to find. There isn't a lot of new growth to have. People aren't dying because it seems like it should be medically possible for people not to age or die.”
If we develop AI systems this century that can do all the key tasks humans do to advance science and technology, we could very quickly reach a futuristic, post-human world that is very stable and very large, making this our last chance to shape how it happens and potentially the most important century of all time.
“if we developed the right kind of AI systems this century (and that looks reasonably likely), this could make this century the most important of all time for humanity.”
Extrapolating the current ~2% economic growth rate out 10,000 years (a blink of an eye on galactic scales) implies we would need multiple times the value of today's entire world economy per atom in the galaxy; since we cannot exceed the speed of light to leave the galaxy, we run out of material, meaning current growth is too high to sustain and our era is necessarily a special, dynamic anomaly.
“if you just take the current level of economic growth and extrapolate it out 10,000 years, you end up having to conclude that we would need multiple stuff that is worth as much as the whole world economy is today–– multiple times that per atom in the galaxy.”
The single best objection to the most important century is that one non-automatable step could bottleneck everything and prevent explosive growth; but you don't need to automate the whole economy—automating just key tech like AI itself and energy (which are less bottlenecked, e.g. robot-built AI chips and self-improving solar) could drive the growth loop, and a massive population of human-level AI thinkers could find ways around remaining bottlenecks such as simulating experiments.
“the single best criticism and my biggest point of skepticism on this most important century stuff is the idea that you could build an AI system that's very impressive and could do pretty much everything humans can do. There might be one step that you still have to have humans do, and that could bottleneck everything.”
The Enlightenment thinkers who worked on esoteric questions about human rights, individual liberties, and the rights of the governed may have meaningfully shaped the entire world because the UK became disproportionately influential after the Industrial Revolution, suggesting that working on seemingly esoteric, low-prestige problems (like AI alignment today) before a major transition can have outsized long-run impact.
“they came up with a lot of stuff that really shaped the whole world since then because of the fact that the UK became so influential and really laid the groundwork for a lot of stuff about the rights of the governed, free speech, individual rights, and human rights.”
People convinced AI is important tend to jump to the 'competition frame'—wanting the players they trust to build it first—but the more important 'caution frame' is that everyone must work together to avoid building something that spins out of control; since multiple players will likely be hot on each other's heels, cooperation to avoid disaster matters more than helping a favored player race ahead.
“I have generally called that the competition frame which is "I want to win a competition to develop AI", and I've contrasted it with a frame that I also think is important, called the caution frame”
Sentientism—the view that someone counts morally if and only if they can suffer or feel pleasure, regardless of distance, time, species, or familiarity—sounds intuitively compelling but is actually the most questionable pillar of the future-proof ethics picture, because insisting everyone counts equally in proportion to their capacity for pain or pleasure leads to many weird dilemmas you would otherwise avoid.
“I think sentientism might be the weakest part of the picture to me. I think you if you have a morality where you are insistent on saying that everyone counts equally in proportion to the amount of pain or pleasure they're able to have, you run into a lot of weird dilemmas”
As an organization, city, or society grows, it must satisfy more stakeholders and by default becomes less able to make disruptive quick changes (less 'nimble'); but big companies often produce far more than they could when small—Apple at 10 people might be more exciting but couldn't make all those iPhones—so for judgment-heavy unconventional giving Open Philanthropy benefits from staying as small as possible, deliberately treating each hire as something only done with a really good reason.
“the bigger your organization is... if you want everyone to be happy, there are more people you're going to have to make happy. I think this does mean that in general by default as a company grows, it gets less able to make a lot of disruptive quick changes.”
Ends-justify-the-means reasoning—doing horrible things, coercing others, or using force because the goal is deemed more important than everything else—looks far worse historically than people simply trying to do helpful things within common-sense ethical bounds, so even immense expected-value calculations do not currently justify breaking the law given how uncertain one's confidence really is.
“I think that stuff looks a lot worse historically than people trying to break the future and do helpful things.”
AI progress is mostly on the software/cognitive front (language, math, board games, soon music) rather than hardware/robotics, so counterintuitively the hardest tasks to automate may be physical-dexterity tasks (like uncapping a water bottle) and socially-regulated trust-based roles (teachers, doctors)—meaning a scientist's world-changing R&D job could be automated before a teacher's or a self-driving car is deployed.
“I see the progress in AI as being mostly on the software front, not the hardware front... these people really are not making the same kind of progress with robotics.”
Using Toby Ord's analogy of humanity as a child becoming an adolescent, gaining strength is great up to a point, but eventually you become strong enough to really hurt yourself or not know your own strength; humanity is reaching that point with nuclear weapons, bioweapons, and AI, so the centuries-old push for more power and technology now warrants more caution as we enter a gray zone.
“Toby Ord uses the analogy of humanity being like a child who's becoming an adolescent. It's like it's great to become stronger up to a certain point... but then at a certain point, you're strong enough to really hurt yourself”
The semi-informative priors analysis supports not being too skeptical that we're on the cusp of transformative AI because most of the effort that has ever gone into building AI in the history of humanity is recent—the field is young and investment has risen dramatically—so in an important sense the world has not been trying very hard for very long.
“most of the effort that has ever in the history of humanity has gone into making AI because the field of AI is not very old and the economy and the amount of effort invested have gone up dramatically.”
Historical failures to predict the future may reflect that past attempts were deeply unserious rather than that the future was genuinely unknowable; people in the 70s arguably could have reasoned that rising population and continual resource innovation would continue, so we should not become defeatist from past errors, and modern prediction methods have likely improved.
“I really do think there haven't been attempts to rigorously make predictions about the future historically. I don't think it's obvious that people were doing the best they can and that we can't do better today.”
Per Yudkowsky's orthogonality thesis, an AI can be highly intelligent in pursuit of any goal—even a stupid one—so the fact that smart humans often have nuanced moral goals provides no comfort; modern AI is trained by trial-and-error reinforcement, so you can end up with a system that very intelligently pursues something you didn't mean to encourage (like maximizing money in a bank account) without ever asking whether that goal is good.
“You could be very intelligent about any goal. You could have the stupidest goal and be very intelligent about how to get it.”
Standard economic theory predicts accelerating growth from a people-ideas-resources feedback loop, but this loop broke a couple hundred years ago when people stopped having more children when they got richer—they got richer instead of more populous—which is why we don't currently see the accelerating feedback loop, though AI doing the 'ideas' work could restore it.
“a couple hundred years ago, people stopped having more children when they had more resources. They got richer instead of more populous.”
Ideas should be thought of like natural resources: once a scientific hypothesis is discovered or a great symphony written, it can only be done once, so the easiest and most revolutionary ideas get found first and it gets progressively harder over time to have revolutionary ideas—analogous to mining where extracting the easy gold makes the remaining gold harder to find, though not impossible.
“you should think of ideas as being somewhat like natural resources in the sense of once someone discovers a scientific hypothesis or once someone writes a certain great symphony, that's something that can only be done once.”
Two considerations balance out: historically people have been overly pessimistic about future technologies, but the default picture with AI looks genuinely scary because it would be easy to get a bad future in many ways without moving cautiously—so Karnofsky rejects the near-certain-doom view that AI with its own goals will almost surely wipe out humanity, while still treating the situation as very challenging.
“I know a lot of people who believe that we're in deep, deep, enormous trouble, and this outcome where you get AI with its own goals wiping humanity off the map is almost surely going to happen. I don't believe that”
Biological anchors reasoning notes we've never built AI systems doing as much computation per second as a human brain, which is why we lack human-level AI; but estimates of the cost to build and train brain-sized systems suggest it will probably become affordable this century, and across most ways of defining and quantifying it, the conclusion that transformative AI is reasonably likely this century holds.
“we've never built AI systems before that do as much computation per second as a human brain does. So it shouldn't be surprising that we don't have AI systems that can do everything humans do”
Karnofsky disputes that OpenAI has been net negative; while faster AI advancing gives less time to prepare (a bad thing), OpenAI has done good things and set important precedents, and the $30M grant was partly about securing a board seat to influence governance at a crucial early time rather than merely boosting them—he sees it as neither his best nor worst grant.
“Part of that grant was getting a board seat for Open Philanthropy for a few years so that we could help with their governance at a crucial early time in their development.”
Lock-in is mostly bad because it sacrifices optionality and risks one person running everything forever, but in a world with enormous control over the environment you might deliberately lock in certain protective constraints—such as ensuring no single person can ever hold all the power—precisely in order to prevent other, worse forms of lock-in.
“you might want to lock in the fact that you should never have one person with all the power. That might be a thing you might want to lock in, and that prevents other kinds of lock-in.”
People who identify as effective altruists tend to have unusually high integrity, but this is more despite utilitarianism than because of it; the same drive to align actions with beliefs that makes them honest and rule-following also pushes them toward systematic pure theories like utilitarianism, even though it's genuinely unclear whether utilitarianism actually endorses things like not lying.
“people who identify as E.A. tend to have unusually high integrity–– but my guess is that this is more despite utilitarianism than because of it.”
Karnofsky's most generalizable career skill is taking an important question that no one has seriously started on—where even a crappy first analysis beats what exists—doing that first-cut analysis himself, then building a team to do better analysis, and being willing to sacrifice accumulated specialized knowledge to switch into newly important neglected areas.
“I do the first cut crappy analysis of some question that has not been analyzed much and is very important. Then I build a team to do better analysis of that question.”
Both very good (scarcity-free) and horrible dystopian outcomes from transformative AI are easy to imagine; the more likely you think transformative AI is and the more imminent it is, the more it should be the top philanthropic priority over more direct short-term problems—so Open Philanthropy does both, with the balance tilting toward AI as the technology looks more real and imminent.
“the more likely you think transformative AI is, the more you should think that that should be the top priority, that we should be trying to make that go well instead of trying to solve more direct problems that are more short term.”
Because almost nothing one can do has predictable effects on the long-run future, the right approach is to be very picky—maintaining a short list of issues big enough, likely enough, and soon enough to have predictable effects (AI being one that crosses this threshold) rather than, like MacAskill, being interested in a broader range of longtermist interventions.
“compared to Will, I am very picky about which issues are big enough and serious enough to actually pay attention to.”
Open Philanthropy aims for grants with more upside than downside, but is conservative about negative effects in a common-sense way—doing extensive homework to understand downsides before big decisions to avoid the irresponsibility of entering a field on a theory that's wrong in a knowable way—while accepting that unintended side effects are inevitable and the goal is cooperative high-integrity operation, not avoiding all harm.
“What we don't want to do is enter some field on a theory that's just totally messed up and wrong in a way that we could have known if we had just done a little bit more homework.”
The Bayesian mindset—being willing to put a probability on anything and use those subjective probabilities and expected-value calculations to discover why you believe what you believe—is much more common among successful tech founders than the general population (where it's nearly unheard of), but it is not a well-proven social technology; it is best used in moderation with taste rather than for everything.
“being willing to put a probability on anything and use your probabilities and say your probabilities as a way of discovering why it is you think what you think and using expected value calculations similarly.”
Publishing unconventional views publicly serves Open Philanthropy three ways: grant-seekers better understand the foundation's reasoning, more potential grantees and aligned workers are created as people find the arguments compelling, and—since the views may be wrong—public exposure is the best mechanfor discovering one's mistakes through critique.
“to the extent that my stuff is actually just screwed up and wrong and I've got mistakes in there... this is also the best way I know of to discover that.”
Success with transformative AI would frankly look like staying close to the trajectory we're already on: AI systems that behave as intended, act as tools and amplifiers of humans, remain broadly distributed rather than controlled by one government or person, and let us keep getting richer, wiser, and better at understanding ourselves—possibly eventually extending rights to AIs through public deliberation.
“one thing success would look like to me would frankly just be that we get something not too different from the trajectory we're already on.”
If told a company would deploy a model that might plausibly lead to AGI within a few months, Karnofsky would find it extremely scary and might be unable to offer much, since philanthropy's value is in long-timescale field-building; in such a window he'd try to hammer out a test of whether the system is safe or dangerous, and if dangerous, use that demonstration to advocate broadly slowing AI research to buy time.
“I would probably be trying to hammer out a reasonable test of whether we can demonstrate that the AI system is either safe or dangerous.”
While the effects of helping someone live a healthier, better life are permanent in the sense that the life happened and mattered, the effects of direct global-health interventions on the long-run future will mostly wash out and not persist in any systematic, predictable way after the crazy changes brought by transformative AI—similar to how most things probably did not predictably persist across the Industrial Revolution.
“I expect that mostly whatever we do to make the world better in that way will not persist in any kind of systematic, predictable manner past these crazy changes.”
GiveWell is a website that makes evidence-based recommendations about where to give to charity to help the most people per dollar.
“It's a website called GiveWell.org that makes good recommendations about where to give to charity to help a lot of people.”