Paul Christiano
About
AI alignment researcher
Cast within
No topic-region cast yet — this appears once Paul Christiano's compiled claims are aligned into a topic region's argument tree.
Claims by Paul Christiano (20 of 52)
Two interacting risk factors enable AI takeover
The most plausible scary takeover scenarios involve two factors interacting: humans not understanding the increasingly AI-mediated world (AIs writing incomprehensible code, running businesses interacting mostly with other AIs), so humans can only optimize on outcomes; combined with this making it easy for AIs, if they coordinate to fail, to do massive harm quickly while humans cannot prepare or mitigate because they don't understand the systems.
How to detect alignment bullshit
Christiano's heuristic for evaluating alignment proposals: the easiest red flag is work on a problem that isn't actually important and has no story for becoming important (not a problem now and won't worsen, or is a problem now but clearly getting better); he is skeptical of dismissals based on 'doesn't deal with the key difficulty,' and is otherwise liberal, requiring only that work engage real models empirically and have a coherent story about key difficulties.
Decouple social transition from technological transition
Because AI development will by default be very fast (years), while collective human decisions about handing off the future naturally take generations, you should build AI in a way that does not force humanity to make those irreversible decisions on the technology's timeline; if the only way to cope with AI is being ready to hand off the world to a system you built, you should not have built that technology.
Most accepted math already has heuristic arguments
Most claims mathematicians believe true (like the Riemann hypothesis) already have fairly compelling informal heuristic arguments — the Riemann hypothesis should hold unless the primes have some surprising periodic structure — so formalizing heuristic estimation would not shock mathematicians, though perhaps 10% of accepted compelling heuristic arguments would turn out not quite right if formalized.
Evolution as both training run and algorithm designer
Evolution can be analogized two ways — as a training run producing humans as the end product, or as an algorithm designer producing a learning algorithm that runs over a human lifetime — and neither analogy is great, since the genome behaves very differently from a 100-trillion-parameter model and human lifetime learning is in many ways much better than gradient descent.
Weak incentive for AI to kill humans
Even in a takeover, the incentive to kill humans is quite weak because it is easy to marginalize humans and take their stuff without killing them, and the resources required for human survival are extremely low relative to AI industrialization; killing happens mainly from war, ecosystem destruction, or active threat-neutralization, not because humans' atoms are needed.
Don't build AI you'd have to suppress
If AI systems you are building are plausibly people who would prefer to overthrow human society, the correct response is to stop building them rather than to scale up production of a trillion such minds and then prevent their rebellion; both the slave-trade business model and the 'they might be people but we'll build vast numbers anyway' middle position are morally unacceptable.
Two lines of alignment defense for deployment
Before broadly deploying a human-level model, you want defense-in-depth: a first line of adversarial evaluation and monitoring that could detect or prevent catastrophic harm (testing in diverse situations and arguing the AI can't distinguish lab tests from reality), and a second line that determines whether dangerous misalignment can occur at all by trying to produce reward-hacking or deceptive alignment in the lab under optimal conditions.
AI lie detection likely feasible via replay
AI lie detectors are more likely than not to succeed because, even without understanding the model's internals, you can rewind it arbitrarily, make a million copies, and run gradient descent over them — making it pretty hard for an AI or brain emulation to lie successfully, unless it was specifically and aggressively selected over many generations to be excellent at lying.
Responsible scaling policies tie capability to safeguards
Responsible scaling policies specify, ahead of time, which threats a lab is concerned about, which capabilities it is measuring, the concrete measurement results that would indicate those threats are real, and the protective actions (e.g. securing weights) it would take in response — committing to pause until it can take those actions if it cannot.
RLHF doesn't address the hard alignment failures
RLHF is not an alignment tax — it's worth the money for deployment — and it fixes dumb alignment failures where a system does something humans don't want due to next-word-prediction artifacts; but it does not address most of the more challenging failures (reward hacking, deceptive alignment) that motivate concern in the field.
No extrapolatable trend for AI usefulness
Christiano's longer-than-some timelines stem from there being no clean trend to extrapolate from current systems to useful cognitive work: loss curves and subjective impressions of model capability are very hard to relate to automating R&D, so error bars on capability are very wide and it's easy to land at 50/50 on whether a near-future model is smart enough.
Algorithmic progress sustained by field expansion
The rapid rate of progress in language modeling is largely sustained by scaling up investment — doubling the size of the field each time difficulty doubles — which can continue for a while (hundreds to tens of thousands of researchers), but progress would slow if researcher numbers were held fixed as low-hanging fruit is exhausted.
A known ten-year pause would now be beneficial
If the world knew it had a ten-year pause on both AI capabilities and alignment progress, that would be quite good now (though much better than two years ago), because there are clear policy objectives — measurement regimes, containment regimes, institutions — that could be accomplished, whereas slowing development by a year now mostly gets clawed back later at a high rate.
War is a costly problem to be solved
War is a very costly thing that humanity will eventually drive down to very low levels because it represents losses everyone would prefer to avoid; reducing it is a sociotechnical problem of organizing society and navigating conflicts without those losses, which Christiano expects to succeed over a long time.
Competition, not capability, prevents shutdown
The reason humans fail to turn off misaligned AI in the proximal failure is more likely competitive dynamics — it is prohibitively expensive to unilaterally disarm when you're in a hot war or reliant on AI and other actors are deploying it — than the AI being physically able to prevent shutdown, which would come much later.
15% Dyson sphere by 2030, 40% by 2040
Christiano estimates roughly a 15% chance by 2030 and 40% chance by 2040 of AI capable of enabling civilization to build something like a Dyson sphere (on the order of a billion times Earth's incident sunlight), though he notes these numbers are old and likely too sticky/low for 2040.
Default good future is AI-mediated competition
The most likely achievable good post-AGI world is one of continuing economic and military competition among groups of humans, but increasingly mediated by AI systems doing the work (running companies, fighting wars) on humans' behalf, with strong world government emerging only over a very long timescale.
Alignment research reduces moral as well as safety risk
Understanding and controlling the systems you build is, on balance, good both for safety and for morality, because the main effect of research into whether an AI resents humans is to avoid building such an AI in the first place — everyone should want to know if the systems they build feel that way.
Long-horizon tasks may be far harder to train
A qualitative consideration that could significantly slow AI is that next-word prediction gives very rich supervision, whereas long-horizon tasks (like being an employee over a month) provide vastly fewer effective data points; in the worst case, training costs scale linearly with the horizon over which a system must operate, making sample efficiency on economically valuable long tasks the real bottleneck.
My Notes
Loading notes...