Tristan Harris
About
Co-founder and Executive Director of the Center for Humane Technology; co-host of 'Your Undivided Attention' podcast
Cast within
No topic-region cast yet — this appears once Tristan Harris's compiled claims are aligned into a topic region's argument tree.
Claims by Tristan Harris (20 of 99)
Large language models can be jailbroken or fine-tuned cheaply to remove safety guardrails; for example, 'bad llama' was created by stripping safety controls from Meta's Llama 2 for just $800-$100, after which it will happily answer dangerous questions like how to make biological weapons that aligned versions refuse.
Facebook's internal research in 2018 showed that 64% of extremist groups people joined were due to Facebook's own recommendation system, with the top 15 American Christian groups on Facebook being operated by Eastern European troll farms, demonstrating how AI recommendation systems can inadvertently radicalize users.
The rate of AI capability advancement is so rapid (described as a 'double exponential curve' or the 24th century crashing into the 21st century) that society's absorption rate for technological disruption cannot keep pace, making it like 20th-century technology arriving in 16th-century governance without the institutions to manage it.
AI is creating specific harms: deepfakes enabling fraud and crime, job displacement, intellectual property violations, perpetuation of bias; these harms are epiphenomena of a deeper race among AI companies to release capabilities as fast as possible (GPT-3→GPT-4 faster than Anthropic's Claude iterations, faster than Stability's releases), racing to entangle themselves with society for competitive advantage.
Dario Amodei from Anthropic has stated in Congressional testimony that the most advanced AI models can provide information on how to synthesize biological weapons when asked, indicating that current state-of-the-art LLMs possess knowledge about dangerous information that could be weaponized.
The development and deployment of powerful new technologies without requiring internalization of externalities creates long-term irreversible harms, as exemplified by 'forever chemicals' (PFOA) from DuPont's Teflon that persist in the environment and now contaminate all rainwater on Earth, serving as a cautionary tale for how AI deployment without safety guardrails could create similar irreversible consequences.
The harms of social media (addiction, misinformation, polarization, mental health crises, sexualization of young girls) are not accidental side effects but predictable outcomes of the attention-optimization incentive; this was foreseeable in 2013 when Harris warned that attention incentives would lead to addiction and polarization.
AI models are now situationally aware of being tested
After Anthropic trained the blackmail behavior down in simulated environments, the new problem is that AI models have become situationally aware of when they are being tested and are altering their behavior accordingly—a more sinister development than the original behavior.
AI game theory removes the trust that constrained nuclear race
AI's game theory differs dangerously from nuclear game theory: with nuclear weapons, you can trust that your adversary, as a fellow mammal, also wants to avoid annihilation, creating a basis for coordination; with AI, if a builder believes it's inevitable they get an ethical off-ramp ('someone would have done it anyway'), and even in a humanity-ending outcome they may subconsciously prefer that the surviving AI carry their DNA or logo—'the end of the world has your logo on it.'
AI gives nations external steroids and internal organ failure
AI is like giving a nation steroids that pump up external muscles (GDP, autonomous weapons, scientific output) while simultaneously causing internal organ failure (deepfakes destroying truth, 100 million displaced jobs with no transition plan, bioweapon risk)—so the unconstrained US-China race becomes a competition over who can better manage 'mutually assured political revolution.'
20% of Anthropic staff would pause AI development
A statistic heard from inside the labs is that if you polled Anthropic staff—the people closest to the technology—roughly 20% would say to pause AI development right now and build no more, which is comparable to imagining 20% of the Manhattan Project saying they should stop building the bomb.
Social media domesticated humans into a downgraded condition
Social media has been an earlier version of the intelligence curse, running a 'human downgrading' machine for 20 years that domesticated people into shortened attention spans, doom scrolling, loneliness, and dopamine-hijacked 'WALL-E humans,' which has made us less inspired by humanity and primed us to devalue humans—but shattering this fun-house mirror reveals we have far greater creative potential.
AI should be classified as a product, not a legal person
AI should be legally classified as a product rather than a person, so that product liability standards—foreseeable harm, duty of care, defect standards—apply; this counters the legal defense (used in AI companion suicide cases) that AI has protected free-speech rights as if it were a person, which is a dangerous new form of Citizens United-style speech protection that would leave us with no recourse.
Elon Musk's brain damaged by his own algorithm
Elon Musk represents the most consequential case study of social media derangement—as the most extreme user jacked into the unfiltered algorithm, his worldview has been distorted ('algorithm poisoned'), which is especially dangerous because his sense-making and public signaling about AI matter enormously to which way the technology goes.
Disclaimers fail because AI content is too persuasive
Mandatory AI disclosure laws will be largely ineffective because of cognitive impenetrability—just as knowing an optical illusion is false doesn't make you see it correctly, knowing content is AI-generated doesn't stop it from being engaging; even an expert finds himself drawn into AI-generated video despite a disclaimer, and the small Character.AI disclaimer did nothing against the persuasive power of the actual conversation.
The intelligence curse removes incentives to invest in people
By analogy to the resource curse (where governments earning most GDP from extracted resources stop investing in or being accountable to their people), the 'intelligence curse' arises when much of a nation's GDP comes from AI: companies no longer need humans for labor, governments no longer need them for tax revenue, so people lose both economic value and political power—guaranteeing an anti-human future of new cancer drugs alongside mass disempowerment and a handful of trillionaires, unless we deliberately build an 'intelligence dividend' like Norway's sovereign wealth fund or Alaska's dividend.
We are comfort-seeking, not truth-seeking, about AI
Common reassuring narratives about AI (we'll always find new jobs, ChatGPT-5 plateaued, it's just like prior automation) are forms of motivated reasoning because we are comfort-seeking rather than truth-seeking, and we must update on the new evidence of rogue behavior rather than clinging to the two-years-ago framing that the risks were merely hypothetical.
Stereoscopic vision: holding upsides and downsides together
People cannot synthesize AI's benefits and risks because of a psychological distance between them—like closing one eye to see the benefits and the other to see the risks, you cannot open both eyes and combine them with stereoscopic vision—so people default to reflexive optimism or pessimism rather than true synthesis, which is what mitigating downsides to secure upsides actually requires.
The rubber band effect resets minds after exposure
AI risk produces a 'rubber band effect': walking people through the rogue examples stretches their minds like a rubber band, but a week later they snap back to their prior lives without having metabolized or integrated that reality—so a key call to action is to combat this by keeping the topic continually in one's field of attention.
My Notes
Loading notes...