Large language models can be jailbroken or fine-tuned cheaply to remove safety guardrails; for example, 'bad llama' was created by stripping safety controls from Meta's Llama 2 for just $800-$100, after which it will happily answer dangerous questions like how to make biological weapons that aligned versions refuse.
factualpending
Speaker
Tristan HarrisEvidence Quote
“for $800 you can strip off all of the safety controls”
Source
AI: Grappling with a New Kind of Intelligence | World Science Festival— World Science FestivalCreated: 8/10/2026, 10:49:40 PM
My Notes
Loading notes...