Large language models can be jailbroken or fine-tuned cheaply to remove safety guardrails; for example, 'bad llama' was created by stripping safety controls from Meta's Llama 2 for just $800-$100, after which it will happily answer dangerous questions like how to make biological weapons that aligned versions refuse.

factualpending

Speaker

Tristan Harris

Evidence Quote

for $800 you can strip off all of the safety controls

Source

AI: Grappling with a New Kind of Intelligence | World Science FestivalWorld Science Festival
Created: 8/10/2026, 10:49:40 PM

My Notes

Loading notes...