The Llama 3.1 405 billion parameter model requires 228GB of download storage and approximately 200GB of RAM to load, but produces only one token every several seconds on a $50,000 Dell workstation with 512GB RAM and dual overclocked Thread Ripper 96-core CPU with RTX 6000 Ada GPU, taking approximately 30 minutes to generate a simple response.
factualpending
Speaker
DaveEvidence Quote
“this is almost as bad as the regular model running on the pie”
Created: 8/13/2026, 9:51:22 AM
My Notes
Loading notes...