The Nvidia RTX 4080 GPU in a consumer gaming PC can run Llama 3.1 (70 billion parameters) at approximately ChatGPT speed with 16GB of host memory usage and averaging 75% GPU utilization with spikes to 100%, making this configuration faster than ChatGPT for local inference.
factualpending
Speaker
DaveEvidence Quote
“it runs local inferences fast or faster than chat GPT while using a reasonably competent model”
Created: 8/13/2026, 9:51:22 AM
My Notes
Loading notes...