The Nvidia RTX 4080 GPU in a consumer gaming PC can run Llama 3.1 (70 billion parameters) at approximately ChatGPT speed with 16GB of host memory usage and averaging 75% GPU utilization with spikes to 100%, making this configuration faster than ChatGPT for local inference.

factualpending

Speaker

Dave

Evidence Quote

it runs local inferences fast or faster than chat GPT while using a reasonably competent model

Source

Run Local LLMs on Hardware from $50 to $50,000 - We Test and Compare!Dave's Garage
Created: 8/13/2026, 9:51:22 AM

My Notes

Loading notes...