causal

Bandwidth per edge area, not bits per wafer, is the key metric

Switching accelerators from HBM to commodity DDR to save wafers fails because chip I/O escapes only on the edges; an HBM4 stack delivers ~2.5 TB/s per ~13mm shoreline versus only ~64-128 GB/s for DDR in the same area—an order of magnitude less bandwidth—and since FLOPS are gated by feeding weights and KV cache, the metric that matters is bandwidth per wafer edge, not bits per wafer.

causalpending

Speaker

Dylan Patel

Evidence Quote

There’s an order of magnitude difference in bandwidth per edge area.

Source

Dwarkesh Patel Podcast - Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI computeDwarkesh Patel Podcast
Created: 6/13/2026, 3:34:26 AM

My Notes

Loading notes...