factual
Public benchmarks are gameable and misleading
Open-source benchmarks like LM Arena are skewed toward a narrow set of use cases that don't match what normal people actually do in a product, and they are easily gameable — Meta's team could easily tune a Llama 4 Maverick version to top the Arena, but the released untuned model ranks lower, so optimizing for these benchmarks leads product development astray.
factualpending
Speaker
Mark ZuckerbergEvidence Quote
“The issue with open source benchmarks, and any given thing like the LM Arena stuff, is that they’re often skewed toward a very specific set of uses cases, which are often not actually what any normal person does in your product.”
Created: 6/13/2026, 3:33:00 AM
My Notes
Loading notes...