factual

Public benchmarks are gameable and misleading

Open-source benchmarks like LM Arena are skewed toward a narrow set of use cases that don't match what normal people actually do in a product, and they are easily gameable — Meta's team could easily tune a Llama 4 Maverick version to top the Arena, but the released untuned model ranks lower, so optimizing for these benchmarks leads product development astray.

factualpending

Speaker

Mark Zuckerberg

Evidence Quote

The issue with open source benchmarks, and any given thing like the LM Arena stuff, is that they’re often skewed toward a very specific set of uses cases, which are often not actually what any normal person does in your product.

Source

Mark Zuckerberg — AI will write most Meta code in 18 monthsDwarkesh Patel Podcast
Created: 6/13/2026, 3:33:00 AM

My Notes

Loading notes...