Joel suspects that METR's task distribution is increasingly becoming a narrower slice of all possible tasks and specifically overlaps more with the exact task distributions used by AI labs for training, meaning METR is measuring progress on tasks optimized for (rather than independent test of) lab capabilities.
causalpending
Speaker
Joel BeckerEvidence Quote
“the tasks that meta is measuring performance on, you know, in some sense, a more and more narrow slice of possible tasks and in particular, a more and more narrow slice that is perhaps similar to the kinds of tasks that you'd expect these major AI companies to be training on”
Created: 8/12/2026, 6:03:24 PM
My Notes
Loading notes...