Joel Becker
About
Member of Technical Staff at METR (leads methodology and technical implementation of time horizon measurement)
Cast within
No topic-region cast yet — this appears once Joel Becker's compiled claims are aligned into a topic region's argument tree.
Claims by Joel Becker (13)
Chinese AI models are approximately 9-12 months behind US models by release timing, and the gap is potentially even larger when measured by time horizon, with some evidence suggesting Chinese models perform better on benchmarks than on held-out problems, possibly due to benchmark overfitting.
Real-world software engineering tasks differ from METR benchmarks in several ways: they involve code quality concerns (elegant vs. messy code), collaboration with other engineers, larger codebases, adversarial scenarios where others modify code you're working on, and require verification overhead because humans must review AI work without context, leading to productivity gaps not captured by benchmark scores.
Time horizon charts plot the difficulty of tasks AI systems can complete over time, where difficulty is measured by how long it takes humans to complete those same tasks under identical conditions, showing an exponential increase in AI capabilities with doublings occurring roughly every 4 months in recent trends.
Human baselines are established by recruiting talented humans with relevant expertise (e.g., software engineers for software tasks, ML engineers for ML tasks) who are not familiar with the specific task beforehand, timing how long it takes them to complete the task successfully using the same tools the AI will use, and averaging across approximately three baselines per task.
The 50% success rate threshold was chosen as the difficulty metric not because it represents operational readiness, but because it is statistically robust (least sensitive to label noise and distribution thickness), appears in prior literature, and represents the point where a model is more likely to succeed than fail at a task given only the human completion time.
Joel suspects that METR's task distribution is increasingly becoming a narrower slice of all possible tasks and specifically overlaps more with the exact task distributions used by AI labs for training, meaning METR is measuring progress on tasks optimized for (rather than independent test of) lab capabilities.
Joel acknowledges that METR's baseline methodology is imperfect and would ideally have 100x more resources (100 baselines per task, top-tier engineers, wider task distributions), but current limitations don't invalidate the core finding of exponential progress because doubling time is robust to baseline variability.
My Notes
Loading notes...