YouTube1h 1m· Apr 2026· cataloged

Understanding the Most Viral Chart in Artificial Intelligence | Odd Lots


What this covers

METR, which stands for Model Evaluation and Threat Researc, is focused on understanding the degree to which AI models can engage in autonomous, complex tasks. METR see this is as a particularly important benchmark, given the risk that AI could one day be engaged in recursive self improvement, taking humans out of the loop. But how do you really gauge a model's ability to do complex problems. And what is being measured for exactly? On this episode we speak with METR's President Chris Painter as well as Joel Becker, a member of the technical staff who works on evaluation methods for the organization. We discuss both the mechanics and the philosophy of METR's work, and what it means when we see a a chart showing that Clause Opus 4.6 can do a task that would take a human nearly 12 hours. Chapters: 00:00:00 - AI Productivity Discussion 00:03:41 - Understanding Time Horizon Charts 00:05:09 - What is Meter? 00:06:46 - Safety Mission vs Public Perception 00:09:26 - How Time Horizon is Actually Measured 00:11:27 - Human Baseline Methodology 00:13:24 - Task Distribution Focus 00:16:54 - Human Baseline Sample Sizes and Challenges 00:19:10 - 50% vs 80% Success Rate Charts 00:23:20 - Investment Interest and Public Information 00:25:29 - When to Worry About AI Autonomy 00:26:26 - Current AI-to-AI Collaboration 00:29:06 - Task-Specific Performance Variations 00:32:12 - Industry Dynamics and Safety Concerns 00:37:30 - Chinese AI Models Assessment 00:39:19 - Stakeholder Interactions 00:41:51 - Capitalism and Safety Tensions 00:44:25 - Compute Costs and Capabilities 00:46:03 - Baseline Methodology Criticisms 00:49:28 - Accelerating Progress Trends 00:52:40 - Team Size and Funding Model 00:54:44 - Talent Bottleneck and Research Priorities Watch Bloomberg Vodcasts: https://www.youtube.com/playlist?list=PLe4PRejZgr0MS0Tqk_zQl1LVVTdWuH2Di

Watch more Odd Lots episodes: https://www.youtube.com/playlist?list=PLe4PRejZgr0OJbRzA6nWybYiThLJd_ouz

Bloomberg's Joe Weisenthal and Tracy Alloway analyze the weird patterns, the complex issues and the newest market crazes. Join the conversation every Monday and Thursday for interviews with the most interesting minds in finance, economics and markets.

Join the conversation: discord.gg/oddlots

Subscribe to Bloomberg Podcasts: https://bit.ly/BloombergPodcasts

Check out more Odd Lots: https://youtube.com/playlist?list=PLe4PRejZgr0MuA6M0zkZyy-99-qc87wKV

Get the Odd Lots newsletter: https://www.bloomberg.com/account/newsletters/oddlots

And for all things Odd Lots, visit https://www.bloomberg.com/oddlots

#Investing #Markets #Finance #Bloomberg #Podcast #OddLots

Visit us: https://www.bloomberg.com/podcasts

For coverage on news, markets and more: http://www.bloomberg.com/video

Visit our other YouTube channels: Bloomberg Television: https://www.youtube.com/@markets Bloomberg Originals: https://www.youtube.com/bloomberg

Source description (no synthesized summary yet).

Sharpest takeaway

METR's time horizon charts measure AI capability progress by task difficulty (human completion time), showing exponential improvements at ~4-month doubling rates, but the metrics have methodological limitations that may overstate real-world productivity gains and warrant scrutiny before informing major investment or policy decisions.

  • Time horizon measures task difficulty by human completion time, not autonomous runtime—a crucial distinction lost in popularization
  • The 50% success threshold is chosen for statistical tractability and prior literature, not because it represents operational readiness, and 80% charts tell a different story
  • Real-world AI productivity lags benchmarks due to code quality, team coordination, adversarial collaboration, and reliability verification overhead

The claims · ranked29 claims · weighted by value

This asset isn't compiled yet

You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.

0.80

Time horizon charts plot the difficulty of tasks AI systems can complete over time, where difficulty is measured by how long it takes humans to complete those same tasks under identical conditions, showing an exponential increase in AI capabilities with doublings occurring roughly every 4 months in recent trends.

factualhigh valueestablishednovelty 2/4durability 4/4· Joel Becker

So fundamentally, you know, in simpler terms, we are plotting the difficulty of tasks. The AI is are able to complete overtime and, you know, the particular way that we measure the difficulty of tasks is in how long it takes humans to complete, to complete those same tasks that we're asking the AI to do.

0.80

The 50% success rate threshold was chosen as the difficulty metric not because it represents operational readiness, but because it is statistically robust (least sensitive to label noise and distribution thickness), appears in prior literature, and represents the point where a model is more likely to succeed than fail at a task given only the human completion time.

factualhigh valueestablishednovelty 2/4durability 4/4· Joel Becker

I think there are reasons for picking the 50%. One in particular. It's the one that's, statistically we're better able to to, to measure for some, for some technical reasons. It's the one that shows up in, in previous literature.

0.69

Human baselines are established by recruiting talented humans with relevant expertise (e.g., software engineers for software tasks, ML engineers for ML tasks) who are not familiar with the specific task beforehand, timing how long it takes them to complete the task successfully using the same tools the AI will use, and averaging across approximately three baselines per task.

factualhigh valueestablishednovelty 1/4durability 3/4· Joel Becker

So so first we come up with the tasks and that's, you know, that's a whole lot of kettle of fish. We can we can talk about exactly how we do that. And then using essentially the same tools that we're about to give the AIS, we take, talented humans, you know, not people who have seen this particular type of task before, but people who have relevant expertise. So if it's a software engineering task, you know, they have software engineering expertise, machine learning task, they have machine learning expertise.

0.65

Many people in the Bay Area who are concerned about AI safety have the intuition that the best way to influence AI's future is to work inside the industry building the technology, but this creates a structural problem: the logic of 'if I don't build it, someone else will' encourages continuous advancement, and combined with competition between labs and geopolitical competition with China, even safety-motivated developers face incentives to build more advanced systems.

causalhigh valuecontestednovelty 2/4durability 4/4· Chris Painter

Many people had this intuition that the thing to do is go and work in the industry, because if you're, like, helping build it, you know, what's the best way to shape the future? It's to build it. And I think that one, there's obviously you could have questions about how sincere that is for for many of the people, who are in the industry, or if there's kind of a mix of different motivations

0.61

METR is a research nonprofit dedicated to measuring whether and when AI systems might pose catastrophic risks to humanity, specifically focusing on threats from AI autonomy by assessing how autonomous AI systems are and the scale, length, and difficulty of tasks they can perform independently.

definitionhigh valueestablishednovelty 1/4durability 3/4· Chris Painter

So METR is a research nonprofit based in the Bay area, like you said, dedicated to advancing the science of measuring whether and when I systems might pose catastrophic risks to humanity as a whole. Focus specifically on threats that come from AI, autonomy or AI systems themselves.

0.61

Chinese AI models are approximately 9-12 months behind US models by release timing, and the gap is potentially even larger when measured by time horizon, with some evidence suggesting Chinese models perform better on benchmarks than on held-out problems, possibly due to benchmark overfitting.

factualhigh valuecontestednovelty 1/4durability 2/4· Joel Becker

The, the models that we, we anticipate being on the frontier and in general, the Chinese models have been something like, you know, 9 to 12 months, let's say, behind the behind the US models.

0.55

Real-world software engineering tasks differ from METR benchmarks in several ways: they involve code quality concerns (elegant vs. messy code), collaboration with other engineers, larger codebases, adversarial scenarios where others modify code you're working on, and require verification overhead because humans must review AI work without context, leading to productivity gaps not captured by benchmark scores.

causalhigh valuespeaker onlynovelty 2/4durability 3/4· Joel Becker

One is that the scoring implicitly is different. In in real problems, I'm scoring based on, you know, something a bit more realistic than these, algorithmic scoring procedures, these automatic scoring procedures that we're using at meta, and many other people are using in the, in the benchmark world, there's some notion of code quality. If you're if you're working in software engineering. But but for other tasks there's, there's beautiful code. Elegant code piece was talk about.

0.52

There is a strange dynamic in the AI industry where lab CEOs and safety researchers make similar alarming claims about AI risks (destroying the world, taking over), creating a 'Baptist and bootlegger' alignment between people building the technology and people warning about it, which is unusual compared to most industries.

factualhigh valuespeaker onlynovelty 2/4durability 3/4· Joe Weisenthal

there's this kind of, sort of Baptist and bootlegger relationship between the AI labs, people who are building this stuff and the sort of alignment safety people, and they sort of go back and forth in like the you have the heads of the lab saying, yes, this might destroy the world and take all your jobs of the safety people and the alignment. People say, yes, this might destroy the world. And like I, it's a very strange industry, right?

0.52

Joel suspects that METR's task distribution is increasingly becoming a narrower slice of all possible tasks and specifically overlaps more with the exact task distributions used by AI labs for training, meaning METR is measuring progress on tasks optimized for (rather than independent test of) lab capabilities.

causalhigh valuespeaker onlynovelty 2/4durability 3/4· Joel Becker

I have some suspicion that the tasks that meta is measuring performance on, you know, in some sense, a more and more narrow slice of possible tasks and in particular, a more and more narrow slice that is perhaps similar to the kinds of tasks that you'd expect these major AI companies to be training on in the first instance.

0.48

Chris Painter's personal view is that if AI self-improvement and AI research automation are plausible future developments, then all of humanity being aware of this trajectory is a precondition for figuring out how to respond to it, and therefore METR prioritizes public information over selective knowledge advantage for any particular group.

normativehigh valuespeaker onlynovelty 1/4durability 3/4· Chris Painter

Maybe this is more a personal view, but I think if this is possible that we will automate AI research, I think all of humanity being aware of it is kind of a aware of where we're heading is sort of a precondition for us all being able to figure out what to do about it. And so I don't kind of want like certain people or one side or one team to kind of like selectively be in the dark because they might invest on the basis of this or something like that.

0.48

The AI labs themselves spend more time and resources thinking about frontier capabilities assessment than government does, but this could change if government dedicates more resources to the problem.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joel Becker

There's some sense in which, like the companies in the industry, spends more time thinking today about frontier capabilities assessment than the government does. Yeah. I think, like one day you could imagine us getting to the point where the government is, like very focused on this and dedicating a lot of resources to it. And at that point, I would expect meteor to be spending more time talking to governments.

0.48

OpenAI and Anthropic were founded with exotic corporate structures (private company owned by nonprofit, etc.) because the founders took seriously the existential risks of AI and wanted to self-limit, suggesting at least some industry actors are not purely motivated by profit.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joe Weisenthal

It was also true that open AI and anthropic, but open a little more. We're like founded with these very exotic corporate structures, like a private company owned by nonprofit, etc., which they presumably did because they took pretty seriously the fact that this technology is a science.

0.48

Choosing the shortest or longest observed baseline time from METR's measurement pool, rather than averaging, probably wouldn't change the final measurement significantly—it would just shift the absolute numbers, not the doubling rate that drives the headline findings.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joel Becker

some of the details, at least for the work we've done so far, you know, aren't going to matter as much as you might naively think. So choosing the, you know, shortest baseline time that we that we, at the we end up observing or the longest time, you know, it's actually not going to make that much difference to the final measurements, you know, of course, we we do think these people, talented software engineers...perhaps we could have found even more talented people. They would have completed it in half the time...but of course, that that wouldn't change. The wouldn't change the doubling time. It would mean you'd get to the same level after another, another four months.

0.48

Joel acknowledges that METR's baseline methodology is imperfect and would ideally have 100x more resources (100 baselines per task, top-tier engineers, wider task distributions), but current limitations don't invalidate the core finding of exponential progress because doubling time is robust to baseline variability.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joel Becker

I think I think it just is true that, baseline methodology or the ways in which we compare to humans in some ways leaves a lot to be desired. That's, you know, ideally we would have invested, you know, a hundred times as many resources in, you know, having 100 baselines and baselines per, task.

0.48

METR's time horizon charts have become the industry standard benchmark for measuring AI progress, and investment decisions are explicitly being made based on these charts as investors interpret them as direct evidence of AI capability and business value.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joe Weisenthal

How much interest you get on these charts from potential investors specifically? And the reason I ask is because, I was just messing around and like googling some stuff. And when the opus chart, the latest opus chart came up, someone posted it on Reddit and I think, like the second comment on it was someone going, how do I invest in open AI?

0.48

There is a potential conflict of interest in METR's baseline methodology: humans are paid to complete tasks and may be incentivized (consciously or unconsciously) to take longer to complete them to maximize payment, though METR mitigates this by paying bonuses for faster completion relative to peers.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joe Weisenthal

You say, Joe, come in and do this task. What is it? The how do you prevent me? Oh, man, this is taking me a long time. Meanwhile, I keep getting $100 an hour for, like, looking at my computer and time. Oh, this is tough. I'm gonna have to come back tomorrow and keep working on this.

0.48

The rapid progress in AI capabilities is creating real evidence that must inform policy, but policymakers today primarily care about data center location and infrastructure questions rather than frontier capability assessment.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joe Weisenthal

when you say, you know, it's easy to imagine or, maybe the government will care more about this. Not so easy for me to imagine. It seems like they mostly care about, you know, data centers and, like, where they located and stuff like that. It would be nice if we had policymakers really looking at, like, frontier capabilities and stuff. Still seems kind of a way off.

0.45

The difficulty translating raw AI benchmark capabilities to real-world productivity shows that benchmark improvements overestimate real-world impact, but the overestimation is not huge and people are getting real utility from modern AI tools.

factualhigh valuespeaker onlynovelty 1/4durability 3/4· Joel Becker

I think there are a number of differences in a number of ways, in particular, in which the benchmark results are overestimating what we might see in the wild, you know, not not hugely overestimating. I think we do see that people are getting real utility out of these modern, AI tools, but overestimating to some extent.

0.45

Capability progress is unlikely to slow in the next few years because data center investments and construction plans are already baked in through 2027-2028, meaning compute input will continue even if investment stops today.

forecasthigh valuespeaker onlynovelty 1/4durability 2/4· Joel Becker

it's hard for me to consider it plausible that it will slow down in the next at least small number of years. Is that a lot of those compute R&D investments, basically already baked in, right? Like the centers have already been built, you know, plans for data centers even beyond, you know, 2027, 20, 28, presumably, you know, coming and coming to fruition, coming, coming about.

0.45

Currently, when AI systems work together (AI to AI), they fail because they generate collaborative hallucinations and tend to devolve into terrible outputs, suggesting AI-to-AI scaling is not yet a near-term threat.

factualhigh valuespeaker onlynovelty 1/4durability 2/4· Chris Painter

What actually happens today when AI is working with AI. Yeah. My sense is that, at some point, you know, a further away points than would have been true some, some time ago. The AI is more or less full on their faces. That's, that you know, there are some things they're not so capable of today, like. Collaborative hallucinations world do just like, you know, just like, devolve into terrible.

0.45

One narrative explanation for why labs are now seeing faster progress (4-month vs 7-month doubling) is that they have now narrowed their focus to software engineering and AI research tasks where they can apply intense optimization pressure, rather than working on more diverse consumer applications like image and video generation.

causalhigh valuespeaker onlynovelty 1/4durability 2/4· Tracy Alloway

You know, you had like OpenAI shutting down its like video efforts etc.. So perhaps part of the story is just this intense focus now on the software engineering side as what these labs are working on. Yeah. And sort of like all these other side quests are not as important. So maybe we will see even more rapid progress on some of these technical benchmarks, because clearly, from the labs perspective, that's where the action is more than some of these consumer things like making making images or videos.

0.42

METR is bottlenecked on technical talent and is unable to conduct research on a large percentage of world-important problems they've identified, with one brainstorm session identifying 20-30 important unsolved problems but only having capacity to work on 1-3 of them.

factualhigh valuespeaker onlynovelty 1/4durability 2/4· Joel Becker

I think clearly the central reason is that we are bottlenecked on, on technical talent, on, on, you know, incredibly capable people to come work on these questions. I was on a meta work rate recently where we were we were brainstorming, you know, 20, 30 of these what seemed like world important problems, problems that we think no one else is going to get to if we do not get to them. And we are able to to, to conduct research on on how many of those problems. I think it's one to, you know, maybe if we do an extraordinary job this quarter, it might be three

0.42

AI systems today can handle the human-directed idea phase (where a person specifies what they want) combined with AI handling software engineering implementation, but cannot yet fully autonomously manage research departments or operate independently from human ideation.

factualhigh valuespeaker onlynovelty 1/4durability 2/4· Chris Painter

the ways in which AI's are autonomous today, or close to autonomous today is the human has the idea and then, you know, submits that idea to, cloud code or Codex or one of these other AI tools, and then they handle the software engineering components. Possibly there's still still some, some intervention after that.

0.42

METR observes that domain-specific time horizon charts show different trajectories, with some showing steep exponential curves and others showing flatter or more irregular patterns, indicating variation in progress across different task domains rather than uniform acceleration.

factualhigh valuespeaker onlynovelty 1/4durability 2/4· Joe Weisenthal

So when you look at the domain specific time horizon charts, so the ones that show like, you know, the I think you call them task suites or something like that, like I guess productivity by specific job and you see these different lines. So sometimes you see like almost horizontal lines and sometimes you see squiggly or steeper lines.

0.30

There exist countervailing forces to safety-capability tension: some safety-promoting techniques make models more useful and compliant with user intentions, creating capitalist incentives to invest in certain kinds of safety research, though this doesn't rule out capabilities progress as a value axis.

factualspeaker onlynovelty 1/4durability 3/4· Chris Painter

I think that there are some forces on the other side, right. Like, you know, some safety promoting technologies quote unquote, or techniques, do make the models more useful, you know, if they're, if they're better complying, better complying with your will in some sense. And so you have capitalist incentives, standard capitalist incentives to, to invest in that kind of research.

0.25

There is a difference between 7-month and 4-month doubling times that matters more rhetorically than practically—if AI is destroying white-collar work in 2 years vs 3 years, the distinction is somewhat academic, but the underlying exponential pace is what matters.

factualspeaker onlynovelty 1/4durability 3/4· Tracy Alloway

it's like, oh, I think like AI is going to destroy all white collar work in two years. And someone else is like, no, no, no, I think it's going to be three years is if that makes like any different whatsoever.

0.22

METR receives a modest amount of inbound interest from venture capital and investment firms, but the organization does not prioritize or deeply engage with the business and investment implications of their work, instead focusing on informing the public and stakeholders about AI capabilities.

factualspeaker onlynovelty 0/4durability 2/4· Chris Painter

I would say, you know, we don't get an enormous amount of inbound from investment firms. I mean, sometimes, you know, VCs or whatever we're based in the Bay area will reach out to us. I think that there's some kind of principle of our goal is to inform the public and give them the best evidence that we can about when we might get to this point of kind of, you know, AI being, you know, fully autonomous or able to improve itself.

0.17

METR prioritizes measuring capabilities on models anticipated to be on the frontier due to limited staff and resources, resulting in deliberate focus on US models and exclusion of Chinese models that are assessed to be behind the frontier.

factualspeaker onlynovelty 0/4durability 2/4· Joel Becker

So we do try to prioritize, just because meta has has limited resources, staff time in particular. The, the models that we, we anticipate being on the frontier

0.17

Many METR team members are people who made significant money in the AI industry and chose to leave high-paid positions to work on independent research questions that are informative to public decision-making, motivated by values rather than financial incentives.

factualspeaker onlynovelty 0/4durability 2/4· Chris Painter

one hope that I have is that more, you know, there will be more and more researchers who have kind of like, made the money that they need from working in the industry. And now we're excited and kind of like lifting all boats by working on kind of like inside of an organization where the North Star can be what is most informative to the rest of the world outside of these, like, you know, relatively small set of companies.