Mitchell wrote an article called 'How Do We Know How Smart AI Systems Are' highlighting that tests of AI capabilities often don't reveal limitations because: (1) test questions may have been in the training data, (2) the vast training data makes it unclear what was memorized vs. learned, and (3) systems may use narrow, non-transferable procedures rather than genuine understanding.

factualpending

Speaker

Melanie Mitchell

Evidence Quote

all these things that we test them on often don't really reveal a lot of their limitations so you give them the bar exam but maybe even if it did well some of those questions were in it training data

Source

The Future of Artificial IntelligenceSanta Fe Institute
Created: 8/12/2026, 6:19:21 PM

My Notes

Loading notes...