Mitchell wrote an article called 'How Do We Know How Smart AI Systems Are' highlighting that tests of AI capabilities often don't reveal limitations because: (1) test questions may have been in the training data, (2) the vast training data makes it unclear what was memorized vs. learned, and (3) systems may use narrow, non-transferable procedures rather than genuine understanding.
factualpending
Speaker
Melanie MitchellEvidence Quote
“all these things that we test them on often don't really reveal a lot of their limitations so you give them the bar exam but maybe even if it did well some of those questions were in it training data”
Created: 8/12/2026, 6:19:21 PM
My Notes
Loading notes...