Mitchell's research group tested abstract reasoning puzzles where humans are shown three demonstrations of a grid transformation and asked to identify which test grid follows the same pattern; humans achieved about 90% accuracy but GPT-4 only achieved about 33% accuracy on 480 such puzzles, suggesting GPT-4 lacks basic abstract reasoning abilities.
factualpending
Speaker
Melanie MitchellEvidence Quote
“on 480 of these little puzzles which are in our Corpus humans were about 90% 90% accurate gp4 only got about 33%”
Created: 8/12/2026, 6:19:21 PM
My Notes
Loading notes...