Mitchell's research group tested abstract reasoning puzzles where humans are shown three demonstrations of a grid transformation and asked to identify which test grid follows the same pattern; humans achieved about 90% accuracy but GPT-4 only achieved about 33% accuracy on 480 such puzzles, suggesting GPT-4 lacks basic abstract reasoning abilities.

factualpending

Speaker

Melanie Mitchell

Evidence Quote

on 480 of these little puzzles which are in our Corpus humans were about 90% 90% accurate gp4 only got about 33%

Source

The Future of Artificial IntelligenceSanta Fe Institute
Created: 8/12/2026, 6:19:21 PM

My Notes

Loading notes...