Benchmarks saturating rapidly on closed-ended tasks (GPQA, Humanity's Last Exam reaching ~25% accuracy) suggests models are approaching human-level performance on extremely difficult specialized tasks, but this says little about automating real-world messy labor that lacks clear verification criteria.

causalpending

Speaker

Sarah Hastings Woodhouse

Evidence Quote

benchmarks seem to be saturating very quickly on a lot of sort of closedended academic tasks

Source

AI Timelines and Human Psychology (with Sarah Hastings-Woodhouse)Future of Life Institute
Created: 8/11/2026, 7:16:09 AM

My Notes

Loading notes...