Benchmarks saturating rapidly on closed-ended tasks (GPQA, Humanity's Last Exam reaching ~25% accuracy) suggests models are approaching human-level performance on extremely difficult specialized tasks, but this says little about automating real-world messy labor that lacks clear verification criteria.
causalpending
Speaker
Sarah Hastings WoodhouseEvidence Quote
“benchmarks seem to be saturating very quickly on a lot of sort of closedended academic tasks”
Created: 8/11/2026, 7:16:09 AM
My Notes
Loading notes...