Claude Sonnet 4.5, after specific training on the FinBench Agent benchmark, outperforms Claude Opus 4.1 by five percentage points, achieving 55% accuracy on the benchmark.
factualpending
Speaker
Nick LynnEvidence Quote
“sonnet 45 outperforms our opus 4.1 model which previously topped the charts by five full percentage points. Yeah. So I think it's 55%”
Created: 8/11/2026, 7:09:15 AM
My Notes
Loading notes...