Large language models will likely get much better when they're multimodal—trained on images as well as words. GPT-4 was trained with images, and it's possible Google is doing the same. When multimodal, these models could learn much more than humans.
forecastpending
Speaker
Geoffrey HintonEvidence Quote
“particularly when they're multimodal they could learn much much more than us”
Created: 8/12/2026, 5:56:09 PM
My Notes
Loading notes...