Large language models will likely get much better when they're multimodal—trained on images as well as words. GPT-4 was trained with images, and it's possible Google is doing the same. When multimodal, these models could learn much more than humans.

forecastpending

Speaker

Geoffrey Hinton

Evidence Quote

particularly when they're multimodal they could learn much much more than us

Source

Geoffrey Hinton - Two Paths to IntelligenceCSER Cambridge
Created: 8/12/2026, 5:56:09 PM

My Notes

Loading notes...