Written text data for training AI systems is approaching saturation, with possibly only 1-10% of reasonably available high-quality written data remaining, and the highest-quality data has already been used with lower-quality data left, potentially putting the field close to a significant bottleneck.
factualpending
Speaker
Yoshua BengioEvidence Quote
“the amount of uh written data is something that we approaching the limit of um maybe we are I don't know the numbers because these are hidden as well but I imagine we are at a few percent maybe you know 10%”
Created: 8/12/2026, 6:45:56 PM
My Notes
Loading notes...