RNNs suffer from short-term memory because they have a shorter window to reference from compared to GRUs and LSTMs, which have greater capacity to achieve longer-term memory with a longer reference window.
causalpending
Speaker
Unidentified Speaker — Illustrated Guide to Transformers Neural Network: A step by… [4Bdc55j80l8]Evidence Quote
“rnns have a shorter window to reference from so when a story gets longer rnns can't access word generated earlier in the sequence this is still true for gr use and L STM's although they do have a bigger capacity to achieve longer term memory”
Created: 8/13/2026, 9:51:15 AM
My Notes
Loading notes...