RNNs suffer from short-term memory because they have a shorter window to reference from compared to GRUs and LSTMs, which have greater capacity to achieve longer-term memory with a longer reference window.

causalpending

Speaker

Unidentified Speaker — Illustrated Guide to Transformers Neural Network: A step by… [4Bdc55j80l8]

Evidence Quote

rnns have a shorter window to reference from so when a story gets longer rnns can't access word generated earlier in the sequence this is still true for gr use and L STM's although they do have a bigger capacity to achieve longer term memory

Source

Illustrated Guide to Transformers Neural Network: A step by step explanationThe AI Hacker
Created: 8/13/2026, 9:51:15 AM

My Notes

Loading notes...