Recurrent neural networks

Before transformers, this is how models processed sequences, one token at a time, and it's why a fully parallel alternative was such a big deal.

01

One token at a time

Thestep 1
h1->
catstep 2
h2->
satstep 3
h3->
matstep 4
ht=f(Whht−1+Wxxt)h_t = f(W_h h_{t-1} + W_x x_t)

A recurrent neural network (RNN) reads a sequence one token at a time. At each step it combines the new token with a hidden state, a running summary vector carried over from the previous step.

That hidden state is the network's only memory of everything it has seen so far. Whatever doesn't fit into that one fixed-size vector is effectively forgotten.

To produce the hidden state after token 3, the network first needs the hidden state after token 2, which needs the one after token 1. Processing is inherently sequential, step by step, with no shortcut.

02

Vanishing gradients

last step
-1
-2
-3
-4
-5

gradient magnitude reaching each earlier step

∂L∂h1=∂L∂hT⋅∂hT∂hT−1⋯∂h2∂h1\frac{\partial L}{\partial h_1} = \frac{\partial L}{\partial h_T} \cdot \frac{\partial h_T}{\partial h_{T-1}} \cdots \frac{\partial h_2}{\partial h_1}

Training an RNN means backpropagating the error at the last step all the way back through every earlier step, multiplying by a derivative at every step along the way. Those per-step derivatives aren't literally identical, they depend on the hidden-state values at each point in the sequence, but they tend to stay in a similar range, so the intuition below (repeated multiplication by roughly the same size number) is a fair approximation of what happens.

Multiply a number smaller than 1 by itself dozens of times and it shrinks toward zero. That's the vanishing gradient problem: the training signal from distant tokens barely reaches the early weights, so the network struggles to learn long-range dependencies.

Variants like LSTMs and GRUs added gating mechanisms specifically to fight this, and they helped a lot in practice, but the sequential, one-step-at-a-time bottleneck itself remained.

03

Why transformers replaced them

RNN: step 1 -> step 2 -> step 3 -> step 4 (sequential)

Transformer: step 1, step 2, step 3, step 4 (all at once)

Self-attention lets every token look directly at every other token in a single step, regardless of how far apart they are. No chain of hidden states in between, no information bottleneck.

Because there's no step-by-step dependency, every token's attention can be computed in parallel on a GPU. That parallelism was a major enabler of today's large models, alongside the compute, data scale, and training-recipe advances that had to grow alongside it.

The tradeoff: attention over a sequence of length n costs O(n^2), since every token compares against every other token, versus O(n) for an RNN, which only ever compares a token against the single running hidden state before it. For very long sequences that quadratic cost adds up. Both approaches have limits, just different ones.