Feed-forward networks

The other half of every transformer block, applied to each token on its own, right after attention has finished mixing information across tokens.

01

Expand, then contract

input: d_model = 512 (shown here as 8)
FFN(x)=activation(xW1+b1) W2+b2\text{FFN}(x) = \text{activation}(xW_1 + b_1)\,W_2 + b_2

After attention mixes information across tokens, each token's vector is passed through a small two-layer network, the same weights reused at every position, but each token processed independently. Nothing mixes between tokens at this step, attention already did that part.

The first layer expands the vector to a wider hidden size. The original paper goes from 512 dimensions to 2048, then applies a non-linearity.

The second layer projects back down to the original size. This expand, bend, contract pattern is where most of a transformer's parameters, and most of its raw compute, actually live.

02

Why you need a non-linearity

ReLU: max(0, x)

Without a non-linearity in between, two linear layers stacked together collapse into one bigger linear layer, mathematically identical to skipping the expansion entirely. Expanding the width alone would add zero expressive power.

The original paper used ReLU, which zeroes out every negative value. Most modern LLMs use GELU instead, a smoothed version of the same idea. Either way, the bend is what lets the network represent curves and thresholds a straight line never could.

ReLU and GELU are two points in a much bigger design space. The activation functions page compares sigmoid, tanh, the whole ReLU family, and Swish side by side, and lets you play with each one.