Activation functions
Every option in the feed-forward block's "bend," compared side by side, and playable: drag the slider on any section below and watch every curve move at once.
Why you need a non-linearity
Without a non-linearity in between, two linear layers stacked together collapse into one bigger linear layer, mathematically identical to skipping the extra layer entirely, no matter how many you stack.
An activation function decides whether, and how strongly, a neuron fires. That bend is what lets a network represent curves and thresholds a straight line never could. Drag the slider below and watch every function bend differently around the same input.
Sigmoid and tanh: smooth, but they saturate
Sigmoid squashes any real number into 0 to 1, which reads naturally as a probability. Its derivative peaks at exactly 0.25 when x is 0, and tapers toward 0 at either extreme, that flattening is called saturation.
Tanh has the same S-shape, rescaled to -1 to 1 instead. Being centered on zero helps the next layer train slightly better, but tanh saturates too, for the same reason: both functions flatten out for large positive or negative inputs.
The ReLU family
ReLU passes positive inputs through unchanged and zeroes out everything negative. Its derivative is exactly 1 for any positive input, not a fraction, so it does not shrink no matter how many layers it passes back through. It's also far cheaper to compute: no exponentials, just a comparison against zero.
The tradeoff is dying ReLU: once a neuron's input goes negative, its gradient becomes exactly 0 and stays there. Leaky ReLU and ELU both fix this by letting a small negative signal through instead of a hard zero, at a small cost in computation.
Swish and GELU
Swish multiplies the input by its own sigmoid, self-gating: sigmoid(x) acts as a gate between 0 and 1 that decides how much of x to let through, using the input itself to control the gate rather than a separate signal.
Unlike ReLU, Swish is smooth everywhere and dips slightly below zero just left of the origin instead of flattening completely, which keeps a small gradient alive for mildly negative inputs. It performs competitively with ReLU on deep networks, at extra computational cost.
GELU (Gaussian Error Linear Unit) is built the same way, weighting the input by roughly how likely it is to be "kept" under a standard normal distribution, rather than by its own sigmoid. It behaves almost identically to Swish, smooth, non-monotonic, a small negative dip, and it's the activation most modern LLMs actually use inside their feed-forward blocks, including the one on this site's transformer page.
Choosing one
Hidden layers default to ReLU: cheap, no vanishing gradient for positive inputs, and it works well in most architectures. If dead neurons show up in practice, switch to Leaky ReLU, PReLU, or ELU, in roughly that order of complexity.
The output layer is a different decision, driven entirely by the task: sigmoid for binary classification, softmax for multi-class classification, and no activation at all (a linear output) for regression, since the prediction needs to be an unrestricted real number.
Comparison
| Function | Output range | Strength | Weakness |
|---|---|---|---|
| Sigmoid | 0 to 1 | Smooth; clean probability-like output | Vanishing gradient; not zero-centered; slow |
| Tanh | -1 to 1 | Zero-centered, unlike sigmoid | Still vanishes for large inputs |
| ReLU | 0 to infinity | Very fast; no vanishing gradient for positive inputs | Dead neurons: negative inputs get a zero gradient |
| Leaky ReLU / PReLU | -infinity to infinity | Fixes dead neurons with a small negative slope | Extra hyperparameter (or learned parameter) |
| ELU | -alpha to infinity | Smooth negative side; closer to zero-centered | More expensive to compute than ReLU |
| Swish | ≈ -0.28 to infinity | Smooth everywhere; competitive with ReLU | More expensive to compute (uses sigmoid) |
| GELU | ≈ -0.17 to infinity | Smooth, probabilistic weighting; the default in most modern LLMs | More expensive to compute than ReLU |
| Softmax | 0 to 1, sums to 1 | Turns raw scores into class probabilities | Output layer only, multi-class problems |
Softmax: a special case
Softmax isn't plotted above because it doesn't act on one number at a time, it takes every output score at once and turns them into a probability distribution. Given raw scores of 2.0, 1.0, 0.1, and -1.0 across four classes, softmax exponentiates each one and divides by the total:
The highest score doesn't just win outright, it gets amplified into the largest probability, but every class keeps some nonzero share. Softmax is effectively sigmoid generalized to more than two classes. It shows up at the output layer of a multi-class classifier, and anywhere else a set of scores needs to become a probability distribution, including inside every attention head of a transformer, turning attention scores into weights that sum to 1.