Activation functions

Every option in the feed-forward block's "bend," compared side by side, and playable: drag the slider on any section below and watch every curve move at once.

01

Why you need a non-linearity

x = 0.0Sigmoid: 0.500Tanh: 0.000ReLU: 0.000Leaky ReLU: 0.000ELU: 0.000Swish: 0.000GELU: 0.000

Without a non-linearity in between, two linear layers stacked together collapse into one bigger linear layer, mathematically identical to skipping the extra layer entirely, no matter how many you stack.

An activation function decides whether, and how strongly, a neuron fires. That bend is what lets a network represent curves and thresholds a straight line never could. Drag the slider below and watch every function bend differently around the same input.

02

Sigmoid and tanh: smooth, but they saturate

x = 0.0Sigmoid: 0.500Tanh: 0.000

Sigmoid squashes any real number into 0 to 1, which reads naturally as a probability. Its derivative peaks at exactly 0.25 when x is 0, and tapers toward 0 at either extreme, that flattening is called saturation.

Tanh has the same S-shape, rescaled to -1 to 1 instead. Being centered on zero helps the next layer train slightly better, but tanh saturates too, for the same reason: both functions flatten out for large positive or negative inputs.

03

The ReLU family

x = 0.0ReLU: 0.000Leaky ReLU: 0.000ELU: 0.000

ReLU passes positive inputs through unchanged and zeroes out everything negative. Its derivative is exactly 1 for any positive input, not a fraction, so it does not shrink no matter how many layers it passes back through. It's also far cheaper to compute: no exponentials, just a comparison against zero.

The tradeoff is dying ReLU: once a neuron's input goes negative, its gradient becomes exactly 0 and stays there. Leaky ReLU and ELU both fix this by letting a small negative signal through instead of a hard zero, at a small cost in computation.

04

Swish and GELU

x = 0.0Swish: 0.000

Swish multiplies the input by its own sigmoid, self-gating: sigmoid(x) acts as a gate between 0 and 1 that decides how much of x to let through, using the input itself to control the gate rather than a separate signal.

Unlike ReLU, Swish is smooth everywhere and dips slightly below zero just left of the origin instead of flattening completely, which keeps a small gradient alive for mildly negative inputs. It performs competitively with ReLU on deep networks, at extra computational cost.

GELU (Gaussian Error Linear Unit) is built the same way, weighting the input by roughly how likely it is to be "kept" under a standard normal distribution, rather than by its own sigmoid. It behaves almost identically to Swish, smooth, non-monotonic, a small negative dip, and it's the activation most modern LLMs actually use inside their feed-forward blocks, including the one on this site's transformer page.

05

Choosing one

x = 0.0Sigmoid: 0.500Tanh: 0.000ReLU: 0.000Leaky ReLU: 0.000ELU: 0.000Swish: 0.000GELU: 0.000

Hidden layers default to ReLU: cheap, no vanishing gradient for positive inputs, and it works well in most architectures. If dead neurons show up in practice, switch to Leaky ReLU, PReLU, or ELU, in roughly that order of complexity.

The output layer is a different decision, driven entirely by the task: sigmoid for binary classification, softmax for multi-class classification, and no activation at all (a linear output) for regression, since the prediction needs to be an unrestricted real number.

Comparison

FunctionOutput rangeStrengthWeakness
Sigmoid0 to 1Smooth; clean probability-like outputVanishing gradient; not zero-centered; slow
Tanh-1 to 1Zero-centered, unlike sigmoidStill vanishes for large inputs
ReLU0 to infinityVery fast; no vanishing gradient for positive inputsDead neurons: negative inputs get a zero gradient
Leaky ReLU / PReLU-infinity to infinityFixes dead neurons with a small negative slopeExtra hyperparameter (or learned parameter)
ELU-alpha to infinitySmooth negative side; closer to zero-centeredMore expensive to compute than ReLU
Swish≈ -0.28 to infinitySmooth everywhere; competitive with ReLUMore expensive to compute (uses sigmoid)
GELU≈ -0.17 to infinitySmooth, probabilistic weighting; the default in most modern LLMsMore expensive to compute than ReLU
Softmax0 to 1, sums to 1Turns raw scores into class probabilitiesOutput layer only, multi-class problems

Softmax: a special case

Softmax isn't plotted above because it doesn't act on one number at a time, it takes every output score at once and turns them into a probability distribution. Given raw scores of 2.0, 1.0, 0.1, and -1.0 across four classes, softmax exponentiates each one and divides by the total:

2.0 -> 0.64
1.0 -> 0.23
0.1 -> 0.10
-1.0 -> 0.03

The highest score doesn't just win outright, it gets amplified into the largest probability, but every class keeps some nonzero share. Softmax is effectively sigmoid generalized to more than two classes. It shows up at the output layer of a multi-class classifier, and anywhere else a set of scores needs to become a probability distribution, including inside every attention head of a transformer, turning attention scores into weights that sum to 1.