Loss functions
The single number training is always trying to shrink, and why regression and classification each need a different kind of number entirely.
Loss vs. cost
loss(example_1), loss(example_2), ..., loss(example_32)
cost = average of all 32 -> one weight update
A loss function measures how wrong a single prediction was. Training rarely works one example at a time, though: it passes a whole batch through the network before updating anything, since a single unusual example would otherwise swing every weight update.
The per-example loss becomes a per-batch cost: the average loss across every example in that batch. Cost is just loss, averaged, before backpropagation runs. A batch size of 32 or 64 is a common compromise, large enough to smooth out noisy examples, small enough to keep each step fast.
Regression losses
Mean Squared Error (MSE) is the default for regression: half the squared difference between the true value and the prediction. Squaring makes it a quadratic, a smooth bowl shape that's especially easy for gradient descent to descend (gradient descent works on far messier, non-convex surfaces too, this is just the friendliest case), but it also means one large outlier contributes disproportionately to the loss.
Mean Absolute Error (MAE) uses the absolute difference instead. It grows linearly, so a large outlier contributes proportionally to its size rather than its square, at the cost of a sharp corner at zero error where the slope is undefined.
Huber loss is a deliberate compromise: quadratic like MSE for small errors, linear like MAE for large ones, switching over at a threshold. Drag the slider and watch MSE curve upward fastest while MAE grows in a straight line.
Classification losses
predicted probability = 0.80, loss = 0.22
Classification losses measure something different: how well a predicted probability matches a true category. Binary Cross-Entropy (BCE) is used whenever the output is a single sigmoid probability, for a yes-or-no prediction.
The shape below is the entire point of using a logarithm instead of a simple difference. When the true label is 1 and the model predicts a probability close to 1, the loss is nearly 0. But as that probability drifts toward 0, the loss doesn't grow gently, it shoots toward infinity: a confident, wrong prediction is punished far more than a hesitant one.
Matching losses to tasks
Categorical Cross-Entropy (CCE) is the multi-class generalization of BCE, paired with a softmax output layer. Because softmax gives a probability for every class and exactly one is correct, the sum in the real formula collapses to a single term: the negative log of the probability assigned to the correct class.
Say a 3-class softmax outputs 0.70, 0.20, 0.10, and the true class is the first one. CCE only looks at the probability the model gave the right answer, everything below is the same -log(p) curve as BCE, just evaluated at whichever class was actually correct.
The output activation and the loss function are always chosen together, as a matched pair, driven entirely by the task, not independently.
See also loss function in the glossary, and how a computed loss actually updates weights on the backpropagation page.