The Perceptron

The single-neuron model that started it all: a weighted sum, a threshold, and a straight line through your data.

01

A single artificial neuron

y=step(w1x1+w2x2+⋯+b)y = \text{step}(w_1 x_1 + w_2 x_2 + \cdots + b)

The Perceptron, introduced by Frank Rosenblatt in 1958, is the simplest possible model of a neuron: take a set of inputs, multiply each by a weight, add them up, and add one more number called a bias.

Then check whether that sum clears a threshold. If it does, the perceptron fires and outputs 1. If not, it stays silent and outputs 0. That's the entire model: a weighted vote, thresholded.

Why bias matters: without it, the weighted sum is always 0 whenever every input is 0, which forces the decision boundary to pass through the origin. The bias shifts that line freely, so it can separate classes that don't happen to straddle the origin, a real geometric limitation, not just a training nicety.

02

A straight line through your data

w1=0.94, w2=0.34, b=-156.47/8 correctly classified

Geometrically, a perceptron's weights and bias define a straight line, or with more inputs, a flat plane, that cuts the space of possible inputs in two. Everything on one side fires; everything on the other doesn't.

Learning means adjusting the weights and bias so that line ends up in the right place. The perceptron learning rule nudges it a little every time it misclassifies an example; correct answers leave it alone.

Try it yourself: drag the two sliders below until every point is classified correctly. That's exactly what training a perceptron is automating, searching for a weight angle and bias that separate the classes.

03

Where it breaks

no straight line separates these

A single perceptron can only separate data that's linearly separable, points some straight line can actually divide. Anything more tangled, most famously the XOR pattern, is mathematically impossible for one perceptron to solve, no matter how it's trained.

That limitation, formalized by Minsky and Papert in 1969, is exactly what stacking multiple layers fixes: each layer draws its own line, and combining them can carve out far more complex regions. That combination is a feed-forward network, covered next.