Basics of deep learning
Before the Perceptron, before backpropagation, before any specific architecture: what actually makes a model "deep," and why that shape works at all.
AI, ML, DL, and where this fits
Data Science cuts across all three, it isn't nested inside them
Artificial Intelligence is the broadest term: an application that performs a task on its own, without a human specifying every step, by learning from data instead. Machine Learning is a subset of that, statistical tools and algorithms for prediction, forecasting, and clustering.
Deep Learning is a subset of Machine Learning, specifically the multi-layered neural networks this whole domain covers. Neural networks trace back to the single-neuron Perceptron in 1958, but "deep" (many-layer) learning didn't become practical until backpropagation matured in the 1980s, and didn't become mainstream until the last decade or so, driven by two things arriving together: an explosion of available data, and GPU hardware that makes training large networks affordable.
Data Science doesn't nest neatly inside this picture. A data scientist's work touches AI, ML, or DL depending on the task, sometimes it's building a predictive model, sometimes training a deep network, sometimes just cleaning and analyzing data. The tool changes; the goal, shipping something useful from data, doesn't.
What deep learning actually is
A deep learning model is a stack of simple mathematical functions, layers, each one transforming its input a little, with the output of one layer feeding the next.
Every one of those transformations has adjustable numbers, weights. Learning means searching for weight values that make the whole stack produce the right output for the training examples.
Nothing in the model is hand-coded rules. What it does is entirely a product of the data it saw and the weights that data produced.
Why stack layers at all
every arrow between layers carries its own adjustable weight
A single layer can only represent fairly simple relationships, a straight-line boundary in the simplest case, as the next page on the Perceptron shows directly.
Stacking layers, each with a non-linearity in between, lets the model build up increasingly complex functions from these simple pieces. That's the core idea behind the field's name: it's deep because there are many layers between input and output.
One sufficiently wide single layer can, in principle, approximate almost anything. In practice, going deep instead of just wide turns out to be a far more efficient way to represent complicated functions.
The training loop, at a glance
<- measure the error, adjust every weight a little, run it again ->
Every layer's weights start out random. Training repeatedly runs examples through the network, checks how wrong its answer was, and adjusts every weight a little to make that answer less wrong next time.
That loop, run the network forward, measure the error, adjust the weights, has a name for each half: forward propagation and backpropagation. A dedicated page covers exactly how those work.
Repeat that loop over enough examples, and the random numbers this page started with become weights that solve a real problem: recognizing images, predicting the next word, or almost anything else.