/glossary

The terms, defined plainly.

29 terms that come up constantly in AI and LLMs, each one short enough to actually read, with a link to where it's explained in full.

Activation function

A function such as sigmoid, ReLU, or tanh applied to a neuron's weighted sum, introducing non-linearity so a network can represent more than a straight line.

AI agent

An LLM-driven loop that reasons about a task, chooses an action, executes it, observes the result, and decides what to do next, repeating until the task is done, instead of answering in one fixed step.

Backpropagation

The algorithm that computes how much each weight in a network contributed to the final error, by applying the chain rule backward from the output layer to the input layer.

Bag of words

Representing a document as a vector of word counts across a shared vocabulary, ignoring the order those words appeared in.

Corpus

An entire text dataset treated as one body of text, every document in it stacked together.

Cosine similarity

The cosine of the angle between two vectors, used to measure how similar two documents' word-count vectors are: 1 for pointing the same direction, 0 for unrelated.

Embedding

A learned, fixed-size vector of numbers that stands in for a token, positioned so that tokens with similar meaning end up with similar vectors.

Feed-forward network

A small two-layer neural network applied independently to each token's vector after attention, expanding it to a larger dimension, applying a non-linearity, then projecting it back down.

Forward propagation

Running input through a network's layers in order to produce an output, the normal direction data flows during prediction, as opposed to backpropagation, which runs error signals in reverse.

Gradient descent

The update rule that nudges each weight a small step in whichever direction would have reduced the error, repeated over many examples until the weights converge.

Layer normalization (LayerNorm)

A step that rescales each token's vector on its own, subtracting its mean and dividing by its standard deviation, so every token enters the next layer on the same predictable scale.

Lemmatization

Reducing a word to its dictionary base form (its lemma) by looking it up against real vocabulary, slower than stemming but always a real, meaningful word.

Logit

A raw, unnormalized score a model assigns to a candidate output before softmax turns it into a probability.

Loss function

A single number that measures how wrong a model's prediction was for one example, the quantity that training tries to minimize.

Multi-head attention

Running several smaller self-attention operations in parallel, each with its own learned projections, so the model can attend to different kinds of relationships (nearby words, subject-verb pairs, position) at once instead of averaging them into one.

N-gram

A feature made of n consecutive words treated as one unit, a bigram is two words, a trigram is three, capturing local word order that single-word counting throws away.

One-hot encoding

Representing a word as a vector the length of the whole vocabulary, all zeros except a single 1 at that word's position.

Positional encoding

A fixed pattern of sine and cosine waves added to each token's embedding so the model can tell tokens apart by position, since attention on its own treats a sequence as an unordered set.

RAG (retrieval-augmented generation)

Retrieving the most relevant chunks of a knowledge base for a given query and injecting them into the prompt, so the model answers from that context instead of from memory alone.

Recurrent neural network (RNN)

An older architecture that processes a sequence one token at a time, carrying a running hidden state forward, which makes it hard to parallelize and prone to forgetting distant tokens, the problem transformers were designed to solve.

Residual connection

A shortcut that adds a layer's input directly onto its output, so a deep stack of layers only has to learn a correction on top of an identity path instead of the whole transformation from scratch, which keeps gradients flowing during training.

Self-attention

A mechanism where every token in a sequence looks at every other token (including itself) and decides how much to weigh each one when building its own updated representation.

Softmax

A function that turns a list of raw scores into a probability distribution: every value becomes positive and the whole list sums to 1, with larger scores getting disproportionately more weight.

Stemming

Reducing a word to an approximate root form by mechanically stripping prefixes and suffixes with fixed rules, fast but not guaranteed to produce a real word.

Stop words

Common, low-information words, articles, pronouns, prepositions, that get filtered out of tokenized text because they rarely help a model tell one piece of text apart from another.

Tokenization

Splitting raw text into the smaller pieces, roughly word-sized or sub-word-sized, that a model actually operates on, the first step before any of those pieces become embeddings.

Tool calling

A model requesting that the surrounding framework run a specific function on its behalf, a search, a calculation, a database query, and feeding the result back into its context.

Vector database

A database built to store embeddings and search them by semantic similarity, so a query finds chunks with a similar meaning instead of just matching keywords.

Vocabulary

The count of unique words appearing anywhere across a corpus, not how many documents it has, how many distinct words show up in them.