29 terms that come up constantly in AI and LLMs, each one short enough to actually read, with a link to where it's explained in full.
A function such as sigmoid, ReLU, or tanh applied to a neuron's weighted sum, introducing non-linearity so a network can represent more than a straight line.
An LLM-driven loop that reasons about a task, chooses an action, executes it, observes the result, and decides what to do next, repeating until the task is done, instead of answering in one fixed step.
The algorithm that computes how much each weight in a network contributed to the final error, by applying the chain rule backward from the output layer to the input layer.
Representing a document as a vector of word counts across a shared vocabulary, ignoring the order those words appeared in.
An entire text dataset treated as one body of text, every document in it stacked together.
The cosine of the angle between two vectors, used to measure how similar two documents' word-count vectors are: 1 for pointing the same direction, 0 for unrelated.
A learned, fixed-size vector of numbers that stands in for a token, positioned so that tokens with similar meaning end up with similar vectors.
A small two-layer neural network applied independently to each token's vector after attention, expanding it to a larger dimension, applying a non-linearity, then projecting it back down.
Running input through a network's layers in order to produce an output, the normal direction data flows during prediction, as opposed to backpropagation, which runs error signals in reverse.
The update rule that nudges each weight a small step in whichever direction would have reduced the error, repeated over many examples until the weights converge.
A step that rescales each token's vector on its own, subtracting its mean and dividing by its standard deviation, so every token enters the next layer on the same predictable scale.
Reducing a word to its dictionary base form (its lemma) by looking it up against real vocabulary, slower than stemming but always a real, meaningful word.
A raw, unnormalized score a model assigns to a candidate output before softmax turns it into a probability.
A single number that measures how wrong a model's prediction was for one example, the quantity that training tries to minimize.
Running several smaller self-attention operations in parallel, each with its own learned projections, so the model can attend to different kinds of relationships (nearby words, subject-verb pairs, position) at once instead of averaging them into one.
A feature made of n consecutive words treated as one unit, a bigram is two words, a trigram is three, capturing local word order that single-word counting throws away.
Representing a word as a vector the length of the whole vocabulary, all zeros except a single 1 at that word's position.
A fixed pattern of sine and cosine waves added to each token's embedding so the model can tell tokens apart by position, since attention on its own treats a sequence as an unordered set.
Retrieving the most relevant chunks of a knowledge base for a given query and injecting them into the prompt, so the model answers from that context instead of from memory alone.
An older architecture that processes a sequence one token at a time, carrying a running hidden state forward, which makes it hard to parallelize and prone to forgetting distant tokens, the problem transformers were designed to solve.
A shortcut that adds a layer's input directly onto its output, so a deep stack of layers only has to learn a correction on top of an identity path instead of the whole transformation from scratch, which keeps gradients flowing during training.
A mechanism where every token in a sequence looks at every other token (including itself) and decides how much to weigh each one when building its own updated representation.
A function that turns a list of raw scores into a probability distribution: every value becomes positive and the whole list sums to 1, with larger scores getting disproportionately more weight.
Reducing a word to an approximate root form by mechanically stripping prefixes and suffixes with fixed rules, fast but not guaranteed to produce a real word.
Common, low-information words, articles, pronouns, prepositions, that get filtered out of tokenized text because they rarely help a model tell one piece of text apart from another.
Splitting raw text into the smaller pieces, roughly word-sized or sub-word-sized, that a model actually operates on, the first step before any of those pieces become embeddings.
A model requesting that the surrounding framework run a specific function on its behalf, a search, a calculation, a database query, and feeding the result back into its context.
A database built to store embeddings and search them by semantic similarity, so a query finds chunks with a similar meaning instead of just matching keywords.
The count of unique words appearing anywhere across a corpus, not how many documents it has, how many distinct words show up in them.