Tokenization

Splitting raw text into the smaller pieces, roughly word-sized or sub-word-sized, that a model actually operates on, the first step before any of those pieces become embeddings.

See it explained in full