Tokenization
The first step in text preprocessing: turning a raw sentence into the individual words a later step can actually work with.
Raw text isn't computable
You won one million dollar
Take a simple spam classifier. Its features are an email's subject and body; its label is spam or ham. But a sentence like “You won one million dollar” isn't something an algorithm can act on directly, there's no computation defined on a string of English words.
Before any of the later preprocessing steps, filtering out low-signal words, reducing words to a root form, eventually turning words into numbers, can happen, the sentence first has to be broken into pieces small enough to handle one at a time. That first cut is called tokenization.
Splitting a sentence into tokens
Tokenization just means splitting a sentence into words. “You won one million dollar” becomes five separate tokens: You, won, one, million, dollar. Nothing is removed or rewritten yet, the sentence is only broken apart.
Every step that follows, filtering, stemming or lemmatizing, eventually vectorizing, operates on these tokens one at a time. Get the split wrong and every later step inherits the mistake.
Next: not every token is worth keeping. Stop words covers filtering the low-signal ones out.