Tokenization

The first step in text preprocessing: turning a raw sentence into the individual words a later step can actually work with.

01

Raw text isn't computable

You won one million dollar

Take a simple spam classifier. Its features are an email's subject and body; its label is spam or ham. But a sentence like “You won one million dollar” isn't something an algorithm can act on directly, there's no computation defined on a string of English words.

Before any of the later preprocessing steps, filtering out low-signal words, reducing words to a root form, eventually turning words into numbers, can happen, the sentence first has to be broken into pieces small enough to handle one at a time. That first cut is called tokenization.

02

Splitting a sentence into tokens

You
won
one
million
dollar

Tokenization just means splitting a sentence into words. “You won one million dollar” becomes five separate tokens: You, won, one, million, dollar. Nothing is removed or rewritten yet, the sentence is only broken apart.

Every step that follows, filtering, stemming or lemmatizing, eventually vectorizing, operates on these tokens one at a time. Get the split wrong and every later step inherits the mistake.

Next: not every token is worth keeping. Stop words covers filtering the low-signal ones out.