Lemmatization
Slower than stemming, and worth it whenever the output has to be a real, meaningful word.
Looking a word up instead of chopping it
Lemmatization reduces a word to its lemma, its dictionary base form, the same job stemming does, but by looking the word up against a real vocabulary (NLTK's WordNetLemmatizer, for example) instead of applying blunt suffix rules.
Run the same words through it: historical and history both resolve to history, a real word, and final, finally, and finalized all resolve to final. Same grouping stemming produced, but every output is something you could actually read.
Stemming vs. lemmatization, by use case
| Stemming | Lemmatization | |
|---|---|---|
| Method | Rule-based suffix stripping | Dictionary lookup (e.g. WordNet) |
| Speed | Fast | Slower |
| Output | Approximate, sometimes not a real word | Always a real, meaningful word |
| Good fit | Spam classification, review-star prediction | Text summarization, translation, chatbots |
That dictionary lookup is why lemmatization is slower than stemming, it has to check candidates against real vocabulary instead of just applying rules, but it's the only one of the two whose output is safe to show a human, or to hand to a downstream step that depends on word meaning rather than just word grouping.
The two aren't competing for the same job. Reach for stemming when speed matters and only the grouping counts, spam or review-star classification. Reach for lemmatization when the actual words matter, text summarization, machine translation, or a chatbot's replies.
That's tokenization, stop words, stemming, and lemmatization, the text-preprocessing basics from lecture 1. Back to NLP for what comes next.