How raw text becomes something a model can compute on: tokenizing it, filtering it, and reducing it to a root form, before it's ever turned into a vector.
The first step in text preprocessing: turning a raw sentence into the individual words a later step can actually work with.
Filtering out the words that carry the least signal, and the one case where that filtering can quietly break your model.
A fast, rule-based way to chop a word down toward its root, even when the result isn't a real word.
Slower than stemming, and worth it whenever the output has to be a real, meaningful word.