Stop Words
Filtering out the words that carry the least signal, and the one case where that filtering can quietly break your model.
Not every token carries the same signal
Take the sentence “Hey buddy, I want to go to your house.” Once it's tokenized, a few of those words, I, to, your, do almost no work: swap them out and the sentence still means roughly the same thing. Others, buddy, go, house, carry most of the actual content.
Stop words are exactly this: a list of common, low-information words, articles, pronouns, prepositions, that show up constantly but rarely help a model tell one piece of text apart from another. Dropping them shrinks the vocabulary a model has to deal with, for free.
A list you can shape yourself
There's no single universal stop word list. NLTK ships a default English one, but nothing stops you from building your own, adding domain-specific filler words, or trimming the default list down.
The one word worth thinking twice about removing is not. “I did go to the house” and “I did not go to the house” are opposite sentences, but a naive stop word list strips not out of both, erasing the difference entirely. Whether stop word removal is even worth using depends on the task: fine for spam or toxicity classification, risky wherever negation or word order changes the meaning, like sentiment analysis, summarization, or translation.
Next: reducing what's left down to a root form, starting with the fast, rule-based way to do it. Stemming.