Common, low-information words, articles, pronouns, prepositions, that get filtered out of tokenized text because they rarely help a model tell one piece of text apart from another.
More terms