One-Hot Encoding
The simplest way to turn a word into a vector: a single 1 in exactly one position, 0 everywhere else, and almost every problem that causes downstream.
Every word becomes a 1 in one position
Take a tiny corpus: “a man eat food”, “cat eat food”, “people watch krish youtube”. Its vocabulary is nine unique words: a, man, eat, food, cat, people, watch, krish, youtube.
One-hot encoding gives every word a vector as long as the whole vocabulary, all zeros except a single 1 at that word's position. “man” is the second word in the vocabulary, so its vector is nine slots long with a 1 in slot two and 0 everywhere else.
Vocabulary size becomes vector size
That vector length is exactly the vocabulary size, nine here, but a real corpus can easily have tens of thousands of unique words. Every one of those words still gets a vector that long, almost entirely zeros. Multiply that across every word in every document and the result is a huge, sparse matrix: expensive to store and slow for a model to train on.
There's a second problem hiding in that same fixed vocabulary: a word that never showed up while it was being built has no column to occupy. Show the model “dog” at test time and there's simply nowhere to put its 1, it can't be represented at all.
No sense of meaning between words
One-hot vectors also carry no notion of closeness. “man” and “eat” appear right next to each other in the same sentence, but their vectors are just as different, one non-overlapping 1 each, as “man” and “krish”, two words that never appear together at all.
Every pair of words is equally far apart, because the representation only encodes which word it is, never anything about how it's used. Whatever comes next has to fix that, or nothing built on top of it can learn which words tend to relate.
Next: Bag of Words fixes the fixed-vocabulary idea into something trainable, by counting instead of just marking a position.