Corpus, Documents & Vocabulary
Four words used constantly once text needs to become vectors, worth pinning down before any of it makes sense.
A dataset, in NLP's own words
| Doc | Text | Label |
|---|---|---|
| d1 | the food is good | 1 |
| d2 | the food is bad | 0 |
| d3 | pizza is amazing | 1 |
| d4 | burger is bad | 0 |
Take a small sentiment analysis dataset: four short pieces of text, each labeled 1 for a positive sentiment or 0 for a negative one. Nothing unusual about it, it looks like any other labeled dataset.
But NLP has its own vocabulary (no pun intended) for the pieces of a dataset like this, and it's worth pinning down before looking at how any of it gets turned into numbers.
Corpus, document, vocabulary, word
the food is good
the food is bad
pizza is amazing
burger is bad
Stack every row together and the whole dataset becomes a corpus, one body of text. Each individual row inside it, “the food is good”, is a document, NLP's word for what would otherwise just be called a sentence or a data point.
Vocabulary is a count, not a row: it's the number of unique words that appear anywhere across the whole corpus. Here that's the, food, is, good, bad, pizza, amazing, burger, eight unique words even though there are only four documents. Each individual entry in that vocabulary, one single word, is just called a word.
Next: turning each document into a vector, starting with the simplest possible way, one-hot encoding.