Corpus, Documents & Vocabulary

Four words used constantly once text needs to become vectors, worth pinning down before any of it makes sense.

01

A dataset, in NLP's own words

DocTextLabel
d1the food is good1
d2the food is bad0
d3pizza is amazing1
d4burger is bad0

Take a small sentiment analysis dataset: four short pieces of text, each labeled 1 for a positive sentiment or 0 for a negative one. Nothing unusual about it, it looks like any other labeled dataset.

But NLP has its own vocabulary (no pun intended) for the pieces of a dataset like this, and it's worth pinning down before looking at how any of it gets turned into numbers.

02

Corpus, document, vocabulary, word

Corpus
document d1

the food is good

document d2

the food is bad

document d3

pizza is amazing

document d4

burger is bad

vocabulary = 8 unique words

Stack every row together and the whole dataset becomes a corpus, one body of text. Each individual row inside it, “the food is good”, is a document, NLP's word for what would otherwise just be called a sentence or a data point.

Vocabulary is a count, not a row: it's the number of unique words that appear anywhere across the whole corpus. Here that's the, food, is, good, bad, pizza, amazing, burger, eight unique words even though there are only four documents. Each individual entry in that vocabulary, one single word, is just called a word.

Next: turning each document into a vector, starting with the simplest possible way, one-hot encoding.