Positional encoding

A fixed pattern of sine and cosine waves added to each token's embedding so the model can tell tokens apart by position, since attention on its own treats a sequence as an unordered set.

See it explained in full