Multi-head attention

Running several smaller self-attention operations in parallel, each with its own learned projections, so the model can attend to different kinds of relationships (nearby words, subject-verb pairs, position) at once instead of averaging them into one.

See it explained in full