Self-attention

A mechanism where every token in a sequence looks at every other token (including itself) and decides how much to weigh each one when building its own updated representation.

See it explained in full