Transformers

Attention, multi-head attention, and the architecture behind every modern large language model.

01Live

Transformers

A visual, step-by-step walkthrough of how a transformer turns text into a prediction, from tokens to attention to the next word.

02Coming soon

Attention Is All You Need, the full paper

A section-by-section dissection of the 2017 paper: the full encoder-decoder architecture, cross-attention, training regime…