Juan Pablo García
All writing
Explainer5 steps · scroll to play

How transformers and attention work

Large language models are transformers. They split text into tokens, let every token look at the others through attention, and predict what comes next, one token at a time.

  1. 1

    Split text into tokens

    The model reads tokens, pieces of words with an ID each, not letters or whole sentences.

  2. 2

    Give each token a vector

    Each token ID is looked up in a table of learned vectors, its embedding. Position information is included so word order counts.

  3. 3

    Attention: decide what to look at

    For each token, attention scores every other token by relevance. Here “it” looks mostly at “cat”.

  4. 4

    Many heads, many patterns

    Attention runs in parallel heads, and each learns a different pattern: what a word refers to, the word just before, what is said about a subject.

  5. 5

    Predict the next token

    After many layers, the model gives a probability to every token in its vocabulary. One is picked, appended, and the process repeats.

Split text into tokens

The model reads tokens, pieces of words with an ID each, not letters or whole sentences.

Give each token a vector

Each token ID is looked up in a table of learned vectors, its embedding. Position information is included so word order counts.

Attention: decide what to look at

For each token, attention scores every other token by relevance. Here “it” looks mostly at “cat”.

Many heads, many patterns

Attention runs in parallel heads, and each learns a different pattern: what a word refers to, the word just before, what is said about a subject.

Predict the next token

After many layers, the model gives a probability to every token in its vocabulary. One is picked, appended, and the process repeats.

In short

  • Everything a model reads or writes is tokens, which is why context limits and prices are counted in tokens.
  • Attention compares every token with every other one, so cost grows quickly with context length.
  • Generation is repeated next-token prediction. Temperature controls how often a less likely token gets picked.
  • The token IDs and percentages in the animation are illustrative.

Have an AI feature to build? Let's talk for 15 minutes.