How transformers and attention work
Large language models are transformers. They split text into tokens, let every token look at the others through attention, and predict what comes next, one token at a time.
- 1
Split text into tokens
The model reads tokens, pieces of words with an ID each, not letters or whole sentences.
- 2
Give each token a vector
Each token ID is looked up in a table of learned vectors, its embedding. Position information is included so word order counts.
- 3
Attention: decide what to look at
For each token, attention scores every other token by relevance. Here “it” looks mostly at “cat”.
- 4
Many heads, many patterns
Attention runs in parallel heads, and each learns a different pattern: what a word refers to, the word just before, what is said about a subject.
- 5
Predict the next token
After many layers, the model gives a probability to every token in its vocabulary. One is picked, appended, and the process repeats.
Split text into tokens
The model reads tokens, pieces of words with an ID each, not letters or whole sentences.
Give each token a vector
Each token ID is looked up in a table of learned vectors, its embedding. Position information is included so word order counts.
Attention: decide what to look at
For each token, attention scores every other token by relevance. Here “it” looks mostly at “cat”.
Many heads, many patterns
Attention runs in parallel heads, and each learns a different pattern: what a word refers to, the word just before, what is said about a subject.
Predict the next token
After many layers, the model gives a probability to every token in its vocabulary. One is picked, appended, and the process repeats.
In short
- Everything a model reads or writes is tokens, which is why context limits and prices are counted in tokens.
- Attention compares every token with every other one, so cost grows quickly with context length.
- Generation is repeated next-token prediction. Temperature controls how often a less likely token gets picked.
- The token IDs and percentages in the animation are illustrative.