小王算钱

Interview Prep · ML Infra

Model architectures:
from N-gram to Transformer

First a timeline showing what each architecture fixed in its predecessor, then the Transformer taken apart in ten steps, each with a diagram, and attention you can compute by hand. After that come shapes and parameter counts, a comparison table, staff-level interview questions, an exercise you can run, and a one-page cheat sheet.

Timeline

From counting to attention: what each step fixed in the one before

  1. Origin

    Google, the paper "Attention Is All You Need", Vaswani and seven co-authors

    The problem it set out to solve

    RNNs cannot be parallelized and are slow to train; information between distant words has to travel step by step. The task at the time was machine translation.
    One Transformer layer (decoder-only, the modern form), data flowing from top to bottom
    • Inputone vector per token
    • First sublayer
      • Norm
      • masked self-attentioneach position sees only itself and earlier ones
      • Add back the inputresidual connection
    • Second sublayer
      • Norm
      • MLPwiden to 4x, then narrow back
      • Add back the inputresidual connection
    • Outputsame shape as the input, passed to the next layer

    How it works

    • Self-attention: every word looks directly at every word in the sentence and takes in their information weighted by relevance. Any two words are one step apart.
    • The computations for different positions do not depend on each other and can be done in one go, which suits GPUs.
    • Attention itself does not know word order, so a positional encoding is added for each position.
    • Each layer is attention followed by an MLP. Both are wrapped in a residual connection and LayerNorm.
    • The original paper is an encoder plus a decoder. The base model has 6 layers, 512-dimensional vectors, 8 heads and about 65 million parameters.

    Training, step by step

    1. Feed in all positions of a passage at once.
    2. Use a mask so that no position sees what comes after it.
    3. Every position predicts the next word.
    4. Add up the error over all positions and update the parameters.

    Inference, step by step

    1. Feed in the text so far and get the last position's probabilities for the next word.
    2. Choose a word by probability.
    3. Append it and compute again.
    4. The keys and values already computed are kept in a cache and not recomputed.

    Strengths

    • All positions are computed at once, so training makes full use of GPUs.
    • Any two words are one step apart, so long-range dependencies are easy to learn.
    • The structure is uniform, and results improve as the model and the data grow.

    Weaknesses

    • The cost of attention grows with the square of the sequence length, so long texts are expensive.
    • It has no built-in notion of order and relies on positional encoding to supply one.
    • It brings few built-in assumptions and needs a great deal of data to show its strength. On images, with little data, it does worse than a CNN.

    Where it is used, and where it still is

    Nearly all large language models today, and most vision and speech models.

    What replaced it, and why

    Not replaced. Newer structures are mostly modifications of it or are mixed with it.

    In an interview, one sentence

    A sequence model using only attention: parallel computation, and any two positions directly connected.

    In an interview, two minutes

    The Transformer removes the RNN and uses only self-attention. Every word looks directly at every other word and takes information by relevance, so any two words are one step apart and all positions can be computed at once. Because attention does not know order, positional encoding is added separately. Each layer is attention plus an MLP, wrapped in residual connections and normalization. The price is a cost that grows with the square of the length. It became mainstream because parallelism let it train on large data, and its results keep improving with scale.

    Worth adding

    A whole section below takes it apart. The three things interviewers most often press on are: why divide by the square root of d, why positional encoding is needed, and why the cost is quadratic.

Five concepts first

Every step on the timeline trades off these five things

Token and embedding
Text is first cut into tokens (words or pieces of words), and each token is looked up in a table to get a vector.
A model can only compute on numbers. All later computation happens on these vectors, and the model's "width", d_model, is the dimension of this vector.
How far the context reaches
When predicting a word, how far back the model can reach for information.
For an N-gram it is a few words; for an RNN it is unlimited in theory, but in practice it forgets; for a Transformer it is the whole context window. This thread runs through the entire timeline.
Whether it can be parallelized
Whether the positions in a sentence must be computed one after another, or can be computed at the same time.
GPUs are good at computing many things at once. An RNN must compute in order and a Transformer can compute them all at once, which is the direct reason the Transformer can make use of large data.
Residual connection
A layer's output = the layer's input + the change the layer computed.
It leaves a straight path for gradients; without it a network of dozens of layers cannot be trained. You can picture the whole model as one main line to which every layer adds a little.
Autoregressive
Generate one token at a time, append it to the input, then generate the next.
This is how GPT-style models generate text, and it is why longer outputs are slower and a KV cache is needed.

Main topic · the figure from the paper first

The classic 2017 structure: encoder plus decoder

This is Figure 1 of "Attention Is All You Need". When interviews and textbooks say "the structure of the Transformer", this is usually what they mean. The task then was translation: the encoder on the left reads the source, and the decoder on the right writes the translation. Read the figure from the bottom up.

encoder: reads the sourcedecoder: writes the translationProbability of the next wordSoftmaxscores → probabilitiesLineara score for every wordRepeated N times (N = 6 in the paper)Add & Normadd residual, normalizeFeed ForwardMLPAdd & Normadd residual, normalizeMulti-Head Attentioncross-attention: sees encoder outputAdd & Normadd residual, normalizeMasked Multi-Head Attentionself-attention: earlier words onlyRepeated N timesAdd & Normadd residual, normalizeFeed ForwardMLPAdd & Normadd residual, normalizeMulti-Head Attentionself-attention: sees all wordsEncoder output: one vector per source word+Positional Encoding+Positional EncodingInput Embeddingsource words → vectorsOutput Embeddingtranslation words → vectorsInput: the source我 爱 猫 (I love cats)Input: translation so far, after a start symbol<start> I love

The figure is wide; scroll sideways to see all of it.

In the paper's figure, each Add & Norm also has a line beside it that goes around the sublayer. That is the residual connection. It is left out here for clarity and written in the boxes' notes instead.

Translating one sentence: how the data moves

  1. The source "我 爱 猫" (I love cats) enters at the bottom left, becomes vectors, and gets positional encoding added.
  2. It passes through 6 encoder layers. Each does self-attention first (the three words look at each other), then an MLP. Out come three vectors, one per word, now carrying information from the whole sentence. The encoder runs only once.
  3. The decoder starts at the bottom right. At the first step its only input is a "start" symbol.
  4. Each decoder layer does three things: self-attention among the words written so far, with a mask so it sees only earlier ones; then cross-attention, looking back at the three vectors the encoder output; and finally an MLP. All 6 decoder layers look at the output of the encoder's last layer.
  5. The Linear and Softmax at the top give the probability of the next word, from which "I" is chosen. The paper uses beam search, keeping several candidates at once and not just the single most probable word at each step.
  6. "I" is appended to the decoder's input, and steps 4 and 5 repeat until the end symbol is output.

How cross-attention differs from self-attention

The formula is exactly the same; the only difference is where Q, K and V come from. In self-attention all three come from the same sequence of words. In cross-attention, Q comes from the vector at each position in this decoder layer (just after the masked self-attention), and K and V come from the encoder's output. The meaning: the decoder takes "what I am about to write" and goes to the source to find the related words. This is the Transformer version of the 2014 attention; follow-up 4 below has an animation.

Where each part of the figure is covered below

Part of the figureWhere in the ten stepsIn today's LLMs
Input / Output EmbeddingStep 2Kept, as a single table
Positional EncodingStep 3Most open models have switched to RoPE
The encoder's Multi-Head AttentionSteps 4 and 5, without a maskNo encoder, so this part is gone
Masked Multi-Head AttentionSteps 4, 5 and 9Kept; the only kind of attention left in the model
The Multi-Head Attention in the middle of the decoderThe subsection aboveNo encoder to look at, so it is gone
Add & NormStep 6Kept; Norm moved to before the sublayer
Feed ForwardStep 7Kept; some models replace it with MoE
Linear and SoftmaxStep 8Kept

Remove the left half of this figure and the attention in the middle of the decoder, and what remains is the skeleton of today's large language models; the changes to the parts are in the table above. The ten steps below follow this reduced structure, because it has fewer parts and is the shape actually in use today. The part each step covers is the same part as in this figure.

Main topic · the Transformer taken apart

Ten steps, from one sentence to the next token

1The whole picture first: a decoder-only Transformer

Nearly all of today's large language models have this shape. Text goes in at the top and flows down, and what comes out at the end is the probability of "what the next token is". The middle block is repeated N times, 12 in GPT-2 small.

Data flows from top to bottom
  • Text"cat chases mouse it"
  • Tokenizecut into tokens, each replaced by an ID
  • token embeddingeach ID is looked up and becomes a vector
  • Add position informationtells the model where each token sits in the order
  • Transformer layer, repeated N times
    • Norm → masked self-attention → add back the inputtakes information from other tokens
    • Norm → MLP → add back the inputeach token processes what it took in, on its own
  • One last Norm
  • Output layerat each position, a score for every token in the vocabulary
  • softmaxturns scores into probabilities; the last position's are the next token's

One way of seeing it is very useful: picture each token's vector as a main line running from top to bottom. No layer replaces it; each only adds a little to it. Attention is responsible for fetching information from other tokens, and the MLP for processing it.

2Tokenization and embedding: turning text into vectors

A model can only compute on numbers. The first step is cutting the text into tokens, each matching an ID in the vocabulary. The ID is then used to look up a vector in the embedding table. The table has "vocabulary size" rows and "d_model" columns, and it is learned in training.

Four tokens, each becoming an ID and then a vector. The IDs are made up for the example.
token
catchasesmouseit
ID
37218921504456
Vector
768 numbers768 numbers768 numbers768 numbers

A token is not necessarily a whole word. A common word is one token, and a rare word is cut into several pieces. GPT-2's vocabulary has 50,257 tokens.

3Positional encoding: telling the model who comes where

Attention does the same thing for every word: look at all the words and weight them by similarity. It does not know word order, and to it "dog bites man" and "man bites dog" are the same set of words. So position information has to be put in separately.

The original paper computes, for each position, a vector as long as the embedding and adds it to the embedding. Each dimension of this vector is a sine or cosine wave, and the later the dimension, the longer the wave.

The original paper's sinusoidal positional encoding, drawn here with 16 dimensions and the first 12 positions. Each cell's color is the value actually computed.
Position 0
Position 1
Position 2
Position 3
Position 4
Position 5
Position 6
Position 7
Position 8
Position 9
Position 10
Position 11
−1+1Each row is one position's 16-dimensional encoding, dimensions 0 to 15 from left to right

Look at the picture: the dimensions on the left change quickly, so neighboring positions already differ in color; those on the right change slowly, and a difference shows only far apart. Together they give every position a unique pattern. Today's models mostly use RoPE, which is not added to the input and acts inside attention, as described later.

4Self-attention: every word looks at the other words

This is the core of the Transformer. Every word has to answer one question: which words in the sentence are related to me, and what information should I take from them?

It does so by turning each word's vector into three new vectors:

NameRoleBy analogy
query (Q)what I am looking forthe words typed into a search box
key (K)what I am, and what I can be found byeach web page's title
value (V)what I offer once foundthe body of the web page

Below, the whole computation is laid out with the four words "cat, chases, mouse, it", and every number shows where it came from.

Three tables: each word's query, key and value. These are all the inputs to the whole computation.
Wordquery (what I am looking for)key (what I can be found by)value (what I offer)
cat[1, 0][2, 0][1, 0]
chases[1, 1][0, 1][0, 1]
mouse[0, 1][1, 1][0.5, 0.5]
it[2, 0.5][0, 0.5][0, 0]

In a real model these three vectors are computed: take the word's current vector x and multiply it by three trained matrices, q = x·W_Q, k = x·W_K, v = x·W_V. Here that step is skipped and the results are written by hand, 2 dimensions each, so the arithmetic can be done in your head.

The formula is softmax(Q·Kᵀ / √d_k)·V, four steps from the inside out. Below we follow just one word, "it", through all four.

Following "it" through. Its query is [2, 0.5].

Step 1: Dot product with each word's keyQ·Kᵀ

A dot product multiplies the numbers in matching positions and adds them up; a larger result means more related.

  • it · cat = 2×2 + 0.5×0 = 4
  • it · chases = 2×0 + 0.5×1 = 0.5
  • it · mouse = 2×1 + 0.5×1 = 2.5
  • it · it = 2×0 + 0.5×0.5 = 0.25

Step 2: Divide by √d_k/ √d_k

Here d_k is 2, and √2 is about 1.414.

  • cat: 4 ÷ 1.414 = 2.828
  • chases: 0.5 ÷ 1.414 = 0.354
  • mouse: 2.5 ÷ 1.414 = 1.768
  • it: 0.25 ÷ 1.414 = 0.177

Step 3: Softmax: weights that add up to 1softmax(…)

First replace each number x with e to the power x, then divide each by the sum of the four results 25.39.

  • cat: e^2.828 = 16.92, 16.92 ÷ 25.39 = 66.6%
  • chases: e^0.354 = 1.42, 1.42 ÷ 25.39 = 5.6%
  • mouse: e^1.768 = 5.86, 5.86 ÷ 25.39 = 23.1%
  • it: e^0.177 = 1.19, 1.19 ÷ 25.39 = 4.7%

Step 4: Add up each word's value by weight· V

0.666×[1, 0] + 0.056×[0, 1] + 0.231×[0.5, 0.5] + 0.047×[0, 0] ≈ [0.782, 0.171]

This is the new information "it" gets from this attention layer, and it comes mainly from the value of "cat".

All four words doing this at once is a matrix operation. In the tables below, each row is one word looking at the others; the row for "it" holds the numbers just computed. No mask is applied yet; with one, each row would have weight only on itself and the words to its left.

Result of step 1, Q·Kᵀ: the dot product of every word's query with every word's key.
Looking ↓ / looked at →catchasesmouseit
cat2010
chases2120.5
mouse0110.5
it40.52.50.25
Result of step 3: the attention weights. Each row adds up to 100%.
Looking ↓ / looked at →catchasesmouseit
cat50.5%12.3%24.9%12.3%
chases35.2%17.4%35.2%12.2%
mouse15.4%31.3%31.3%22.0%
it66.6%5.6%23.1%4.7%
Result of step 4: each word's new vector.
WordOutput
cat[0.630, 0.247]
chases[0.528, 0.350]
mouse[0.311, 0.469]
it[0.782, 0.171]

What changes with a different word, or with the causal mask on? The table below is yours to click.

"it" puts 67% of its attention on "cat"

The query of "it" is [2, 0.5]. It takes a dot product with each word's key, divides by √2, then goes through softmax.

Word looked atIts keyDot product÷ √2Weight
cat[2, 0]42.83
66.6%
chases[0, 1]0.50.35
5.6%
mouse[1, 1]2.51.77
23.1%
it[0, 0.5]0.250.18
4.7%

Finally the values of all the words are added up by weight, and the new vector of "it" is [0.782, 0.171].

Look from another word
causal mask

The vectors here are hand-written 2-dimensional numbers, so the arithmetic can be done in your head. In a real model they have hundreds to thousands of dimensions and are learned in training.

Why divide by √d: the d here is the dimension of the key in each head, written d_k in the paper; it is 2 in this example and 64 in GPT-2 small. The more dimensions, the more the dot products swing. When the values are too large, softmax turns into "one is 1 and the rest are 0", the gradient is almost 0, and the model cannot learn. Dividing by √d pulls the values back into a normal range.

5Multi-head: looking in several ways at once

One set of Q, K and V can express only one kind of "related". But words relate to each other in many ways: which is the subject, which one a pronoun refers to, which one is being modified. So 12 different sets of projection matrices each project the same 768-dimensional vector into 64-dimensional Q, K and V, and each set runs attention once. Each is called a head.

One layer of GPT-2 small: 12 heads, each projecting 768 dimensions down to 64, each looking in its own way, joined back into 768 dimensions at the end.
  1. Input

    768 dimensions per token

  2. 12 heads

    each head uses its own matrices to project 768 dimensions into 64-dimensional Q, K and V

  3. Concatenate

    12 × 64 = 768 dimensions

  4. One more linear layer

    outputs 768 dimensions

The total computation is about that of one large 768-dimensional head, but 12 heads can each attend to something different.

6Residual connections and normalization: what lets dozens of layers stack

The output of attention does not replace the original vector; it is added back to it. This is the residual connection, and it comes from ResNet. Normalization brings each token's vector to around mean 0 and variance 1, followed by a learnable scale and shift, to stop the values growing or shrinking layer after layer.

Two ways of writing the same sublayer. The only difference is where the normalization goes.

Post-Norm (the original 2017 paper)

  1. Input x

  2. Sublayer

    attention or MLP

  3. Add x

  4. Norm

    normalization stands on the main line

Pre-Norm (today's practice)

  1. Input x

  2. Norm

    normalizes only the copy sent into the sublayer

  3. Sublayer

    attention or MLP

  4. Add x

    nothing stands on the main line

Under Pre-Norm, the main line x is clear from the first layer to the last and gradients pass back unchanged, so very deep networks are easier to train.

7MLP: each word processes its information on its own

After attention comes a small two-layer network, also called feed-forward. It does the same thing to each token separately: widen the vector to 4 times its width, pass it through a nonlinear function, then narrow it back to the original width.

The MLP of GPT-2 small. It does not look at other tokens and processes only its own.
  1. Input

    768 dimensions

  2. Linear layer, widen

    3072 dimensions

  3. Nonlinearity

    GELU

  4. Linear layer, narrow

    768 dimensions

The division of labor is this: attention moves information between tokens, and the MLP computes on the information that was moved. In each layer the MLP has about twice the parameters of attention, two thirds of the layer. Counting the embedding, the MLPs are about 46% of the whole GPT-2 small model, and the larger the model, the closer to two thirds.

8Output: the probability of the next token

After all the layers, each position is still a d_model-dimensional vector. A final linear layer turns it into "vocabulary size" scores, called logits, one per token. Softmax turns the scores into probabilities. In generation only the last position is used.

How one token is chosen from the probabilities is set by the sampling method. The parameter adjusted most often is the temperature.

After "I love eating", the probability of "apples" is 62.1%

  • apples3
    62.1%
  • bananas2
    22.9%
  • rice1.5
    13.9%
  • rocks-1
    1.1%
temperature

The number after each word is the score the model gave it (the logit), hand-written for this example. Scores are divided by the temperature and then go through softmax: a low temperature makes high-scoring words stand out more, and a high temperature brings the words' probabilities closer together.

With the temperature close to 0, the highest-scoring token is chosen almost every time, and the output is stable but rigid. Raise it and low-scoring words also get a chance, and the output is more varied and more prone to error.

9Training: every position predicts the next token

In training a whole passage is fed in at once, and every position predicts the token that follows it. The answers are the text itself, shifted by one position.

Inputs and answers. A passage of 4 tokens provides 3 questions.
Input
catchasesmouse
Answer
chasesmouseit

There is a problem here: when predicting the 2nd token, a model that could see the 2nd token would be copying the answer. The causal mask solves it: inside attention, later positions are hidden, and each token sees only itself and earlier ones.

The causal mask. Each row is a token; blue is what it can see, and × is what is hidden.
catchasesmouseit
catsees×××
chasesseessees××
mouseseesseessees×
itseesseesseessees

With the mask, all positions can be computed at once without leaking to each other. This is why a Transformer trains fast: an RNN goes step by step, while a Transformer computes the whole passage in one go. The error is cross-entropy: the lower the probability the model gives the correct answer, the larger the error.

10Generation: writing one token at a time, and the KV cache

In generation there are no answers to check against. The model outputs a token, appends it to the input, and runs again to get the next. This is called autoregression.

The naive approach recomputes the whole passage every time. But the keys and values of the earlier tokens are exactly the same as last time, because those tokens see only what is before them. Store them and compute only the new one each time: that is the KV cache.

Generating the 4th token. The thick border marks what actually has to be computed at this step.

Without a cache: all four are recomputed

key, value
catchasesmouseit

With a cache: the first three are fetched, and only the new one is computed

From cache
catchasesmouse
Computed
it

The MLPs of the earlier positions need not be recomputed either, because their output at every layer is unchanged. The price is GPU memory: every layer has to store every token's keys and values, and the longer the context, the more it takes.

Four follow-ups

Points the timeline passed over in a sentence, expanded

1How is smoothing done in an N-gram model?

The problem is this: a combination that never appears in the corpus gets a counted probability of 0. One such combination in a sentence makes the whole sentence's probability 0. Smoothing takes a little probability from the combinations that were seen and gives it to those that were not.

The simplest is add-one smoothing: pretend every combination was seen one more time. Suppose the vocabulary has only three words, and after "I love" they appear 30, 10 and 0 times:

Add-one smoothing. Add 1 to the numerator and the vocabulary size (3 here) to the denominator.
Word after "I love"CountUnsmoothedAdd-one
you3030 ÷ 40 = 75%31 ÷ 43 = 72.1%
food1010 ÷ 40 = 25%11 ÷ 43 = 25.6%
running00 ÷ 40 = 0%1 ÷ 43 = 2.3%

Add-one smoothing is too crude: a real vocabulary has tens of thousands of words, there are too many unseen combinations, and too much probability is given away. Two other ideas are used in practice. One is backoff: "I love running" was never seen, so step back to "love running", and failing that to just "running". The other is interpolation: mix the 3-gram, 2-gram and 1-gram probabilities in proportion. When backing off, a little probability first has to be deducted from the seen combinations and then distributed in proportion, or the total would exceed 1. Among the classic methods the best performer is Kneser-Ney smoothing (in Chen and Goodman's comparison, its modified version). It first subtracts a small fixed number from the count of every seen combination and gives the probability saved to the lower-order model; and the lower-order model looks not at how many times a word appears but at how many different words it follows.

2Gradient clipping when training an RNN: how exactly is it done?

The gradient is a vector that tells each parameter which way to move and by how much. When gradients explode this vector becomes extremely long, the parameters take too large a step, and training collapses. Clipping keeps the direction and only squeezes the length to within an upper limit.

Clipping by norm. An example with the threshold set to 1.
  1. Compute the gradient

    [3, 4]

  2. Compute its length

    √(3² + 4²) = 5

  3. Above the threshold of 1

    Multiply each number by 1 ÷ 5

  4. The clipped gradient

    [0.6, 0.8], with length 1

A gradient whose length is not above the threshold is left alone. The length here is computed over the gradients of all the parameters joined together. Note that it deals only with exploding gradients, not vanishing ones: a vanished gradient cannot be rescued by scaling it up, what eases it is the straight path of the LSTM's cell state. Many published Transformer training configurations use gradient clipping too, often with a threshold of 1, LLaMA for example.

3I learned convolution in signals and systems. What problem does it actually solve here?

The operation is the same: slide a small "kernel" along the signal, and at each position multiply the matching numbers and add them up. The difference is where the kernel comes from. In signals and systems the kernel is the system's impulse response, set by the system itself, or a filter someone designed, such as a low-pass filter. In a neural network the numbers in the kernel are learned in training.

The kernel is [−1, 1], sliding along a row of numbers. It outputs a non-zero value where the numbers change, which makes it an "edge finder".
000555
(−1)×0 + 1×0 = 0
000555
(−1)×0 + 1×0 = 0
000555
(−1)×0 + 1×5 = 5
000555
(−1)×5 + 1×5 = 0
000555
(−1)×5 + 1×5 = 0

This kernel has only two numbers and yet finds edges at any position. That is the advantage of convolution, and there are three parts to it:

AdvantageWhat it meansExample
Few parametersThe same kernel is reused at every positionA layer of 64 kernels of 3×3 on a color image needs only 1,728 weights (not counting biases)
Shared across positions (translation equivariance)Whether the cat is on the left or the right of the image, the same detector is used; when the cat moves, the detected features move with itNo need to learn it separately for every position
Uses localityNeighboring pixels are the most closely related, so look at a small area firstStacked layer upon layer, the area seen grows larger and larger

Compare the approach without convolution: a 224×224 color image has about 150,000 numbers, and if every neuron in the next layer were connected to all of them, 1,000 neurons would need 150 million weights. One small detail: "convolution" in neural networks does not flip the kernel and is strictly cross-correlation, but since the kernel is learned, flipping makes no difference.

4What does "for every word it generates, the decoder scores each position of the source and takes a weighted sum" mean?

Take translating "我 爱 猫" (I love cats). The encoder reads the source, and each of the three words leaves a vector, which you can think of as "what this word means in this sentence". The decoder has to write the translation one word at a time.

The earliest approach handed the decoder only the single fixed-length vector the encoder ended with, the equivalent of translating from an overall impression of the sentence. Attention's approach is: before writing each word, look back over all three source words and decide which to rely on most for this step.

Writing "I", the decoder puts 85% of its attention on "我"

The encoder has read the source; each word leaves a stateThe decoder writes the translation one word at a time我 · I85%爱 · love10%猫 · cats5%I??

Reference for this step = 0.85 × state of "我" + 0.10 × state of "爱" + 0.05 × state of "猫"

The thicker the line, the more this step relies on that source word. The weights here are hand-written for illustration, not the output of a real model.

For every word written, three things happen:

StepWhat is doneWhen writing "love"
ScoreFeed the decoder's state before writing this word, together with the state of each source word, into a small network, which gives a score for each我: low, 爱: high, 猫: low
Turn into weightsApply softmax to the scores so they add up to 18%, 84%, 8%
Weighted sumAdd up the states of the source words by weight, giving one vector0.08×"我" + 0.84×"爱" + 0.08×"猫"

The vector that results is "the reference for this step". The decoder combines it with its own state and outputs "love". For the next word the scoring is done again. The attention in the middle of the decoder in the Transformer diagram does exactly this, except that scoring by a small network is replaced by a dot product, and it is done once in every layer.

A walk through the shapes

What shape a 5-token sentence takes inside GPT-2 small

Getting the shape right at every step in an interview shows real understanding. GPT-2 small: 12 layers, d_model 768, 12 heads, vocabulary 50,257.

StepShapeMeaning
Token IDs(5)5 integers
embedding(5, 768)one 768-dimensional vector per token
Add position(5, 768)shape unchanged
Q, K and V of each head12 × (5, 64)each head projects 768 dimensions to 64
Attention weights12 × (5, 5)in each head, one weight from every token to every token
Attention output(5, 768)the 12 heads joined back together
Middle of the MLP(5, 3072)widened 4×
Output of one layer(5, 768)the same as the input, which is why layers stack
After 12 layers(5, 768)still this shape
logit(5, 50257)at each position, a score for every token
Probability of the next token(50257)take only the last position and apply softmax

Note the row for attention weights: it is 5×5. With a sentence of length n it is n×n. That is where "the cost grows with the square of the length" comes from.

Parameter count

Where the parameters are

Pick a model size and see how the parameters are distributed. For all four sizes this formula gives exactly the parameter count obtained from Hugging Face's GPT-2 implementation; small is 124,439,808. OpenAI's paper says 117M, which differs from what the released weights count to.

Total parameters

124,439,808

About 124 M, of which the MLPs are 46% and attention 23%

  • MLP56,669,184
  • attention28,348,416
  • token embedding38,597,376
  • position embedding786,432
  • normalization38,400

A way to do it in your head: each layer is about 12 × d² (attention is 4 matrices of d×d and the MLP the equivalent of 8), plus the embedding of vocabulary × d.

Why it won

Side by side with RNNs and CNNs

This table comes from the original paper. n is the sequence length, d the vector width, and k the size of the convolution window.

Layer typeComputation per layerSteps that must be done in orderFurthest distance between two words, in steps
self-attentionn² · d11
RNNn · d²nn
Convolutionk · n · d²1log_k(n)

The last two columns are the key. "Steps that must be done in order" being 1 means it can be parallelized; "furthest distance" being 1 means even the most distant two words influence each other directly. The price is in the first column: when n is very large, the computation of self-attention exceeds that of RNNs and convolution. The log_k(n) in the convolution row refers to dilated convolution; ordinary convolution needs n/k stacked layers for the two most distant words to meet.

Three ways of using it

encoder-onlydecoder-onlyencoder-decoder
ExamplesBERTGPTThe original Transformer, T5
What it can seeBoth sidesOnly what came beforeThe encoder sees both sides; the decoder sees what came before, and the encoder
Training taskFill in the blankPredict the next tokenGiven the input, generate the output
Good atUnderstanding, retrieval, classificationGeneration, general-purpose tasksTasks with one thing in and another out, such as translation and speech to text

The 2017 original and today's LLMs

PartOriginal paperCommon practice now
Overallencoder + decoderDecoder only
Where normalization sitsAfter the sublayer (Post-Norm)Before the sublayer (Pre-Norm)
Kind of normalizationLayerNormRMSNorm
Position informationSinusoidal encoding, added to the inputRoPE, rotating Q and K inside attention
Activation in the MLPReLUGELU or SwiGLU
Attention headsEach head has its own K and VGQA: several heads share one set of K and V
MLPOne per layerOften replaced by MoE in large models

Side by side

The same dimensions, in one table

ArchitectureHow far it seesParallelCostHow it knows orderBest for
N-gramThe previous N−1 wordsNot applicableTable lookupBy the fixed windowA very small, very fast baseline
Word2VecA few words either side in training; no context in useNot applicableTable lookupDoes not know orderVector representations of words
RNNEverything in theory, a few dozen words in practiceNoProportional to lengthReads in order, so it knowsVery small sequence models
LSTM / GRUA hundred or so wordsNoProportional to lengthReads in order, so it knowsTime series, small models
CNNOne layer sees only nearby; many layers are stackedYesProportional to lengthBy the window's positionImages, edge devices
Seq2Seq + AttentionThe decoder sees every source position directlyNoSource length × translation lengthReads in order, so it knowsMachine translation before 2017
TransformerThe whole context window, in one stepYes, in trainingProportional to the square of the lengthPositional encoding added separatelyLarge-scale language, vision and multimodal models
BERTThe whole sentence, both directionsYesProportional to the square of the lengthLearned position vectorsRetrieval, classification
GPTEverything beforeYes in training; one at a time in generationProportional to the square of the lengthLearned position vectors in GPT-1 and GPT-2; later open decoder-only models mostly use RoPEGeneration, general-purpose models
ViTAll patches of the image, both directionsYesProportional to the square of the number of patchesLearned position vectorsImage understanding, the vision part of multimodal models
Modern LLMsEverything before; windows reach hundreds of thousands of tokensYes in training; one at a time in generationProportional to the square of the length, with a smaller constantRoPEGeneral-purpose large language models
Diffusion / DiTThe whole imageYes within a step; no between stepsA full pass of the network per step, for many stepsDecided by the denoising networkImage, video and audio generation
MambaEverything, but compressed into a state of fixed sizeYes, in trainingProportional to lengthReads in order, so it knowsVery long sequences

Where the industry is heading · checked October 2026

What people are working on now

  • MoE has become common practice in large models

    Replace each layer's MLP with many experts, and send each token through only a few. Total parameters can be very large while each computation uses few. DeepSeek-V3 and the largest Qwen3 model, at 235B, both have this structure.

  • Mixing attention with linear-cost layers

    For longer contexts, some models no longer use full attention in every layer; most layers use a linear-cost structure, with a full attention layer every few layers. Sebastian Raschka's list of papers from June 2026 says this alternating hybrid structure was one of the more popular directions that year; the examples in his November 2025 article are Qwen3-Next and Kimi Linear: three linear attention layers to one full attention layer.

  • Making attention itself cheaper

    GQA lets several heads share keys and values and shrinks the KV cache. FlashAttention avoids writing the whole attention matrix to GPU memory. These do not change the quadratic order, but they make long contexts feasible in practice.

  • Generating text with diffusion

    Autoregression produces one token at a time. Diffusion language models aim to generate many tokens in parallel in one go and then revise repeatedly. They have not replaced autoregression, and they are an active research direction.

Common misconceptions

These sound right but are not

  • Misconception: The Transformer invented attention.

    Attention was already used in RNN-based machine translation in 2014. The Transformer's contribution was removing the RNN and using only attention.

  • Misconception: Attention weights are the model's explanation; a large weight means that word is important.

    A weight only says where this one head in this one layer took information from. A model has many layers and many heads, with MLPs and residuals after them, and one head's weights cannot be taken as an explanation of the model's decision.

  • Misconception: The Transformer is parallel, so generation is parallel too.

    Parallelism refers to training, and to processing input that is already given. Generation still goes one token at a time, each token depending on the one before.

  • Misconception: BERT and GPT are two different architectures.

    They use the same kind of layer. The difference is whether attention can see later words, and whether the training task is filling in blanks or predicting the next word.

  • Misconception: Diffusion is an alternative to the Transformer.

    One is a method of generation and the other a network structure. DiT uses a Transformer as the denoising network in diffusion.

  • Misconception: Most of the parameters are in attention.

    In a standard layer the MLP has about twice the parameters of attention: attention is 4 matrices of d×d, and the MLP is the equivalent of 8.

  • Misconception: Positional encoding is a fixed vector added to the input.

    That is the original paper. The mainstream today is RoPE: nothing is added to the input, and the query and key are rotated inside attention at every layer.

Staff-level interview questions

Answer first, then open

▶Why does attention divide by the square root of d?
  • The d here is d_k, the dimension of the key in each head, not d_model. In GPT-2 small the division is by √64 = 8, not √768.
  • A dot product multiplies d_k pairs of numbers and adds up the products. If each dimension has mean 0 and variance 1, the variance of the dot product is d_k, so the more dimensions, the larger the swings.
  • When dot products are large, the softmax output approaches "one is 1 and the rest are 0". The gradient is then almost 0 and the model cannot learn.
  • After dividing by √d_k the variance is back to 1 whatever the dimension, and the softmax does not saturate.
  • This is only one way to make the starting state reasonable. Whether and how much to scale depends on the scale of the vectors: some models normalize the query and key first and then multiply by a learnable coefficient, to the same end.
▶Why does a Transformer need positional encoding, and an RNN not?
  • Self-attention does the same thing for every word: look at all the words and weight them by similarity. Shuffle the input words and each word gets the same result, just moved to its new place. By itself it cannot tell "dog bites man" from "man bites dog".
  • An RNN reads in order, so order is already in the computation.
  • So a Transformer has to put position information in explicitly. The original paper adds it to the input vectors; RoPE, common today, acts on the query and key inside attention.
  • One addition: the argument above holds only for attention without a mask. With a causal mask each position sees a different number of words, and the model can infer some position information from that. Which positional encoding to use depends on the need: to extrapolate to longer contexts, relative position (RoPE) is better than absolute; when the context is short and fixed, the difference is small.
▶What is the cost of self-attention, and where is the bottleneck?
  • With length n and dimension d, attention costs n squared times d: every position is computed against every position. The MLP costs n times d squared.
  • With constants: the part of attention that depends on n² is about 2n²d, the MLP is 8nd², and attention's projection matrices are 4nd². The two sides are equal at roughly n between 4 and 6 times d. For GPT-2 small (d = 768) that is three to five thousand tokens; for a model with d = 4096, sixteen to twenty-four thousand. Shorter than that, the MLP dominates; longer, attention does.
  • With long contexts the real bottleneck is often memory: the n×n attention matrix, and the KV cache at inference.
  • Which optimization to choose depends on where the bottleneck is: FlashAttention reduces GPU memory reads and writes, GQA shrinks the KV cache, and sparse or linear structures change the quadratic order at some cost in accuracy.
▶Why are today's large models decoder-only and not encoder-decoder?
  • Unified tasks: every task can be written as "given what came before, continue", with no need to separate input and output.
  • Dense training signal: every position predicts the next word.
  • A simple structure that is easy to scale. Input and output are in the same sequence, one KV cache covers both, and an identical beginning can be reused across requests.
  • Encoder-decoder is not wrong. Where input and output are clearly two different things (speech to text, for example), it still makes sense. Which to choose depends on whether the task is naturally "one thing in, another out" and on whether one model has to do everything.
▶What is the difference between Pre-Norm and Post-Norm, and why is Pre-Norm used now?
  • Post-Norm is the original paper's approach: compute the sublayer, add the residual, then normalize. Pre-Norm normalizes first, computes the sublayer, then adds the residual.
  • Under Pre-Norm no normalization stands on the residual main line, so gradients pass back unchanged to the very front, and deep networks are easier to train.
  • Post-Norm holds up well when the network is shallow and carefully tuned. Which to choose depends mainly on depth: above a few dozen layers, Pre-Norm is used almost universally.
▶What is the KV cache? What does it save, and what does it cost?
  • When generating word t, the outputs of the previous t−1 positions at every layer are exactly the same as at the previous step, because those positions see only the words before them. So their keys and values need not be recomputed, and neither do their MLPs.
  • Store them, and each step only computes the new word's query, key and value and attends once over the cached keys and values. The computation per step falls from "proportional to the square of the length" to "proportional to the length".
  • The cost is GPU memory. The cache's size is: 2 (one copy each for keys and values) × layers × tokens × heads that store keys and values × dimension per head × bytes per number, once for every sequence in the batch. The longer the context, the more it takes, and with long contexts it is often the largest share of memory.
  • Hence GQA (several query heads share one set of keys and values, directly reducing the number of heads to store) and the various methods of compressing the cache.
  • Whether to compress the cache, and by how much, depends on whether the bottleneck is compute or memory, and on how many concurrent requests there are.
▶Given a model's configuration (layers, width, vocabulary size), how do you estimate its parameter count?
  • Each layer is about 12 times d squared: attention is 4 matrices of d×d, and the MLP is two matrices of d×4d and 4d×d, 8 of d×d in all.
  • The embedding is vocabulary size times d.
  • The total is roughly 12 × layers × d² + vocabulary × d. Check it on GPT-2 small: 12 × 12 × 768² is about 85 million, plus 50257 × 768, about 39 million, gives about 124 million, in line with the published figure.
  • The actual number also depends on a few details: whether the output layer shares parameters with the embedding, whether positional encoding is learned, whether the MLP is 4× wide, and whether MoE is used.
▶Growing the context window from 8K to 1M: where do problems appear?
  • Computation: processing the input, attention is quadratic, so with length up 125 times this part grows more than fifteen thousand times (125²). With a KV cache, each token generated afterwards costs in proportion to the length.
  • GPU memory: the KV cache is proportional to the length, stored for every head of every layer.
  • Positional encoding: positions this far were never seen in training, and the model may not know how to use them. Relative schemes such as RoPE extrapolate somewhat better, but only so far.
  • Quality: a window that is long enough does not mean the model uses it well, and content in the middle is easily ignored.
  • The solution depends on the goal: to merely fit so long an input in memory, use a more memory-efficient attention implementation and cache compression; to make it cheap, use retrieval to put in only the relevant parts; to solve it structurally, mix in linear-cost layers.
▶Design question: deploying a latency-sensitive generation service, what do you look at in the architecture?
  • First separate two measures: time to first token depends on the speed of processing the input and so on its length; the latency of each later token depends on the speed of a single generation step.
  • Processing the input can be parallelized and done in one pass. Generation can only go one at a time, and a KV cache is essential.
  • GPU memory decides how many requests one machine can serve at once. The KV cache's size is proportional to the number of layers, the context length and the number of heads that store keys and values; GQA and quantization reduce it directly.
  • An MoE model uses few parameters per token, but all the parameters have to be in memory.
  • The trade-off depends on what requests look like: with long input and short output, optimize input processing and cache reuse; with long output, optimize single-step speed and batching.

Hands-on exercise · about 5 minutes

Write self-attention once in NumPy

Requires Python and NumPy. Save the code below as attention.py and run it. The core is only four lines.

import numpy as np

np.set_printoptions(precision=3, suppress=True)

tokens = ["cat", "chases", "mouse", "it"]

# The query, key and value vector of each word. These are hand-written
# 2-dimensional vectors so the arithmetic can be done in your head; in a real
# model they come from the embedding times three trained matrices.
Q = np.array([[1, 0], [1, 1], [0, 1], [2, 0.5]])
K = np.array([[2, 0], [0, 1], [1, 1], [0, 0.5]])
V = np.array([[1, 0], [0, 1], [0.5, 0.5], [0, 0]])


def softmax(x):
    e = np.exp(x - x.max(axis=-1, keepdims=True))
    return e / e.sum(axis=-1, keepdims=True)


def attention(Q, K, V, causal):
    d = K.shape[-1]
    scores = Q @ K.T / np.sqrt(d)          # step 1: similarity of every query with every key
    if causal:                              # step 2: a word sees only itself and earlier words
        hidden = np.triu(np.ones_like(scores), k=1).astype(bool)
        scores = np.where(hidden, -np.inf, scores)
    weights = softmax(scores)               # step 3: each row becomes weights that sum to 1
    return weights, weights @ V             # step 4: add up the values by weight


weights, output = attention(Q, K, V, causal=False)
print("Weights without a mask (each row is what one word attends to)")
print(weights)
print('The row for "it":', dict(zip(tokens, weights[3].round(3).tolist())))
print('Output vector for "it":', output[3])

weights, output = attention(Q, K, V, causal=True)
print("\nWeights with the causal mask")
print(weights)
print("Sum of each row:", weights.sum(axis=1))

Actual output on NumPy 1.26:

Weights without a mask (each row is what one word attends to)
[[0.505 0.123 0.249 0.123]
 [0.352 0.174 0.352 0.122]
 [0.154 0.313 0.313 0.22 ]
 [0.666 0.056 0.231 0.047]]
The row for "it": {'cat': 0.666, 'chases': 0.056, 'mouse': 0.231, 'it': 0.047}
Output vector for "it": [0.782 0.171]

Weights with the causal mask
[[1.    0.    0.    0.   ]
 [0.67  0.33  0.    0.   ]
 [0.198 0.401 0.401 0.   ]
 [0.666 0.056 0.231 0.047]]
Sum of each row: [1. 1. 1. 1.]
  • The numbers here are exactly those in the worked attention example in step 4 above, from the same vectors. You can check them cell by cell against that table.
  • With the mask, the upper right of the matrix is all 0: each word sees only itself and earlier words. The first row has only itself left, so its weight is 1.
  • "it" is the last word and nothing comes after it, so its row is the same with or without the mask.
  • Try changing things: remove the division by np.sqrt(d) and see whether the row for "it" concentrates more on "cat"; change the vector for "it" in Q to [0.5, 2] and see which word it turns to.

Cheat sheet

Ten minutes before the interview

N-gram
Counting. Sees only the previous few words; no similarity between words.
Word2Vec
Words become vectors. One vector per word, ignoring context.
RNN
Reads in order with one hidden state. Vanishing gradients, no parallelism.
LSTM / GRU
Gates and a straight path; remembers further. Still no parallelism.
Seq2Seq + Attention
The encoder reads, the decoder writes, attention looks back at the source.
ResNet
Output = input + change. What makes deep networks trainable.
Transformer
Attention only. Parallel, any two words one step apart, at a cost that grows with the square of the length.
self-attention
softmax(QKᵀ / √d_k) · V. Dividing by √d_k keeps the softmax from saturating and the gradient from vanishing.
BERT
Encoder-only, bidirectional, fill-in-the-blank pre-training. Understanding and retrieval.
GPT
Decoder-only, causal mask, predict the next word. Generation.
Modern LLMs
Pre-Norm, RMSNorm, RoPE, SwiGLU, GQA, MoE, KV cache.
Parameter count
About 12 × layers × d² + vocabulary × d.
Diffusion
Denoise step by step. A method of generation; the network can be a Transformer.
Mamba
Cost proportional to length; chooses what to keep in its state based on the input. Often mixed with attention.

Sources

"Where the industry is heading" was checked in October 2026. It may have changed since. The interactive numbers on the page (attention weights, positional encoding, parameter counts) are all computed on the spot and not hard-coded.

DISCUSSION

Join the discussion

Share a thought or question. Comments publish immediately; add your email only if you would like reply notifications.

Be the first to start the conversation.