AI fundamentals, part 3: Building a tiny GPT — the transformer block in NumPy
Part 3from the AI fundamentals series · 8 parts in all
Parts 1–2 gave us tokens and attention. Now we assemble a complete GPT-style transformer block — embeddings, attention, the feed-forward network, residual connections and layer normalization — and train it on a toy corpus. ~120 lines of NumPy, no frameworks, and by the end it generates text.
The block, precisely
A GPT block is two sub-layers with highways around them:
x ──► LayerNorm ──► Multi-head Attention ──(+ x)──┐
│◄────────────────────────────────────────────────┘
└──► LayerNorm ──► MLP (expand 4x, gelu, project) ──(+ )──► out
The residual connections (the +) are the unsung heroes:
gradients flow through them untouched, which is what lets networks go 100 layers deep. The
MLP holds most of the parameters — attention moves information between
positions; the MLP processes each position. GELU is the smooth activation modern
models use.
The components in NumPy
def gelu(x):
# Gaussian Error Linear Unit: a smooth ReLU. 0.5x(1+tanh(...)) is the
# standard approximation used in GPT-2.
return 0.5 * x * (1 + np.tanh(np.sqrt(2/np.pi) * (x + 0.044715 * x**3)))
def layer_norm(x, g, b, eps=1e-5):
# Normalize each token's vector to mean 0/std 1, then rescale with learned
# g (gain) and shift with learned b. Keeps activations in a sane range so
# deep stacks train at all.
mu = x.mean(-1, keepdims=True)
var = x.var(-1, keepdims=True)
return g * (x - mu) / np.sqrt(var + eps) + b
def mlp(x, W1, b1, W2, b2):
# Position-wise feed-forward: expand to 4x, nonlinearity, project back.
# This is where the model stores its learned facts.
h = gelu(x @ W1 + b1)
return h @ W2 + b2
def block(x, p):
# One transformer block: attention sub-layer, then MLP sub-layer,
# each wrapped in pre-LayerNorm + residual.
attn_out, _ = self_attention(layer_norm(x, p["ln1g"], p["ln1b"]),
p["Wq"], p["Wk"], p["Wv"], p["Wo"])
x = x + attn_out # residual 1: keep the highway open
x = x + mlp(layer_norm(x, p["ln2g"], p["ln2b"]),
p["W1"], p["b1"], p["W2"], p["b2"]) # residual 2
return x
The whole model: embeddings to logits
def gpt_forward(token_ids, params, n_layers=4):
# token_ids: (T,) integers from our BPE tokenizer.
# Returns logits (T, vocab): one score per vocabulary entry, per position.
T = len(token_ids)
# Token embeddings (lookup rows) + learned positional embeddings, added.
x = params["tok_emb"][token_ids] + params["pos_emb"][:T]
for i in range(n_layers): # stack the blocks
x = block(x, params["blocks"][i])
x = layer_norm(x, params["lnfg"], params["lnfb"])
# Weight tying: reuse the token-embedding matrix as the output classifier
# (huge parameter saving, slightly better results — a free lunch).
return x @ params["tok_emb"].T
def loss_and_grads(token_ids, params):
# Next-token cross-entropy over every position at once.
logits = gpt_forward(token_ids[:-1], params) # predict tokens[1:]
targets = token_ids[1:]
# softmax cross-entropy, averaged over positions
logprobs = logits - logits.max(-1, keepdims=True)
logprobs = logprobs - np.log(np.exp(logprobs).sum(-1, keepdims=True))
nll = -logprobs[np.arange(len(targets)), targets].mean()
return nll
Training loop and what to expect
Full backprop by hand is beyond one article (we use numeric gradients or a tiny autograd engine — the series repo has one), but the loop itself is five lines and shows the ritual that trains every LLM:
# Adam would be used in practice; SGD with decay shown for clarity.
for step in range(2000):
batch = random_window(train_ids, 64) # random 64-token snippet
loss = loss_and_grads(batch, params) # forward
grads = autograd(loss, params) # backward (engine or numeric)
for k in params: params[k] -= 0.01 * grads[k] # update
if step % 200 == 0: print(step, loss) # should fall ~2 -> <1.5
With a small corpus (a few hundred KB of text), 4 layers, 128 dims and ~30 minutes on a
laptop CPU, loss drops from ln(vocab) ≈ 8 (uniform guessing) to around
1.5–2.0. The model then generates locally-plausible text: real words, repetition,
no world knowledge. Scale — data, parameters, compute — is what turns that into GPT-4, and
that's exactly the next part's subject.