Trust is earned, not given

A different perspective

2023-06-08 · Projects

AI fundamentals, part 1: From bytes to BPE — how language models read text

Part 1from the AI fundamentals series · 8 parts in all

This begins a series on how large language models actually work, written for programmers who know Python but not the internals of AI. We follow the same arc as the university courses on language modeling — tokenization, architecture, training, scaling, inference — but build every piece ourselves in small, readable Python. No frameworks, no copied code. Part 1: how raw text becomes numbers.

Why tokenization exists at all

A neural network computes with floating-point vectors. It has no concept of "text". So the first question any language model must answer: what are the units? Characters are too many steps apart (a 20-token thought becomes 20+ inputs); words are too many distinct things (English alone has hundreds of thousands, plus every typo ever typed). The industry compromise is the subword: chunks like un, believ, able — a few hundred thousand of them, covering everything with no out-of-vocabulary failures.

Step 1: everything is bytes

Modern models (GPT-2 onward, Llama, Mistral) start by treating text as raw UTF-8 bytes — 256 possible values, guaranteed to handle any language or emoji without special cases:

text = "héllo world"
raw = text.encode("utf-8")          # b'h\xc3\xa9llo world'
ids = list(raw)                     # [104, 195, 169, 108, 108, 111, ...]
print(len(text), len(ids))          # 11 characters -> 12 bytes (é is two bytes)

Byte-level is the floor: a model trained this way can ingest any file ever written. The cost is efficiency — common words like the would burn 3 tokens instead of 1. That's what the next step fixes.

Step 2: BPE — compression by counting

Byte-Pair Encoding is greedy data compression applied to language: count which adjacent pair of symbols occurs most often in your training corpus, merge that pair into a new symbol, repeat vocab_size - 256 times. Frequent sequences become single tokens; rare text stays as plain bytes.

def get_pair_counts(ids):
    # Count every adjacent pair in the id sequence.
    counts = {}
    for a, b in zip(ids, ids[1:]):
        counts[(a, b)] = counts.get((a, b), 0) + 1
    return counts

def merge(ids, pair, new_id):
    # Replace every occurrence of `pair` in `ids` with `new_id` (left to right).
    out, i = [], 0
    while i < len(ids):
        if i + 1 < len(ids) and (ids[i], ids[i+1]) == pair:
            out.append(new_id)      # fused token
            i += 2
        else:
            out.append(ids[i])      # untouched token
            i += 1
    return out

# --- training ---
ids = list("the cat sat on the mat the end".encode("utf-8"))
merges = {}                          # (a, b) -> new_id
next_id = 256
for step in range(10):               # real models do ~50k-100k of these
    pairs = get_pair_counts(ids)
    top = max(pairs, key=pairs.get)  # most frequent adjacent pair
    merges[top] = next_id
    ids = merge(ids, top, next_id)
    next_id += 1

print("t h  merged early because 'th' is everywhere in English")

After training, encoding new text is just replaying the merges in order of priority:

def encode(text, merges):
    ids = list(text.encode("utf-8"))
    for pair, new_id in sorted(merges.items(), key=lambda kv: kv[1]):  # learned order
        if pair in zip(ids, ids[1:]):  # this merge still applies
            ids = merge(ids, pair, new_id)
    return ids

print(encode("the", merges))   # short! 'the' got fused during training
print(encode("thé", merges))   # 'é' is non-ASCII: stays raw bytes (2 tokens)

And decoding — tokens back to text — is a lookup, with one famous subtlety:

vocab = {i: bytes([i]) for i in range(256)}          # 0-255 are the raw bytes
for pair, new_id in merges.items():                  # rebuild fused tokens
    vocab[new_id] = vocab[pair[0]] + vocab[pair[1]]

def decode(ids):
    raw = b"".join(vocab[i] for i in ids)
    return raw.decode("utf-8", errors="replace")     # a token can end MID-character!

print(decode(encode("Hello, émigrés!", merges)))

That errors="replace" is why broken model output sometimes shows <|rental|>-style boxes or ? characters: truncation split a multi-byte character in half. The tokenizer is not a string API — it's a byte API.

Why this matters more than beginners expect

In the next part we take these integer sequences and finally do what the "T" in GPT stands for: transform them. The whole series builds toward a model that writes text token by token — and every piece will be code you could write yourself.