AI fundamentals, part 1: From bytes to BPE — how language models read text
Part 1from the AI fundamentals series · 8 parts in all
This begins a series on how large language models actually work, written for programmers who know Python but not the internals of AI. We follow the same arc as the university courses on language modeling — tokenization, architecture, training, scaling, inference — but build every piece ourselves in small, readable Python. No frameworks, no copied code. Part 1: how raw text becomes numbers.
Why tokenization exists at all
A neural network computes with floating-point vectors. It has no concept of "text". So the
first question any language model must answer: what are the units? Characters are too
many steps apart (a 20-token thought becomes 20+ inputs); words are too many distinct things
(English alone has hundreds of thousands, plus every typo ever typed). The industry compromise
is the subword: chunks like un, believ, able
— a few hundred thousand of them, covering everything with no out-of-vocabulary failures.
Step 1: everything is bytes
Modern models (GPT-2 onward, Llama, Mistral) start by treating text as raw UTF-8 bytes — 256 possible values, guaranteed to handle any language or emoji without special cases:
text = "héllo world"
raw = text.encode("utf-8") # b'h\xc3\xa9llo world'
ids = list(raw) # [104, 195, 169, 108, 108, 111, ...]
print(len(text), len(ids)) # 11 characters -> 12 bytes (é is two bytes)
Byte-level is the floor: a model trained this way can ingest any file ever written.
The cost is efficiency — common words like the would burn 3 tokens instead of 1.
That's what the next step fixes.
Step 2: BPE — compression by counting
Byte-Pair Encoding is greedy data compression applied to language: count
which adjacent pair of symbols occurs most often in your training corpus, merge that pair into
a new symbol, repeat vocab_size - 256 times. Frequent sequences become single
tokens; rare text stays as plain bytes.
def get_pair_counts(ids):
# Count every adjacent pair in the id sequence.
counts = {}
for a, b in zip(ids, ids[1:]):
counts[(a, b)] = counts.get((a, b), 0) + 1
return counts
def merge(ids, pair, new_id):
# Replace every occurrence of `pair` in `ids` with `new_id` (left to right).
out, i = [], 0
while i < len(ids):
if i + 1 < len(ids) and (ids[i], ids[i+1]) == pair:
out.append(new_id) # fused token
i += 2
else:
out.append(ids[i]) # untouched token
i += 1
return out
# --- training ---
ids = list("the cat sat on the mat the end".encode("utf-8"))
merges = {} # (a, b) -> new_id
next_id = 256
for step in range(10): # real models do ~50k-100k of these
pairs = get_pair_counts(ids)
top = max(pairs, key=pairs.get) # most frequent adjacent pair
merges[top] = next_id
ids = merge(ids, top, next_id)
next_id += 1
print("t h merged early because 'th' is everywhere in English")
After training, encoding new text is just replaying the merges in order of priority:
def encode(text, merges):
ids = list(text.encode("utf-8"))
for pair, new_id in sorted(merges.items(), key=lambda kv: kv[1]): # learned order
if pair in zip(ids, ids[1:]): # this merge still applies
ids = merge(ids, pair, new_id)
return ids
print(encode("the", merges)) # short! 'the' got fused during training
print(encode("thé", merges)) # 'é' is non-ASCII: stays raw bytes (2 tokens)
And decoding — tokens back to text — is a lookup, with one famous subtlety:
vocab = {i: bytes([i]) for i in range(256)} # 0-255 are the raw bytes
for pair, new_id in merges.items(): # rebuild fused tokens
vocab[new_id] = vocab[pair[0]] + vocab[pair[1]]
def decode(ids):
raw = b"".join(vocab[i] for i in ids)
return raw.decode("utf-8", errors="replace") # a token can end MID-character!
print(decode(encode("Hello, émigrés!", merges)))
That errors="replace" is why broken model output sometimes shows
<|rental|>-style boxes or ? characters: truncation split a
multi-byte character in half. The tokenizer is not a string API — it's a byte API.
Why this matters more than beginners expect
- Cost and context limits are counted in tokens. A "4,000-token context" is roughly 3,000 words of English — but far fewer of CJK text or base64 blobs, which tokenize densely into bytes.
- Tokenization causes model "stupidity". Ask a model how many
rs are in "strawberry" and it can miscount — it never saw letters, it saw token IDs. Arithmetic struggles partly trace to digit tokenization too. - Security lives here. Token boundaries can hide prompt-injection payloads; unusual unicode splits in surprising ways. Anyone building on LLMs should own a tokenizer.
In the next part we take these integer sequences and finally do what the "T" in GPT stands for: transform them. The whole series builds toward a model that writes text token by token — and every piece will be code you could write yourself.