Trust is earned, not given

A different perspective

2023-01-19 · Projects

AI Frontiers, part 1: Attention is all you need β€” reading the transformer paper five years later

Part 1from the AI Frontiers series · 65 parts in all

This is the first article in a series I am calling AI Frontiers. The premise is simple: the field of machine learning moved faster between 2017 and 2023 than in the two decades before it, and most working engineers β€” myself included β€” consumed that change through headlines, product launches, and Twitter threads rather than through the primary literature. The primary literature is where the actual ideas live. So in this series I am going to read the papers that marked real turning points, write down what they actually said, and check what has held up. Some entries will be about single papers; some will be about a cluster of work that landed together. All of them will be written from the desk of someone who builds software for a living, not from a research lab.

There is no better place to start than the paper whose title became a shrug and then a clichΓ©: Attention Is All You Need, published by eight researchers at Google Brain and Google Research in June 2017 (Vaswani et al.). Five years later, virtually every large language model you have heard of β€” GPT-4 included β€” is a descendant of the architecture described in its eleven pages. And yet the paper is widely misremembered. It did not invent attention. It did not, by itself, create large language models. What it did was more surgical and, in hindsight, more consequential: it removed the recurrent network from sequence modeling entirely and showed that a particular kind of attention mechanism, given enough width and depth, could do everything recurrence used to do β€” and do it in a way that parallelizes across whole sequences instead of walking them one token at a time.

The problem the paper was actually solving

To understand what changed, you have to remember what sequence modeling looked like in 2016. The dominant architectures for translation were recurrent neural networks β€” LSTMs and GRUs β€” arranged in encoder-decoder pairs, sometimes enhanced with attention. The recurrent setup has a structural property that sounds innocuous and is actually disastrous at scale: the computation for token n depends on the completed computation for token nβˆ’1. That dependency chain means you cannot process a sentence in parallel during training. With a GPU that can perform thousands of operations simultaneously, you are using it like a very expensive single-core processor, feeding tokens through one at a time, depth-first through time.

The 2016-era solution to the quality problem was to make the networks deeper and to bolt on attention, which had been introduced by Bahdanau, Cho, and Bengio in 2014 as a way for a translation model to look back at the entire source sentence when producing each output word, rather than squeezing the whole sentence through a single fixed-length vector (Bahdanau et al.). Attention worked so well that by 2016 the field was stacking more and more recurrence underneath it. The Attention Is All You Need authors asked the obvious-in-hindsight question: if attention is doing the heavy lifting, what is the recurrence for? Their answer, demonstrated empirically across four WMT translation benchmarks: nothing. Delete it.

What remains is the transformer: a stack of blocks, each containing multi-head self-attention β€” every token attends to every other token, with several distinct attention "heads" running in parallel so the model can track different relationships at once β€” followed by a small position-wise feed-forward network, with residual connections and layer normalization around both, and positional encodings injected at the bottom so the model can distinguish "the dog bit the man" from "the man bit the dog." Positional encoding is the detail people forget: self-attention itself is permutation-invariant. It sees a bag of vectors. The sin and cosine encodings the paper chose were, in the authors' words, chosen partly because they might extrapolate to sequence lengths longer than those seen in training β€” a claim whose complications would resurface years later as context windows grew.

Why removal mattered more than addition

The consequences of removing recurrence are easy to state and hard to overstate. First, training became parallel across the sequence dimension. A 512-token sentence is one big batch of matmuls instead of 512 sequential steps. Second, the architecture is essentially homogeneous: the same block, stacked. That sounds like a small aesthetic point, but homogeneous stacks are what GPUs and, later, dedicated accelerators eat for breakfast. The transformer is, almost accidentally, the perfect shape of model for the hardware the world was already mass-producing. There is a decent argument that the transformer won not because it was architecturally superior in every respect but because it was the best fit for the silicon it would run on β€” a lesson the field keeps relearning whenever someone proposes a clever new block that runs at half the speed.

Third β€” and this is the point that mattered most for what came later β€” self-attention gave every position direct access to every other position with a path length of one. In a recurrent network, information from token 1 must survive a chain of hundreds of intermediate states to influence token 500; gradients fade or explode along that chain. In a transformer, token 500 attends to token 1 directly. Long-range dependencies stop being a physics problem. The paper demonstrates this with a subject-verb agreement task where accuracy stays nearly flat as the distance between subject and verb grows, while the recurrent baselines degrade with distance (Vaswani et al.).

What the paper did not settle is anything about scale. The original model had 65 million parameters and trained on a few million sentence pairs in about three and a half days on eight GPUs β€” a serious but ordinary 2017 budget. The idea that you could take this architecture, remove the supervision (predict the next word instead of translating), and train it on a terabyte of internet text was not in the paper. That came from the line of work that runs through ELMo's deep contextual embeddings (Peters et al.), GPT's generative pre-training (Radford et al.), and BERT's bidirectional masking (Devlin et al.) β€” all of which arrived within a year and a half of the transformer paper and converted it from a translation model into a general-purpose text engine.

The vocabulary trap

Reading the paper in 2023 requires unlearning some of its own vocabulary, because words changed meaning on the way to the present. Two examples stand out.

"Transformer" then and now. The 2017 model is an encoder-decoder machine built for conditional generation: it reads a source sentence and emits a translation, with cross-attention connecting the two stacks. When people say "transformer" today, they usually mean a decoder-only stack like GPT β€” no encoder, no cross-attention, just one tower that predicts the next token. The paper itself contains an ablation table showing the decoder-only configuration underperforming the full model on translation β€” a result that was true for translation with 2017-scale data and irrelevant to what decoder-only models would become when given three orders of magnitude more data. If you read the paper expecting to find GPT in it, you will find it only in embryo.

"Attention" is several mechanisms wearing one name. The paper's multi-head self-attention is the headline, but the same word covers the cross-attention that connects encoder to decoder and the feed-forward layers that everyone forgets. It is worth being precise: in modern decoder-only models, the attention layers move information between positions (that is the communication), while the position-wise feed-forward networks process each position independently (that is the computation). Practitioners sometimes describe a transformer as tokens communicating via attention, then thinking via MLPs, in alternating rounds. Nothing in the 2017 paper frames it that way, but the components are all there.

What has aged well, and what has not

Five years on, most of the paper's choices look startlingly durable. Multi-head scaled dot-product attention survives essentially unchanged (the scaling factor, dividing by the square root of the key dimension, was justified in a one-line footnote about suspected dot products growing large in magnitude β€” that footnote became standard practice everywhere). Residual connections and layer normalization survive. The training recipes β€” Adam with a warmup schedule whose learning rate rises linearly for the first 4,000 steps and then decays with the inverse square root of the step number β€” are still recognizable inside modern training runs.

Some things aged poorly. Positional encodings are the clearest case: the sinusoidal scheme did not, in practice, extrapolate gracefully to much longer sequences, and the entire field of position-representation research β€” learned embeddings, relative encodings such as Transformer-XL's and T5's, RoPE's rotary embeddings, ALiBi's linear biases β€” exists because the original scheme hit a wall the moment people wanted 30,000-token contexts instead of 512-token sentences. Layer normalization placement moved: later work (Xiong et al.) showed that pre-norm β€” normalizing inside the residual branch rather than after it β€” trains deeply stacked models far more stably, and every serious large model uses pre-norm today even though the 2017 paper used post-norm. And the full encoder-decoder layout has largely lost to decoder-only for generative language models, though it remains the right tool for some translation and speech tasks.

The deepest thing that aged well is neither a component nor a trick. It is the paper's demonstration that architectural priors matter less than people believed, provided you replace them with something that lets data and compute flow. Before 2017, model design was heavily inductive: convolutions encoded locality, recurrence encoded order, and tree structures encoded syntax. The transformer bet that if you give a model direct access to all pairs of positions and let attention learn what matters, the data will teach the structure β€” order, locality, syntax, and all. That bet kept paying out at every scale increase since. It is the reason the field's core research question went from "what architecture?" to "what data, what objective, how much compute?" β€” which is where the rest of this series will spend most of its time.

What an application engineer should take from it

I build line-of-business software, not foundation models. The value of reading this paper carefully, for someone like me, is not the ability to implement it β€” plenty of libraries do β€” but the ability to reason about the systems we now depend on. Three practical takeaways survived my re-read.

First, quadratic attention cost is not an implementation detail; it is the load-bearing constraint of everything downstream. Because every token attends to every other, memory and compute grow with the square of sequence length, and every product decision about context windows, retrieval, chunking, and cost traces back to that fact. When a vendor advertises a 32K context and a competitor advertises 100K, the interesting question is never "which is bigger" but "what did they trade away" β€” in position encoding, in attention approximation, in training data β€” to get there.

Second, the parts of the architecture that do the "thinking" β€” the position-wise feed-forward layers β€” are also where most of the parameters live, which is why facts stored in a model are widely believed to live in those layers rather than in attention. That maps onto practical questions like fine-tuning (which I will take up in part 3, where LoRA's insight is precisely that you can leave most of those weights frozen) and onto why editing a model's knowledge is so much harder than editing a database.

Third, the history is a corrective against presentism. In January 2023, with ChatGPT dominating every feed, it is easy to imagine that the current moment was inevitable. It was not. The transformer spent its first year as a better translation model; it took a separate insight β€” unsupervised pre-training on raw text β€” plus a few years of scale to get to where we are. The next transformation is probably sitting in a paper published last month whose significance nobody, including its authors, fully appreciates yet. That is the working assumption this series will run on.

Works Cited

Bahdanau, Dzmitry, Kyunghyun Cho, and Yoshua Bengio. "Neural Machine Translation by Jointly Learning to Align and Translate." arXiv.org, 2014, arxiv.org/abs/1409.0473. Accessed 19 Jan. 2023.

Devlin, Jacob, et al. "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding." arXiv.org, 2018, arxiv.org/abs/1810.04805. Accessed 19 Jan. 2023.

Peters, Matthew E., et al. "Deep Contextualized Word Representations." Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, ACL, 2018, aclanthology.org/N18-1202. Accessed 19 Jan. 2023.

Radford, Alec, et al. "Improving Language Understanding by Generative Pre-Training." OpenAI, 2018, openai.com/research/improving-language-understanding-by-generative-pre-training. Accessed 19 Jan. 2023.

Vaswani, Ashish, et al. "Attention Is All You Need." arXiv.org, 2017, arxiv.org/abs/1706.03762. Accessed 19 Jan. 2023.

Xiong, Ruibin, et al. "On Layer Normalization in the Transformer Architecture." Proceedings of the 37th International Conference on Machine Learning, PMLR, 2020, proceedings.mlr.press/v119/xiong20b.html. Accessed 19 Jan. 2023.