Trust is earned, not given

A different perspective

2023-07-06 · Projects

AI Frontiers, part 29: The open-weight leap — LLaMA, Alpaca, and the leak that built an ecosystem

Part 29from the AI Frontiers series · 65 parts in all

On 24 February 2023 Meta published LLaMA, a family of language models from 7B to 65B parameters, alongside a paper whose argument was specifically about efficiency: a model trained on more tokens than its size would suggest could outperform much larger models, and the whole thing could be trained on publicly available data (Touvron et al.). Access was gated. Researchers applied, and approved applicants received weights under a research-only license that prohibited commercial use.

Within about a week, the weights were circulating as a torrent. Within a month there were several fine-tuned derivatives, one of them produced for a reported $600. By the summer, every few weeks brought another model — some genuinely open, some open in the sense that the weights were downloadable and the story about the data was not. What happened between February and July of 2023 is the origin story of the open-weight ecosystem, and it is worth understanding as two separate events: a technical result that made small models viable, and a failed distribution strategy whose failure changed the industry.

Why LLaMA was the right model at the right moment

The technical contribution was a scaling decision. Chinchilla had established that most models of the period were badly undertrained — too many parameters for the number of tokens they saw — and that for a fixed compute budget you should train a smaller model on substantially more data (Hoffmann et al.). LLaMA applied that lesson directly: 1.4 trillion tokens, most of it publicly available text, with the 13B model reported to outperform the 175B GPT-3 on most of the benchmarks the paper evaluated, and the 65B model competitive with much larger systems.

Two consequences followed immediately. First, the size of a useful model collapsed. A model that fits on one or two accelerators and runs in a few gigabytes at reduced precision is a model that a university group, a startup or a determined hobbyist can actually work with. Second — and this is the part that mattered more for what came next — the demonstration that competitive results were achievable on public data removed the most convenient excuse for keeping weights closed. If the recipe is published techniques applied to a scraped corpus, then the contribution is execution and compute rather than data monopoly, and the case for a download gate gets much weaker.

LLaMA was not the first open model. BLOOM (BigScience), OPT (Meta), GPT-NeoX-20B (EleutherAI) and Pythia had all been released with open weights and, in several cases, open training data and code. What LLaMA added was a quality bar that made the open models serious competitors rather than honorable experiments, and a distribution method that guaranteed attention by trying to prevent it.

The leak, and the license that stopped meaning anything

The story of how the weights leaked is unremarkable: someone posted a magnet link, mirrors appeared, and the file propagated. The interesting part is what it did to the license.

Meta's terms for LLaMA restricted use to research, prohibited commercial deployment, and required that redistribution follow the same terms. Those terms could be enforced against an approved researcher who violated them. They could not be enforced against the tens of thousands of people who obtained the weights from a torrent with no relationship to Meta at all. The license did not disappear; it simply became ineffective for the population that actually had the files, which is a failure mode worth remembering for anyone designing a controlled release. A gate that holds against honest actors and leaks against everyone else produces exactly the wrong outcome: compliance costs borne by the well-behaved, with none of the control the gate was meant to buy.

Two months later Meta shipped Llama 2 with weights downloadable by anyone, a permissive license permitting commercial use with a carve-out for very large platforms, and a written argument that open release makes models safer because more eyes are on them. It is impossible to prove which consideration drove the change, and the sequence — gate the weights, watch them leak, then open them deliberately — is a reasonably strong hint that the experiment had already been run for them.

The ecosystem, month by month

What followed was the fastest expansion of a model ecosystem the field has seen, and the sequence has a clear internal logic.

Alpaca (Stanford, March) applied Self-Instruct to LLaMA-7B: 52,000 instructions generated by a larger model, fine-tuning costing under $600, and an assistant-like model produced from a base model in hours (Taori et al.). The specific capability was modest. The demonstration — that instruction-following could be manufactured cheaply on top of a leaked base model — was not.

Vicuna (LMSYS, March) fine-tuned LLaMA-13B on conversations collected from users of an early chat interface and reported that a GPT-4-based judge preferred its outputs to other open models in the large majority of comparisons. Those judgments are weak evidence, as the hallucination entry implies and as everyone learned later — a language model grading its own kind is a proxy with a measurable bias — but the perception of parity with commercial assistants was what moved the field.

Koala (Berkeley, April) trained on web-collected dialogue data and made the point explicit in its title: it advertised itself as a model tuned on public data rather than model-generated data, arguing that distillation from a commercial model is a different thing from imitating human conversation. Databricks answered the licensing question the same month with Dolly 2.0, which released an instruction-following model, the 15,000 training examples, and the data-generation process — the first genuinely open instruction-tuned model, because every component of it could be reproduced without a gate.

RedPajama (Together, April) went after the data instead. LLaMA's paper described a 1.2-trillion-token corpus built from public sources; RedPajama published an open reproduction of that recipe, which meant the final unreproducible component was gone. MPT-7B (MosaicML, May) followed with an Apache-licensed 7B model trained on a trillion tokens on a documented budget, and StarCoder (BigCode, May) released a 15.5B code model trained on permissively licensed source code from 80 languages with an opt-out mechanism for developers (Li et al.). By June there were enough models that the interesting question had shifted from "can we do this?" to "which one, and on what terms?"

Open weights, open data, and why the distinction matters

The most useful idea to come out of this period is that "open" is not a binary, and the component that matters most is usually not the weights.

A model is reproducible only if you can obtain the weights, the training data, the code, and enough documentation of the pipeline to know what was done. Of those four, the data is the one that is essentially never available for scraped corpora, because nobody knows precisely what is in them, and in many cases nobody is legally permitted to redistribute them. That is why EleutherAI's Pythia suite (Biderman et al.) mattered more than its benchmark positions: 16 models spanning sizes with 154 training checkpoints each, all released, explicitly intended as an instrument for studying how models change during training rather than as a product. Open does not always mean competitive, and for research the second is not the point.

The licensing layer underneath all of this was, and remains, inconsistent in ways that surprise people. BLOOM shipped with the Responsible AI License with use restrictions attached. Dolly 2.0 released its data under a permissive creative-commons license. StarCoder's training data handled code licensing through an opt-out for rights holders. Lloyd's LLaMA license restricted commercial use; Llama 2's permitted it with a large-scale carve-out. Two models that both say "open weights" can carry obligations that differ in every dimension a company cares about, and the practical advice for anyone building on open models has not changed since: read the license, and assume you will eventually have to explain your provenance decisions to somebody who is not sympathetic.

Arena ratings, and why the judge mattered as much as the models

The evaluation problem this period created is worth dwelling on, because the claims being made were strong and the instruments were weak. Vicuna's headline result came from a language model judging conversations, and the paper documenting that methodology (Zheng et al.) is candid about the reasons to distrust it: judges favour longer answers, prefer their own family's style, are sensitive to the order in which two responses are presented, and are unreliable on questions requiring reasoning or mathematics. Those are not fatal flaws. They are the ordinary biases of a proxy, and the previous entry is a long argument about what happens when you optimize against one without measuring the gap.

The chatbot arena, launched in May, was the field's attempt to route around benchmark contamination and self-judging bias at once. Anonymous pairs, real user votes, an Elo rating computed over thousands of comparisons. It is a better instrument than a static benchmark and it introduced its own distortions — a pleasant style wins votes that a correct answer might not, and popularity among a self-selected audience is not quality. What it did decisively was make the comparison public and continuous, which changed the incentives for everyone releasing a model. Once there is a live scoreboard, being unmeasurable stops being an option.

What the leak taught about distribution

Three lessons from this period generalize well beyond models.

A gate you cannot enforce is worse than no gate. Meta's research license was honored by the people who asked for permission and irrelevant to everyone else, which meant the compliance burden landed entirely on the actors who were already behaving. If you cannot constrain access, controlling the terms of access mostly selects for naivety.

Weight release is irreversible and cheap to copy. There is no version of this where a released model's use can be audited at scale, and the second and third derivative models have no relationship with the original publisher at all. Any governance approach that assumes otherwise is building on sand, which is why the serious proposals in this area focus on the training compute and on deployment practices rather than on the file.

Openness is a strategy, not a virtue. Releasing weights converts a capability into an ecosystem, and it also converts a research investment into a commodity. Some labs are willing to make that trade because it buys talent, adoption and goodwill; some are not, because their business depends on the capability remaining scarce. Both are coherent positions, and the parties holding them were decided in the months covered here.

What an open ecosystem actually decided

Four things were settled in these months, and they set the terms for everything that followed.

Fine-tuning became a commodity. Once a competent 7B base model was freely available, the marginal cost of a specialized model dropped to a rented GPU and a dataset. That is the direct precondition for LoRA and QLoRA being interesting, and for the specialization strategies that fill the rest of this series.

Independent evaluation became necessary and inadequate. The LMSYS chatbot arena, launched in May, was the right instinct: let real users vote on anonymous pairs and compute an Elo rating, rather than trusting a benchmark that models may have seen. The Hugging Face Open LLM Leaderboard supplied the other half, a standardized public comparison. Both immediately ran into the problem of comparing systems on tasks that had become training targets, which is the subject of a later entry.

The moat question stopped being rhetorical. A memo reportedly circulated inside Google in May arguing that neither Google nor OpenAI had a durable advantage over open models, because the techniques diffuse within months and the cost of reproduction keeps falling. Whatever one thinks of the conclusion, the premise was correct and this period is where it became visible. A closed lab's advantage is capability and pace, not secrecy.

Regulation became inevitable. Once capable weights were downloadable by anyone, the question of who is responsible when they are misused became concrete rather than theoretical, and no technical mechanism — licenses, gates, use policies — had proven capable of constraining it. That is the argument that eventually produced legislation, and it started here, with a torrent of files that a research license was never going to contain.

Works Cited

Biderman, Stella, et al. "Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling." arXiv, 2023, arxiv.org/abs/2304.01373. Accessed 6 July 2023.

Black, Sidney, et al. "GPT-NeoX-20B: An Open-Source Autoregressive Language Model." arXiv, 2022, arxiv.org/abs/2204.06745. Accessed 6 July 2023.

Databricks. "Introducing Dolly 2.0." Databricks Blog, 12 Apr. 2023, www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm. Accessed 6 July 2023.

Hoffmann, Jordan, et al. "Training Compute-Optimal Large Language Models." arXiv, 2022, arxiv.org/abs/2203.15556. Accessed 6 July 2023.

Köpf, Andreas, et al. "OpenAssistant Conversations — Democratizing Large Language Model Alignment." arXiv, 2023, arxiv.org/abs/2304.07327. Accessed 6 July 2023.

Li, Raymond, et al. "StarCoder: May the Source Be with You!" arXiv, 2023, arxiv.org/abs/2305.06161. Accessed 6 July 2023.

MosaicML. "Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs." MosaicML Blog, 5 May 2023, www.mosaicml.com/blog/mpt-7b. Accessed 6 July 2023.

Scao, Teven Le, et al. "BLOOM: A 176B-Parameter Open-Access Multilingual Language Model." arXiv, 2022, arxiv.org/abs/2211.05100. Accessed 6 July 2023.

Taori, Rohan, et al. "Alpaca: A Strong, Replicable Instruction-Following Model." Stanford Center for Research on Foundation Models, 13 Mar. 2023, crfm.stanford.edu/2023/03/13/alpaca.html. Accessed 6 July 2023.

Together Computer. "RedPajama: An Open Source Recipe to Reproduce LLaMA Training Dataset." GitHub, 2023, github.com/togethercomputer/RedPajama-Data. Accessed 6 July 2023.

Touvron, Hugo, et al. "LLaMA: Open and Efficient Foundation Language Models." arXiv, 2023, arxiv.org/abs/2302.13971. Accessed 6 July 2023.

Zhang, Susan, et al. "OPT: Open Pre-trained Transformer Language Models." arXiv, 2022, arxiv.org/abs/2205.01068. Accessed 6 July 2023.

Zheng, Lianmin, et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv, 2023, arxiv.org/abs/2306.05685. Accessed 6 July 2023.