Trust is earned, not given

A different perspective

2023-07-20 · Projects

AI Frontiers, part 4: Llama 2 and the open-weights turning point

Part 4from the AI Frontiers series · 65 parts in all

In part 3 I described the May 2023 moment when fine-tuning came to the people: LoRA and QLoRA made adaptation cheap, and a chaotic ecosystem of LLaMA derivatives β€” Alpaca, Vicuna, Guanaco β€” grew up around a base model that had, strictly speaking, leaked. That ecosystem ran on legal fog. The weights were out, but nobody building anything serious could sign off on a derivative of a leak. Two days ago, on July 18, Meta ended the fog with Llama 2: officially released, commercially licensed, free, and good. This entry is about why I think the license, not the benchmark scores, was the actual event β€” and about what the release tells us about where the open-weights world goes from here.

What shipped

Llama 2 came in three sizes β€” 7B, 13B, and 70B parameters β€” each in two flavors: a base pre-trained model and a chat-tuned variant. The base models replace LLaMA's training recipe with roughly double the tokens (2 trillion, against GPT-3's ~300 billion and LLaMA's ~1 trillion), a longer 4,000-token context, and grouped-query attention in the 70B to cut inference memory. The chat models are the more interesting artifact scientifically: they were built with the pipeline the LLaMA 2 paper describes as standard for assistants β€” supervised fine-tuning on human-written examples, followed by reinforcement learning with human feedback (RLHF), using both the PPO algorithm and a rejection-sampling variant (Touvron et al.). This was the first time a major lab documented its alignment pipeline in that much detail for an openly available model, reward model annotations included, and reading that section remains the fastest way to understand what "RLHF" concretely means as engineering rather than acronym.

The benchmarks landed the 70B chat model in GPT-3.5 territory β€” clearly behind GPT-4, competitive with or ahead of other open alternatives, with the paper itself noting that GPT-4 still leads substantially. The human preference evaluations put Llama-2-70B-chat winning roughly 36% of comparisons against ChatGPT (with ~32% ties) β€” a genuinely remarkable number for a freely downloadable model, whatever you think about human-eval noise. On the safety side, the paper is unusually specific: red-teaming across seven risk categories, safety-specific fine-tuning stages, and context-distilled refusal triggers. For the first time, an open release carried a safety report whose structure mimicked the closed labs' system cards.

The license was the product

Here is the part I under-appreciated until I read it three times: the technical release mattered less than the legal instrument wrapped around it. Meta's Llama 2 Community License grants commercial use β€” the thing every leaked-weights derivative could not offer β€” with two notable strings. First, an acceptable-use clause barring a specific list of misuses. Second, and famous within days: the moat clause. Any entity using the models "in connection with" a service with over 700 million monthly active users needs a separate license from Meta. That number is not subtle. It is a carve-out shaped exactly like Google, and the message was read that way everywhere: the open-weights play is a competitive weapon against closed rivals, not an act of uncompensated generosity.

Also note what open means here, because the distinction ran through every discussion that month. Llama 2 is open-weights, not open-source: training data undisclosed, training code unavailable, license terms attached, and the OSI β€” protective of its definition β€” pointedly declined to endorse it. A purist could reasonably call it a free-to-use download with rules. But the functional difference from the leaked LLaMA was night and day: enterprises, cloud platforms, and startups could adopt it without counsel reaching for the antacids. Within days, Azure's model catalog, AWS Bedrock, and Google's Vertex all announced hosted Llama 2 β€” each cloud hedging its closed-model bets with the open alternative, exactly as the license intended. The ecosystem effect was less like a paper publication and more like a platform launch, because that is what it was.

The benchmark wars that followed

The weeks after release produced a benchmark fight worth remembering, because it teaches something durable about how we evaluate models. The infrastructure of evaluation had itself just matured: Hugging Face's Open LLM Leaderboard had launched in June 2023 as a reproducible harness for static benchmarks (ARC, HellaSwag, MMLU, TruthfulQA), AlpacaEval offered GPT-4-as-judge scoring of instruction-following, and the LMSYS Chatbot Arena β€” profiled in part 3 β€” collected head-to-head human votes and rated models by Elo. Llama 2 variants flooded all three within days, and the three instruments immediately disagreed in informative ways: static-leaderboard scores that looked stellar translated into middle-of-the-pack Arena Elo once real users voted; AlpacaEval deltas swung with answer verbosity, exactly the judge bias part 3 warned about. Community quantizers such as TheBloke published 4-bit GGUF and GPTQ builds of every variant within hours, and llama.cpp ran them on laptops β€” adoption at a speed no closed model has ever matched.

The durable lesson: after GPT-4, no single benchmark number moves the market. Triangulation does β€” human arena votes, static benchmarks, and task-specific evals disagreeing informatively with each other. The Arena became the de-facto referee precisely because it was the hardest instrument to game, and its post-release picture was consistent: Llama-2-70B-chat solidly in GPT-3.5 territory, GPT-4 still a class apart.

The 7B and 13B tier: where the economic damage happens

The 70B chat model took the headlines, but I expect the smaller tiers to do more economic damage, and the first 48 hours already show why. The 7B and 13B bases cost nothing to download, and the part 3 pipeline β€” QLoRA on a consumer card, merge, quantize, serve β€” applies to them unchanged. TheBloke's 4-bit GGUF and GPTQ builds were on Hugging Face within roughly a day of release, llama.cpp ran them on laptops the same evening, and Hugging Face's transformers supported the architecture out of the box. The parts for the fine-tune explosion are all on the shelf; the only remaining ingredient is calendar time. My working prediction, for the record: by autumn there will be thousands of public Llama 2 derivatives β€” instruction-tuned, domain-tuned, merged β€” most of them worse than the base outside their niche, a few of them defining new task categories. The option is what changes procurement conversations: "build on the open base" stops being a fringe position the moment the license lets counsel sign.

The first days: the self-hosting math

The practical question every team is asking this week: does self-hosting beat the API? The arithmetic has three inputs, and two days is enough to see their shape. Throughput: vLLM β€” the open serving stack with paged KV-cache management that shipped in June β€” runs the Llama architecture unchanged, and a single A100 holds dozens of concurrent sessions; quantized 7B and 13B builds already run on single consumer GPUs. Cost: at low volume, the API wins on total cost of ownership β€” always, because nobody pays themselves back for GPU operations at three requests a minute. At sustained volume, the crossover arrives fast: a busy 13B on one rented GPU undercuts per-token API pricing by roughly an order of magnitude. Latency and data residency: the unpriced deciders β€” p50 latency you control, and documents that never leave your VPC, which closed deals no benchmark could. The pattern I expect to emerge β€” open weights for the steady high-volume middle of the workload, frontier APIs for the hard edges β€” is the hybrid architecture part 9 makes explicit.

The first weeks: the safety disclosures

The Llama 2 paper's safety section is the most detailed public account of an assistant alignment pipeline to date, and it is worth reading as engineering rather than PR. The pipeline runs in two stages: safety-specific supervised fine-tuning built from adversarial prompt collections gathered by red teams across enumerated risk categories, followed by RLHF with both helpfulness and safety reward models, plus context distillation so refusal behavior generalizes. The paper reports the resulting safety evaluations per risk category and, unusually, discusses the helpfulness–safety trade-off explicitly β€” more safety training measurably degrades helpfulness on benign prompts.

Two critiques will stick, and one is already visible. First, the safety metrics are self-reported β€” same as closed labs, but open claims invite open audits, and independent evaluators will find both over-refusal on benign prompts and remaining jailbreak surfaces within weeks. Second, and more structurally: open weights make safety measurable by anyone. Whatever Meta's internal numbers say, the community can β€” and already does β€” run its own refusal suites, bias probes, and attack collections against the exact shipped weights. That inversion, from trust-the-vendor to test-the-artifact, is the open ecosystem's real safety contribution, and it is worth more than any single model's benchmark row.

A month in production: the self-hosting math

The practical question every team asked in August: does self-hosting beat the API? The arithmetic had three inputs. Throughput: vLLM β€” the open serving stack with paged KV-cache management that had shipped in June β€” made a single A100 hold dozens of concurrent Llama 2 sessions, and quantized builds (TheBloke's GGUF and GPTQ exports were up within hours of release) brought the 7B and 13B tiers to single consumer GPUs and llama.cpp on laptops. Cost: at low volume, the API wins on total cost of ownership β€” always, because nobody pays themselves back for GPU operations at three requests a minute. At sustained volume, the crossover arrives fast: a busy 13B on one rented GPU undercuts per-token API pricing by an order of magnitude. Latency and data residency: the unpriced deciders β€” p50 latency you control, and documents that never leave your VPC, closed deals that no benchmark number could. The pattern that emerged β€” open weights for the steady high-volume middle of the workload, frontier APIs for the hard edges β€” is the hybrid architecture part 9 will make explicit.

What it changed β€” and what it did not

For builders, the July release changed the default answer to a question every team had been punting: "can we self-host?" Before July, the answer required legal gymnastics; after July, it was a yes with an asterisk. The pattern that emerged over the following weeks β€” open weights for baseline capability and cost control, closed APIs for the frontier capability you cannot match β€” is now the standard architecture of the industry, and it was effectively ratified by this release plus its cloud distribution deals.

What did not change: the frontier. GPT-4 remained a class apart, and the release's own evaluations were honest about the gap. The open-weights ecosystem had become an industry, but it remained a fast-follower industry β€” adopting architectures and techniques the closed labs proved first, usually within months. That dynamic β€” frontier in the closed world, diffusion through the open world β€” has held ever since, through Mistral's September release and everything after. It is the single most useful mental model I can offer for predicting what the open ecosystem will look like next year: look at what closed models did last year.

A note on how little changed β€” deliberately

One more observation before closing, because it surprised me on the re-read: almost everything technically novel in Llama 2 was already in circulation. The architecture is LLaMA 1's β€” RMSNorm, SwiGLU, rotary embeddings, grouped-query attention in the 70B being the one addition. The alignment pipeline is the industry-standard SFT-plus-RLHF recipe, documented with unusual candor but not invented here. The training-scale insight β€” two trillion tokens beats one trillion β€” is the scaling-laws result from part 2's lineage, applied with discipline. Even the safety stack assembles known parts. What was genuinely new was the distribution decision, and that is the deepest lesson of the release: by July 2023, the capability gap between what labs could do and what labs would publish had become the binding constraint on the whole field. Meta did not need a research breakthrough to reshuffle the industry β€” releasing existing capability under a signature- ready license was sufficient. The open-weights story that follows, from Mistral's Apache- 2.0 entry (part 8) to the small-model commoditization of part 9, is downstream of that single decision.

It also sets up the question the rest of this series keeps circling: when distribution, rather than research, is the bottleneck β€” what does that do to the pace of capability change? The next two years suggest an answer: the pace of diffusion accelerates while the pace of frontier progress stays gated by compute and data. Both halves of that sentence were visible, if you knew where to look, in the forty-eight hours after July 18.

Works Cited

OpenAI. "GPT-4 Technical Report." arXiv.org, 2023, arxiv.org/abs/2303.08774. Accessed 20 July 2023.

Touvron, Hugo, et al. "Llama 2: Open Foundation and Fine-Tuned Chat Models." arXiv.org, 2023, arxiv.org/abs/2307.09288. Accessed 20 July 2023.

"Introducing Llama 2: The Next Generation of Open Source Large Language Models." Meta AI, 18 July 2023, ai.meta.com/blog/lama-2-updates/. Accessed 20 July 2023.

Gerganov, Georgi. "llama.cpp." GitHub, 2023, github.com/ggerganov/llama.cpp. Accessed 20 July 2023.

"vLLM: Efficient Memory Management for Large Language Model Serving." GitHub, 2023, github.com/vllm-project/vllm. Accessed 20 July 2023.

Chiang, Wei-Lin, et al. "Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference." LMSYS Org, 2023, lmsys.org/blog/2023-05-03-arena.html. Accessed 20 July 2023.

Zhou, Chunting, et al. "LIMA: Less Is More for Alignment." arXiv.org, 2023, arxiv.org/abs/2305.11206. Accessed 20 July 2023.

Hu, Edward J., et al. "LoRA: Low-Rank Adaptation of Large Language Models." arXiv.org, 2021, arxiv.org/abs/2106.09685. Accessed 20 July 2023.