AI Frontiers, part 37: The long tail of licensing — what open weights actually permit
Part 37from the AI Frontiers series · 65 parts in all
Two models can both be described as "open" and differ in ways that determine whether a company can ship with them. One permits commercial use without limit; another permits it below a monthly user threshold and requires a separate agreement above it. One has no acceptable-use restrictions beyond the law; another prohibits specified applications and requires downstream users to pass those restrictions on. One ships with training data that was redistributed alongside the model; another ships weights and a description of data nobody can reproduce.
By early 2024 this was no longer an academic distinction. Models were being embedded into products, procurement teams were asking for the terms, and the answers were a mess of custom licenses, use-restriction addenda, and documentation gaps. This entry is a survey of what the licenses actually say, why the ecosystem ended up this way, and which parts of the problem are genuinely unresolved rather than merely unfamiliar.
The shapes a model license takes
Four archetypes cover almost everything in circulation.
Permissive open-source licenses. Apache 2.0 is the clearest case: use, modify, redistribute, commercial deployment, no thresholds, no use restrictions, with a patent grant and attribution requirements that are familiar to any legal department. A model released this way is the easiest to adopt because the review is trivial. The tradeoff for the publisher is that there is no mechanism to restrict anything downstream, which is precisely why some labs avoided it.
Community licenses with thresholds. Meta's approach with Llama 2 permits commercial use and adds a clause requiring a separate license above a very large monthly active user count, plus acceptable-use terms that flow down to users and a requirement to display attribution. The threshold is written to be irrelevant to almost every developer and relevant to the handful of companies that could compete with the publisher, which is a coherent business position and not a statement about openness.
Open licenses with use restrictions. The Responsible AI License family, applied to BLOOM and later to other models, adds behavioral restrictions — no medical advice, no legal advice, no generating disinformation — and requires that the restrictions propagate. These are the most confusing to adopt because the restrictions are behavioral rather than technical, so compliance depends on what your application does with the model rather than on what you do with the weights. A lawyer evaluating a chatbot that might answer a health question has to reason about a clause that a weights-only analysis never surfaces.
Research-only releases. Weights available under terms that prohibit commercial use altogether, often with a gated download requiring an application. These are common for models whose training data could not be licensed for commercial use, and the honest reading is that the publisher is telling you the weights are commercially radioactive rather than merely restricted.
The data problem underneath every model license
Every license above governs the weights. None of them resolves the question that actually creates legal exposure, which is where the training data came from and what the publisher was entitled to do with it.
For models trained on a web scrape, three separate questions arise. Was the collection lawful under the terms of the sites it came from? Was the use of copyrighted material as training data permitted, whether by fair use, by an exception, or by some other doctrine? And does the resulting model incorporate or reproduce protected expression in a way that creates liability for whoever deploys it? Those are questions for courts, and in early 2024 the courts had not answered them — a set of lawsuits against model developers and image-generation companies were in early stages, and the field's working assumption that training on publicly available text is permissible had not been tested to judgment.
The provenance problem is worse than the legal one in the short run because it is an information problem. Nobody can say precisely what is in a trillion-token corpus assembled from crawls over several years. The Data Provenance Initiative (Longpre et al.) audited thousands of publicly available datasets and found pervasive inconsistencies in licensing, attribution and documentation — including datasets whose stated license did not match their contents, and datasets with no discoverable provenance at all. Anyone trying to build a model from "open" data discovers that the openness stops at the point where you need to know what you have.
This is why some models are research-only: the publisher concluded that it could defend training on the corpus but did not want to extend a warranty to commercial users. It is also why several releases in this period carefully avoided any claim about the data, describing it in the vaguest legally survivable terms. That is not evasion for its own sake; it is the only accurate statement anyone could make.
The documentation layer, which should exist and mostly does not
The infrastructure for describing datasets and models existed years before the models did. Datasheets for Datasets (Gebru et al.) proposed a structured questionnaire covering motivation, composition, collection process, preprocessing, uses and distribution. Model Cards (Mitchell et al.) did the same for models, with a section for intended use and evaluation results broken out by demographic group. Both are cited constantly and followed inconsistently, usually in the direction of listing an evaluation section and omitting everything uncomfortable.
What a useful model card should contain, and what most omitted in this period: the composition of the training data by source and by license category; the languages and domains covered and the ones that are systematically underrepresented; evaluation results broken out by language and by sub-population rather than only in aggregate; known failure modes including the ones the publisher observed during development; the compute used, since that is a governance-relevant fact; and an explicit statement of what the model was not evaluated for.
The last item is the most valuable and almost universally absent. A statement that a model was never tested on medical questions is more useful to a downstream developer than any benchmark score, because it tells them exactly where the unmeasured risk sits. The publishing incentives run the other way, which is why the field's best documentation practices have mostly been adopted by academic releases rather than commercial ones.
What businesses actually did about it
The practical response across organizations I have watched is a four-step review that has become more or less standard.
First, classify the license: permissive, threshold-based, use-restricted, or research-only. That determines whether the conversation ends immediately.
Second, map the use restrictions onto your actual product surface, not onto your intentions. If the license prohibits medical advice, does your system answer questions that a reasonable person would call medical? A restriction that your product arguably violates is not cured by the fact that you did not mean to violate it.
Third, ask for the data provenance and accept that the answer will often be incomplete. Where it is incomplete, the decision becomes a risk appetite question — how much exposure is acceptable for this feature — which is a business judgment rather than a legal conclusion.
Fourth, instrument for the thing that would change the answer. A lawsuit that establishes a rule about training data, a publisher that changes its terms, a restriction that turns out to be enforced — each of these would invalidate an earlier approval, and the only defense is knowing which models are deployed where so the answer can be revisited. Most organizations cannot answer that question, which is the same gap that makes incident response hard in every other domain.
Reading a license in ten minutes
Most of the work in a model review can be done quickly if you ask the right five questions in the right order.
First, what happens above a scale threshold? A clause that requires a separate agreement past some number of monthly users is irrelevant to a startup and existential to a platform, and it is usually well hidden. Second, is it a license or a contract? A license purports to grant permissions under intellectual property law; a contract imposes obligations, and the two have different remedies — which matters when a term is unenforceable, because the contract may fail in ways the license does not. Third, do the acceptable-use restrictions flow down to my users? Many licenses require redistribution of the restrictions, which means your own terms of service have to incorporate them, and most products never do.
Fourth, what am I required to disclose? Attribution requirements range from nothing to displaying the model name prominently in the user interface, and the second kind is a product requirement that often reaches engineering late. Fifth, what warranties are disclaimed? For model licenses, all of them, which is worth knowing precisely because it means the liability for a bad output is entirely yours.
A sixth question is worth asking even though it is not in the license: what would make this approval wrong? If the answer is a court ruling, a change in the publisher's terms, or the publication of data provenance that contradicts the model card, then you have a monitoring requirement, and writing it down is the difference between a documented risk and an unmanaged one.
The practice of compliance, which is smaller than it sounds
For all the complexity above, the operational burden for a typical team is modest if it is handled as engineering rather than as legal negotiation.
Keep a registry: which model, which version, which license, which data provenance claim if any, where it is deployed, who approved it and on what date. This is a table, and it is the single artifact that makes every subsequent question answerable. When a publisher changes terms, when a lawsuit produces a ruling, when a customer asks, the registry turns a fire drill into a filter.
Automate the notices. Attribution and restriction-carry-through obligations are mechanical, and mechanically satisfying them is far cheaper than discovering them during due diligence. The same goes for the data-provenance disclosures that regulation was beginning to require: if the obligation is to publish a summary of training content, then having assembled that summary early is the difference between a document and a crisis.
And read the license at the moment of adoption rather than at the moment of scale. The failure mode I have seen most often is a model adopted casually under a research license, embedded into a production feature, and discovered a year later when someone reads the terms. The technical work was done; the compliance work was skipped, and it is much harder to unwind an integration than it is to choose a differently licensed dependency.
Where this is going
Three structural pressures were visible by early 2024 and have only sharpened since.
Licensed data as a product. If unlicensed scraped data carries unresolved risk, then cleanly licensed corpora become valuable, and a market for them was forming — publisher agreements, licensed archives, synthetic generation with clear provenance. The cost of a model rises and the legal risk falls, and where the balance settles is a commercial question.
Regulation arriving to fill the gap. The EU's AI Act reached a provisional political agreement in December 2023 and includes documentation and transparency obligations for general-purpose models, including a requirement to publish summaries of training content. That is a disclosure regime arriving before the courts resolved the underlying rights questions, and its practical effect is to make the provenance problem something a developer must document rather than something they can ignore.
Standardization pressure from buyers. Enterprises do not want to read a new custom license for every model. The predictable outcome is that licenses converge toward a small number of recognizable patterns, and that documentation converges toward a standard structure because procurement demands it. The OSI's effort to define what "open source AI" should mean was an early sign of that pressure, and the argument about whether open weights without open data qualify is the argument the field will have.
The thing to hold onto is that none of this is a technical limitation being worked around. Model licensing is a set of commercial and legal choices about rights in a new kind of artifact, and they are being decided by practice — how publishers write terms, how buyers negotiate them, how courts rule — rather than by principle. Engineers who treat the license as a checkbox will eventually discover it is not one. Engineers who read it, record what they concluded, and keep a list of what would change their mind will be fine, which is the same advice that applies to every dependency decision they were already making.
Works Cited
BigScience. "The BigScience Responsible AI License (RAIL)." Hugging Face, 2022, huggingface.co/spaces/bigscience/license. Accessed 7 Mar. 2024.
European Parliament and Council. "Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence." Official Journal of the European Union, 2024. Accessed 7 Mar. 2024.
Gebru, Timnit, et al. "Datasheets for Datasets." arXiv, 2018, arxiv.org/abs/1803.09010. Accessed 7 Mar. 2024.
Henderson, Peter, et al. "Foundation Models and Fair Use." arXiv, 2023, arxiv.org/abs/2303.15715. Accessed 7 Mar. 2024.
Kocetkov, Denis, et al. "The Stack: 3 TB of Permissively Licensed Source Code." arXiv, 2022, arxiv.org/abs/2211.15533. Accessed 7 Mar. 2024.
Longpre, Shayne, et al. "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing and Composition." arXiv, 2023, arxiv.org/abs/2310.16787. Accessed 7 Mar. 2024.
Meta. "Llama 2 Community License Agreement." Meta, 2023, ai.meta.com/llama/license/. Accessed 7 Mar. 2024.
Mitchell, Margaret, et al. "Model Cards for Model Reporting." arXiv, 2018, arxiv.org/abs/1810.03993. Accessed 7 Mar. 2024.
Open Source Initiative. "The Open Source AI Definition." Open Source Initiative, 2024, opensource.org/deepdive. Accessed 7 Mar. 2024.
Scao, Teven Le, et al. "BLOOM: A 176B-Parameter Open-Access Multilingual Language Model." arXiv, 2022, arxiv.org/abs/2211.05100. Accessed 7 Mar. 2024.