The post has been translated automatically. Original language: Russian
Let's introduce the team a week before the round. The product works: an open model, trained on its own support data, closes applications faster than people. There is a slide titled "Proprietary AI model" in the presentation. The investor reads it and asks one question: what exactly is yours here?
Then there is usually a pause. Not because the team is hiding something. Because the word "model" in this question means not one thing, but a stack of different ones, and the rights to each live by their own rules.
I analyzed how large AI companies issue model rights: a public description of the OpenAI and Microsoft agreement, the WIPO guidelines for the protection of AI systems, license texts from Apache 2.0 to Llama Community License, patent registries. The research question is simple: what exactly does a company own when it "owns a model"?
The smaller the company, the faster this question comes to it from the outside: from an investor, an enterprise customer, or the site it enters.
Even the owners of frontier models do not consider the model to be a single asset.
In the public description of the agreement between OpenAI and Microsoft, confidential research methods, model architecture, weights, inference code, and retraining code are named as separate items. Separate lines. Separate rights.
This is not legal pedantry. WIPO uses the same modes: copyright in ML systems protects primarily the code, the model is broadly based on a combination of copyright, patents and trade secrets, and weights and parameters are given as an example of information that is protected by a confidentiality regime.
If those with the most models think so, the phrase "this is our model" should be taken apart before someone else's lawyer does.
Decompose the model into layers
| Layer | What is it | What is protected by |
| Architecture and methods | Network structure, learning methods, and information | Patent (where conditions are met), secret, informed publication |
| The source code | Training code, inference, SDK, safety-tools | Copyright, secret, licenses |
| Weights and checkpoints | Numerical parameters, model states, adapters | Secrecy, access control, contract; status in copyright is not defined |
| The recipe for learning | Data composition and proportions, hyperparameters, synthetics | It's almost entirely a mystery. |
| Data and evals | Training kits, markup, benchmarks, red-team kits | Contracts, licenses, secrecy |
| Alignment and promptness | Reward models, safety rules, system promptes, skills | Copyright for text and code, secret, licenses |
| Infrastructure and API | Serving, protocols, SDK | Patents, copyrights, open licenses |
| Brand | Model names, compatibility signs | Trademarks, ecosystem rules |
Copyright here protects a specific expression: written code, text, documentation. It does not protect the method: WIPO explicitly describes the possibility of legally reproducing the same functionality with other code. Therefore, one layer without the rest guarantees almost nothing.
An inconvenient rule follows from the table. You can't own a model more than you own every layer of it.
This line looks trivial right up to the first due diligence.
Consider weights a mystery, not a work of art.
Weights are the most valuable file in the company and the weakest point of the entire "our model" design.
Does libra have an author? Is there any originality? Is an independently trained duplicate that behaves almost the same way considered a copy? There is no single global answer to these questions.
Therefore, in practice, closed companies protect weights not as a work, but as a secret: access restrictions, contractual prohibitions, physical non-alienation of the file. The model is sold through an API, and the API here is not the delivery method, but the architecture of the fence: the user receives the result, but does not receive weights, code, recipe, and internal routing.
The most valuable file is the one with the most unclear legal status. This is not a paradox, this is the structure of the industry.
Read the license under your fine-tune
Now the main layer for practice. The team has completed the open model and calls the result its own. Legally, layered ownership appears there: the base weights live under the terms of the upstream license, new adapters and delta weights can belong to the developer, the data is used on its own grounds, the model name rests on trademarks.
And the licenses give very different answers.
Apache 2.0 — broad rights to use, modify, and commercially distribute, plus a separate patent clause. Llama 4 Community License — proprietary Meta terms: acceptable use policy included in the contract, audience threshold for large services, requirements for naming derivatives. The terms of Gemma go further: the concept of Model Derivatives also covers models created by transferring patterns through conclusions — by distillation or synthetic data — with the intention of repeating Gemma's behavior; however, the conclusions themselves are not considered derivatives. DeepSeek-R1 is distributed under MIT with direct distillation permission, but distilled variants inherit the licenses of the basic Qwen and Llama.
What does your license say about derivatives? About distillation? About commercial use? About the name? About using inferences to train another model?
The answer "we have completed our training, so it means ours" does not stand up to the first reading of the license.
Don't mistake open weights for open source
The market is constantly mixing several different modes. The proprietary model behind the API: access is available, there are no weights and no code. Open-weight: weights can be downloaded, but the data is not disclosed, the code may be missing, and the license may contain restrictions. Source-available: part of the code is shown, but without the freedoms of an open license. Open source AI in the strict sense: the definition of OSI requires modifiability not only parameters, but also code, and sufficiently detailed information about the data. And an open standard: the interface is open, and the implementations and models around it can be closed.
"The company is open" is almost an empty statement until three things are named.: what exactly is open, under what license, and with what restrictions.
Collect five responses to someone else's account
The phrase "our model" is divided into five questions:
Whose architecture? Whose learning code? Whose weights are and what is written in the license of the base model? Whose data is it, and on what basis did it get into the training? Whose adapters and advanced layers are yours or are they derived under the terms of the upstream license?
I have not yet found a model on the market where all five answers sound like "completely ours." Even frontier companies have licensed data and open components in their stack. For a fully trained model, the matrix of rights is a normal condition, not a diagnosis.
The same five questions are asked by the due diligence investor and the enterprise customer before the contract, almost verbatim: show what your model consists of and on what basis you use each layer.
Five questions — check for an hour with the team. The answer "I need to see the license" is the normal result of the first iteration. It's worse when the question is first asked in someone else's office.
Closing note
This is an operational diagnosis, not a legal opinion. The answers to the five questions depend on the jurisdiction — the status of libra, the patentability of methods, and the scope of data rights look different in different legal systems. For a specific transaction, the answer must be given by a specialist in applicable law.
Представим команду за неделю до раунда. Продукт работает: открытая модель, дообученная на собственных данных поддержки, закрывает заявки быстрее людей. В презентации есть слайд «Proprietary AI model». Инвестор читает его и задаёт один вопрос: что именно здесь ваше?
Дальше обычно наступает пауза. Не потому, что команда что-то скрывает. Потому, что слово «модель» в этом вопросе означает не одну вещь, а стопку разных, и права на каждую живут по своим правилам.
Я разбирала, как оформляют права на модели крупные AI-компании: публичное описание соглашения OpenAI и Microsoft, руководства WIPO по охране AI-систем, тексты лицензий от Apache 2.0 до Llama Community License, патентные реестры. Исследовательский вопрос простой: что именно принадлежит компании, когда ей «принадлежит модель»?
Чем меньше компания, тем быстрее этот вопрос приходит к ней снаружи: от инвестора, enterprise-заказчика или площадки, на которую она выходит.
Даже владельцы фронтир-моделей не считают модель одним активом
В публичном описании соглашения OpenAI и Microsoft отдельными позициями названы конфиденциальные методы исследований, архитектура модели, веса, inference-код и код дообучения. Отдельные строки. Отдельные права.
Это не юридическая педантичность. WIPO разводит те же режимы: авторское право в ML-системах защищает прежде всего код, модель в широком смысле держится комбинацией авторского права, патентов и коммерческой тайны, а веса и параметры приведены как пример информации, которую охраняют режимом конфиденциальности.
Если так считают те, у кого моделей больше всего, фразу «это наша модель» стоит разобрать на части раньше, чем это сделает чужой юрист.
Разложите модель на слои
| Слой | Что это | Чем охраняется |
| Архитектура и методы | Устройство сети, способы обучения и инференса | Патент (где выполнимы условия), секрет, осознанная публикация |
| Исходный код | Код обучения, inference, SDK, safety-инструменты | Авторское право, секрет, лицензии |
| Веса и чекпойнты | Численные параметры, состояния модели, адаптеры | Тайна, контроль доступа, договор; статус в авторском праве не определён |
| Рецепт обучения | Состав и пропорции данных, гиперпараметры, синтетика | Почти целиком тайна |
| Данные и evals | Обучающие наборы, разметка, бенчмарки, red-team наборы | Договоры, лицензии, тайна |
| Alignment и промпты | Reward-модели, правила безопасности, системные промпты, skills | Авторское право на текст и код, секрет, лицензии |
| Инфраструктура и API | Сервинг, протоколы, SDK | Патенты, авторское право, открытые лицензии |
| Бренд | Имена моделей, знаки совместимости | Товарные знаки, правила экосистемы |
Авторское право здесь охраняет конкретное выражение: написанный код, текст, документацию. Метод оно не охраняет: WIPO прямо описывает возможность законно воспроизвести ту же функциональность другим кодом. Поэтому один слой без остальных почти ничего не гарантирует.
Из таблицы следует неудобное правило. Нельзя владеть моделью сильнее, чем вы владеете каждым её слоем.
Эта строка выглядит банальной ровно до первого due diligence.
Считайте веса тайной, а не произведением
Веса — самый ценный файл в компании и самое слабое место всей конструкции «наша модель».
Есть ли у весов автор? Есть ли оригинальность? Считать ли копией независимо обученный дубль, который ведёт себя почти так же? Единого мирового ответа на эти вопросы нет.
Поэтому закрытые компании на практике охраняют веса не как произведение, а как тайну: ограничение доступа, договорные запреты, физическое неотчуждение файла. Модель продаётся через API, и API здесь — не способ доставки, а архитектура ограждения: пользователь получает результат, но не получает веса, код, рецепт и внутреннюю маршрутизацию.
Самый ценный файл — с самым неясным правовым статусом. Это не парадокс, это устройство отрасли.
Прочитайте лицензию под своим fine-tune
Теперь главный для практики слой. Команда дообучила открытую модель и называет результат своим. Юридически там появляется слоёное владение: базовые веса живут по условиям upstream-лицензии, новые адаптеры и дельта-веса могут принадлежать разработчику, данные используются на своих основаниях, имя модели упирается в товарные знаки.
И лицензии дают очень разные ответы.
Apache 2.0 — широкие права на использование, изменение и коммерческое распространение, плюс отдельная патентная оговорка. Llama 4 Community License — собственные условия Meta: политика допустимого использования включена в договор, порог по аудитории для крупных сервисов, требования к наименованию производных. Условия Gemma идут дальше: понятие Model Derivatives охватывает и модели, созданные переносом паттернов через выводы — дистилляцией или синтетическими данными — с намерением повторить поведение Gemma; при этом сами выводы производными не считаются. DeepSeek-R1 распространяется под MIT с прямым разрешением дистилляции, но дистиллированные варианты наследуют лицензии базовых Qwen и Llama.
Что ваша лицензия говорит про производные? Про дистилляцию? Про коммерческое использование? Про имя? Про использование выводов для обучения другой модели?
Ответ «мы дообучили, значит наше» не выдерживает первого чтения лицензии.
Не принимайте открытые веса за open source
На рынке постоянно смешивают несколько разных режимов. Проприетарная модель за API: доступ есть, весов и кода нет. Open-weight: веса можно скачать, но данные не раскрыты, код может отсутствовать, а лицензия — содержать ограничения. Source-available: часть кода показана, но без свобод открытой лицензии. Open source AI в строгом смысле: определение OSI требует для модифицируемости не только параметры, но и код, и достаточно подробную информацию о данных. И открытый стандарт: открыт интерфейс, а реализации и модели вокруг него могут быть закрыты.
«Компания открытая» — почти пустое утверждение, пока не названы три вещи: что именно открыто, под какой лицензией, с какими ограничениями.
Соберите пять ответов до чужого кабинета
Фраза «наша модель» раскладывается на пять вопросов:
Чья архитектура? Чей код обучения? Чьи веса — и что написано в лицензии базовой модели? Чьи данные — и на каком основании они попали в обучение? Чьи адаптеры и дообученные слои — ваши или производные по условиям upstream-лицензии?
Я пока не нашла на рынке модели, где все пять ответов звучат как «полностью наши». Даже у фронтир-компаний в стеке лежат лицензированные данные и открытые компоненты. Для дообученной модели матрица прав — нормальное состояние, а не диагноз.
Эти же пять вопросов задаёт инвестор на due diligence и enterprise-заказчик перед контрактом, почти дословно: покажите, из чего состоит ваша модель и на каком основании вы используете каждый слой.
Пять вопросов — проверка на час с командой. Ответ «надо посмотреть лицензию» — нормальный результат первой итерации. Хуже, когда вопрос впервые звучит в чужом кабинете.
Закрывающая заметка
Это операционная диагностика, а не правовое заключение. Ответы на пять вопросов зависят от юрисдикции — статус весов, патентоспособность методов и объём прав на данные в разных правопорядках выглядят по-разному. По конкретной сделке ответ должен давать специалист по применимому праву.