AI Vendor Terms, Training Rights, and What the Contract Has to Say

A business that builds a product on a hosted language model usually accepts several layers of terms. The provider’s commercial agreement, data processing addendum, product terms, and account settings can govern different parts of the relationship.

Together, they determine whether the provider may train on what your customers type, how long it keeps prompts and outputs, who owns the output, and whether it will defend you against an infringement claim. Those terms also affect the promises you can make in your privacy policy and customer contracts.

Where the No Training Commitment Lives

A data processing addendum governs the provider’s handling of personal data, but an express commitment against model training may appear elsewhere. You should review the addendum together with the commercial terms and any terms specific to the product you use.

Anthropic’s data processing addendum limits processing to providing the services and following the customer’s documented instructions. Section B of its commercial terms states the training restriction directly. Anthropic may not train models on customer content from the covered services. The addendum’s purpose restrictions and the commercial agreement’s training restriction should be read together.

OpenAI also makes a contractual commitment. Under § 4.2 of its Services Agreement, OpenAI may use customer content for specified purposes, including providing the services, complying with law, and preventing abuse. It may not use that content to develop or improve the services unless the customer agrees to that use.

Microsoft’s Azure data privacy statement states that customer prompts and completions aren’t available to OpenAI and aren’t used to train generative foundation models without the customer’s permission or instruction. Amazon Bedrock uses a different hosting arrangement, isolating model providers from customer prompts and completions. In each case, you should identify both the model developer and the company hosting the service.

Google’s Gemini API terms distinguish paid and unpaid services. The unpaid provisions permit improvement of products and machine learning technologies using submitted content and responses, with potential human review. The paid provisions prohibit using prompts or responses to improve Google’s products.

Billing configuration and geography affect that distinction. Some free AI Studio access qualifies as paid through an active billing relationship or a Workspace enterprise account. Paid data treatment also applies in the European Economic Area, Switzerland, and the United Kingdom even when access is free. You should confirm the provisions governing your account before submitting customer information.

Consumer accounts require a separate review. Anthropic’s consumer privacy policy permits training on inputs and outputs unless the user opts out, with exceptions for specified safety review and feedback uses. That policy excludes content processed for business customers. An employee who pastes a customer contract into a personal account may place it under different terms from those negotiated for the company’s commercial account.

Retention Is a Separate Question From Training

A provider that never trains on your data may retain it for abuse monitoring or to support a product feature. OpenAI’s data controls documentation describes default abuse monitoring retention of up to 30 days, subject to exceptions. Azure also describes abuse monitoring that can involve storage and human review, while Google’s paid Gemini terms permit logging for a limited period to detect prohibited uses and meet specified legal requirements.

Providers offer different alternatives to their defaults. OpenAI’s zero data retention eligibility depends on the endpoint and feature, while approved Modified Abuse Monitoring excludes customer content from abuse monitoring logs. That distinction requires attention because a feature may separately store application data even when abuse monitoring logs exclude customer content.

Microsoft accepts applications for modified abuse monitoring, which changes storage and human review while automated review continues. Anthropic’s zero data retention arrangements include exceptions for legal compliance and misuse, and retention settings for models accessed through Bedrock depend on the service and model. A business should confirm the arrangement covering each feature it uses.

A privacy policy promising that prompts are never retained can mislead customers if a provider stores them. The policy should accurately describe relevant retention practices and any exceptions. A commitment against training supports a statement about training; retention requires its separate answer.

Output Ownership and the Limits of an Assignment

Provider terms commonly address ownership of inputs and outputs. Anthropic assigns whatever rights it has in outputs to the customer under its commercial agreement, and OpenAI makes a similar assignment. Google’s Gemini terms state that Google won’t claim ownership over generated content and warn that other users may receive similar content.

An assignment transfers the rights the provider holds, which may be limited or nonexistent. U.S. copyright protection requires human authorship, so an assignment alone doesn’t make a model’s unaided output copyrightable. A business reviewing ownership should distinguish the provider’s contractual claims from copyright protection and from the possibility that an output infringes someone else’s rights.

Open weight models, whose model parameters are available for download, have their own license conditions. Meta’s Llama 4 Community License requires specified attribution and acceptable use compliance. When the license’s distribution conditions apply, a product must display “Built with Llama,” and certain models developed using Llama materials or outputs must have “Llama” at the beginning of their names.

The license also requires separate permission for licensees exceeding its 700 million monthly active user threshold, measured at the Llama 4 release date using the preceding calendar month. Meta disclaims warranties, limits liability, and requires the licensee to indemnify Meta for specified claims. A business that fine tunes and distributes a model should review these conditions before promising customers unrestricted ownership or use.

Indemnity for Infringing Output, and the Conditions Attached

Some providers offer to defend customers against specified intellectual property claims, subject to exclusions and conditions. The agreement governing the particular service determines which claims qualify and what the customer must do to preserve coverage.

Section K of Anthropic’s commercial terms covers specified claims arising from paid use of its services or outputs. Its exclusions include practicing a patented invention contained in an output and trademark claims based on using an output in trade or commerce. A general statement that the provider offers intellectual property indemnity doesn’t describe all those limitations.

Google announced training data and output indemnities for specified Cloud services. Its list of covered services identifies the eligible offerings, and the applicable agreement contains the conditions and exclusions. Paying for Gemini Developer API access doesn’t, by itself, establish entitlement to Google Cloud indemnity. A business should identify the service and agreement before relying on that protection.

Microsoft’s Customer Copyright Commitment requires specified safeguards for Azure OpenAI offerings. These include a system instruction against infringement and evaluations designed to detect reproduction of third party content. Customers must address significant ongoing reproduction found through those evaluations and retain the results and mitigation report for a claim.

Additional requirements depend on the use case. Text generation requires the protected material text filter in filter mode. Code generation requires the code filter in annotate or filter mode, with compliance with cited licenses when using annotate mode. Both require the jailbreak protection in filter mode.

Disabling a required safeguard can defeat coverage for the affected output. Your technical team should document compliance with the applicable requirements at launch and after configuration changes so the business can substantiate a request for defense.

What Your Contracts Must Contain

Privacy statutes can require particular terms when a covered business engages a processor or service provider. Applicability comes first. A business’s location, customers, activities, size, and the information involved can determine which requirements apply.

Texas generally exempts small businesses as defined by the U.S. Small Business Administration, although an exempt small business must obtain consent before selling sensitive personal data. For covered controllers and processors, Business and Commerce Code § 541.104 requires a contract addressing processing instructions, purposes, data types, duration, and each party’s rights and obligations. It also specifies confidentiality, deletion or return, assessments, information access, and subcontractor requirements.

Effective January 1, 2026, House Bill 149 amended the processor provisions to address assistance with security requirements for personal data handled by AI systems and breach notification involving the processor’s system. Businesses subject to these provisions should account for those duties in the vendor relationship.

California’s requirements apply to businesses within the CCPA’s statutory coverage, subject to its exemptions. Under Civil Code § 1798.100(d), covered contracts must specify limited purposes, require the applicable level of privacy protection, and permit reasonable steps to confirm compliant use. They must also require notice when the recipient can no longer meet its obligations and permit steps to stop and remedy unauthorized use.

Regulation § 7051 supplies additional service provider and contractor requirements. A standard vendor addendum may address these requirements, but the business should compare the relevant provisions before assuming that it does.

Children’s Data and Provider Restrictions

COPPA applies to covered collection of personal information online from children under 13. Information supplied by adults about children doesn’t, by itself, trigger the Rule, although other privacy laws may apply. The audience, source of the information, and operator’s knowledge determine the COPPA analysis. FTC guidance.

Under the amended Rule, covered operators must obtain separate verifiable parental consent before disclosing children’s personal information to a third party for AI training or development. Parents must have the option to consent to collection and use without consenting to that nonintegral disclosure. The FTC addressed AI training in its explanation of the amendments.

A binding restriction on training is one way to prevent that use. A business seeking to permit the disclosure must address the separate consent requirement and the Rule’s other obligations, including security and retention. If a provider requires training rights that the operator can’t reconcile with those obligations, the business may need a different provider or arrangement.

Provider restrictions require a separate check even when parental consent is available. For example, the Gemini Developer API terms prohibit use in applications directed toward, or likely to be accessed by, people under 18. Parental consent under COPPA doesn’t override that contractual restriction.

State Statutes That Address Training

Several states now address AI training directly. Their requirements differ, and a business should determine whether it acts as a controller, developer, deployer, or another regulated participant under each applicable statute.

Connecticut’s Public Act 25-113 amended the privacy notice requirement effective July 1, 2026. A covered controller must disclose whether it collects, uses, or sells personal data for training large language models. The amendments also expanded coverage, including circumstances in which a nonexempt business processes a single Connecticut consumer’s sensitive data.

The notice should reflect the business’s processing purposes and the provider arrangement. A vendor’s contractual permission to train, account settings, and the uses of data require review before the business can describe them accurately. A free subscription or consumer account label alone doesn’t answer that question for every provider.

California’s risk assessment regulations address specified training activities by covered businesses. Section 7150(b)(6) includes processing personal information to train automated decisionmaking technology for significant decisions about consumers, as well as specified technologies for physical or biological identification or profiling.

Section 7153 also requires a business making certain technology available to another business to provide information needed for the recipient’s risk assessment. The regulations took effect January 1, 2026, with transition provisions for existing processing, and the first assessment submissions are due April 1, 2028. A business should identify both the assessment obligation and the deadline applicable to its processing.

California separately requires training data disclosures under Civil Code § 3111, enacted through Assembly Bill 2013. Subject to statutory exceptions, developers making generative AI systems or services publicly available to Californians must publish documentation summarizing the training datasets. The requirements applied by January 1, 2026 and before subsequent releases or substantial modifications.

The required documentation includes sources or owners, types and volume of data, protected intellectual property, personal information, licensing or purchase, and synthetic data use. A business that fine tunes a model and publicly offers the resulting system may qualify as a developer, so this review extends beyond companies building foundation models.

On March 4, 2026, the district court denied xAI’s request for a preliminary injunction against the statute in X.AI LLC v. Bonta. That ruling addressed preliminary relief and did not finally resolve the constitutional challenge.

Colorado replaced its 2024 AI act before it took effect. Senate Bill 26-189, signed May 14, 2026, governs covered automated decisionmaking technology used in consequential decisions beginning January 1, 2027.

The replacement law requires covered developers to provide deployers with documentation that includes categories of training data, intended and known inappropriate uses, limitations, and instructions for appropriate use and human review. The attorney general enforces the law, which includes a 60 day cure period subject to exceptions for knowing or repeated violations. Businesses buying or selling covered scoring and screening tools should address the required documentation in their agreements.

Texas’s House Bill 149 also amended the biometric identifier statute, Business and Commerce Code § 503.001. It added an exception for specified training, processing, or storage involved in developing AI models or systems, except when the system is used or deployed to identify a specific individual uniquely. Repurposing identifiers captured under that exception for a commercial use outside it can expose the possessor to statutory penalties.

The Texas amendments also address publicly available biometric identifiers. An individual’s appearance in a public source doesn’t establish consent to capture merely because the source is public, unless the individual made the identifier public.

Arkansas addressed ownership through Act 927 of 2025, which added Arkansas Code § 18-4-101. Under its terms, the person providing the input or directive owns the generated content, subject to existing rights. The person providing lawfully acquired training data owns the resulting trained model unless ownership has been transferred by contract. The statute also provides for employer ownership when an employee acts within the scope of employment and under the employer’s direction and control.

Those state ownership provisions should be considered alongside the provider agreement and federal intellectual property law. They don’t eliminate the separate question of whether particular output qualifies for federal copyright protection.

Subprocessors and Changes to Terms

A subprocessor is another company a provider uses to process customer personal data. Your agreement should address how the provider may engage these companies, what protections apply, and how you receive notice of changes.

Anthropic’s commercial terms generally allow updates to take effect 30 days after posting or notice, with an exception for changes required by law. Its data processing addendum requires notice before a new subprocessor processes customer personal data and provides a 15 day objection period. A business should monitor the applicable notices and assess whether a change requires updates to its contracts or disclosures.

Changing your privacy policy also requires care. In a February 2024 staff post, the FTC warned that quietly changing terms to permit new uses of previously collected data, including AI training, may be unfair or deceptive. Publishing revised language doesn’t necessarily authorize a use that conflicts with earlier promises.

The remedy can extend to models developed from improperly collected data. In In re Everalbum, Inc., the FTC’s 2021 order required deletion of specified data and models or algorithms developed using the covered biometric information. A business should address lawful collection and use before treating customer information as training material.

Trade Secrets and the Prompt Window

Trade secret protection depends in part on reasonable measures to preserve secrecy. Under 18 U.S.C. § 1839(3), the information must also derive independent economic value from its secrecy.

Employees who submit source code, pricing models, or customer lists to a consumer chatbot under terms permitting training or reuse can create an argument that the business failed to take reasonable protective measures. The consequences depend on the terms, disclosures, safeguards, and surrounding facts.

Commercial restrictions on training and reuse, confidentiality obligations, and an internal policy directing employees to approved accounts can support the business’s protection of confidential information. Their adequacy depends on how the business implements and enforces them.

Reviewing the Provider Relationship Before Launch

You should identify each provider handling customer information, including the company hosting the model, and review the terms governing training, retention, ownership, indemnity, and changes to the service. Product restrictions, account settings, and any separately negotiated terms belong in the same review.

You should then determine which privacy and AI statutes apply to your business and processing. The relevant vendor contracts and notices should address those requirements, including any parental consent, training disclosures, assessments, or developer documentation.

Your privacy policy and customer contracts should describe the practices and protections the business can support. When a provider changes its terms, adds a subprocessor, or changes the service or account arrangement you use, you should assess whether those changes require different disclosures, agreements, or consent.

This article is general information about the law, not legal advice, and reading it does not create an attorney-client relationship. Laws change and how they apply depends on your specific facts. For advice on your situation, consult a qualified attorney.

Need advice tied to your business issue?

Share the issue. Get direct attorney review. Receive a concrete recommendation.

Submit an Inquiry