US-State

California AI Training Data Transparency Act

The short answer. AB 2013 (2024) added Title 15.2 to the California Civil Code. Section 3111 requires the developer of a generative AI system or service made available to Californians to post documentation about the data used to train it, on or before January 1, 2026 and before each subsequent release or substantial modification. It is a public summary of the dataset, not a filing of the dataset, and the duty falls on the developer rather than on the data vendor who supplied the recordings.

The law

Cal. Civ. Code §3111

On or before January 1, 2026, and before each time thereafter that a generative artificial intelligence system or service, or a substantial modification to a generative artificial intelligence system or service, released on or after January 1, 2022, is made publicly available to Californians for use, regardless of whether the terms of that use include compensation, the developer of the system or service shall post on the developer's internet website documentation regarding the data used by the developer to train the generative artificial intelligence system or service.

The trigger is making the system available to Californians, not having a California entity. The required contents are listed in §3111(a) and they include facts only the data supplier holds: the sources or owners of the datasets, whether the datasets were purchased or licensed, whether they contain personal information as defined in Section 1798.140(v), whether synthetic data generation was used, and the time period during which the data was collected. Each of those is a question a buyer should be able to answer from the purchase file without calling the vendor.

Cal. Civ. Code §3110(b)

'Developer' means a person, partnership, state or local government agency, or corporation that designs, codes, produces, or substantially modifies an artificial intelligence system or service for use by members of the public.

Fine-tuning is inside the definition, because §3110(d) treats a substantial modification — including the results of retraining or fine tuning — as a new release. If you buy a corpus to fine-tune a model you distribute, the disclosure obligation can be yours, and the dataset you bought becomes part of a public document on your own website.

EU AI Act Article 53(1)(d)

draw up and make publicly available a sufficiently detailed summary about the content used for training of the general-purpose AI model, according to a template provided by the AI Office

The EU counterpart, aimed at providers of general-purpose models. The two regimes overlap on the questions they ask: the source of the data, the collection period, and whether synthetic data was used. A buyer preparing one summary to satisfy §3111(a) is most of the way to the AI Office template, and vice versa.

Who it applies to

Section 3111 has no size exemption and no open-source exemption in the enacted text. The only carve-outs, in §3111(b), are purpose-based: a generative system whose sole purpose is to help ensure security and integrity, a system whose sole purpose is the operation of aircraft in the national airspace, and a system developed for national security, military or defense purposes that is made available only to a federal entity.

The obligation is a posting, not a filing. Nothing in the title requires the developer to publish the data, the individual records, or the training set itself, and nothing in it sets a quality standard for the data. What has to be public is a high-level summary, at the dataset level, in the categories §3111(a) enumerates.

The duty attaches to the developer, so a data vendor has no direct obligation under the title. That is worth being precise about, because it is where buyers misread the law. The vendor is not the regulated party, but the developer's summary has to state facts about the vendor's data — who owns it, whether it was licensed, whether it contains personal information. A vendor that will not answer those questions in writing makes the buyer's own posting impossible to complete accurately.

What it costs to get wrong

Section 3111 contains no damages provision and no private right of action. Nothing in the title sets a fine. A reader looking for a penalty figure in the statute will not find one, and any page that quotes one is quoting something else.

What the title creates is a dated, public document on the developer's own website listing the datasets behind the model, whether they contained personal information, whether they were purchased, and when the data was collected. That document can be compared against what the vendor said during the sale, by a plaintiff's lawyer, by a regulator, or by a journalist. The mismatch is the exposure, and it is documentary rather than financial.

The practical consequence for a buyer is that the purchase file becomes evidence. If the posted summary says a dataset contained personal information and the contract said it did not, the buyer has to explain the difference. The defense is a procurement record that shows what was asked, what was represented, and when — which is the same file that supports every other compliance question on this page.

How to comply when you are buying data

The work is to buy data in a form that can be described publicly. That means writing the disclosure fields into the purchase contract, because they are facts about the vendor's data that the buyer cannot verify independently.

  • Require the source or owner of each dataset in writing, at the granularity §3111(a)(1) needs. "Proprietary" is not a description that can be posted.
  • Get the purchase and licensing status of each dataset, and the collection time period, including a statement of whether collection is ongoing.
  • Get a written answer on personal information: whether the dataset contains personal information as defined in Section 1798.140(v), and whether it contains aggregate consumer information. If the answer is no, ask what basis the vendor has for saying so.
  • Ask whether any part of the dataset was synthetically generated, and keep the answer in the dataset record. §3111(a)(12) asks specifically about synthetic data generation, including continuous use.
  • Decide in advance who writes the summary if the model is fine-tuned and distributed. If the answer is you, the facts you need are the ones above, and they have to be collected at purchase rather than reconstructed later.

How we handle consent and licensing →

Related compliance topics

Not legal advice

We are a sourcing company, not a law firm. Nothing on this page is legal advice, and it does not create a lawyer–client relationship. Whether a particular dataset is permissible in your jurisdiction depends on your use case, where you operate, and where the people in the recordings are located. Our role is to document the chain of consent accurately so that your counsel can assess it.

Sourcing data under California AI Training Data Transparency Act?

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com