When a customer asks what your model learned from, the base model is the easy half: point to your vendor's published training data summary and its indemnity. The half that fails the questionnaire is your own fine-tuning, retrieval and synthetic data, which no vendor indemnity covers and which most teams cannot list. Write one machine-readable manifest per dataset in Croissant, use the Data & Trust Alliance's 22 provenance fields as your checklist of what goes in it, and get lineage from the pipeline tool you already run. Do not buy a dedicated provenance platform for this, and do not hand a customer a dataset card as your evidence.
The workload this guide assumes: one product team fine-tuning a vendor base model on 30 to 50 datasets (licensed corpora, scraped pages, customer records, synthetic examples) and answering a provenance questionnaire every few weeks. Prices were checked on each vendor's live page on October 5, 2026.
| Option | What it records | Answers "what did it learn from?" | Cost (checked Oct 5, 2026) | Verdict |
|---|---|---|---|---|
| Croissant 1.1 + RAI 1.0 (MLCommons) | Per-dataset manifest: source, licence, collection method, PROV-O lineage, ODRL usage terms | Yes, if you fill it in | Free, open standard | Pick it as the manifest format |
| Data & Trust Alliance Provenance Standards 1.0.0 | 22 fields across Source, Provenance and Use | It tells you what to answer; it stores nothing | Free | Use it as the field checklist |
| OpenLineage + Marquez | Automatic job and dataset lineage from Spark, Airflow, dbt, Flink | Shows transformations, says nothing about rights | Free, Apache-2.0 | Pick it if you are not on Databricks |
| Databricks Unity Catalog lineage | Table and column lineage, notebooks, jobs, model serving | Same as above, plus whatever tags you add | No separate lineage charge in the docs; you pay for Databricks | Pick it if training already runs there |
| Weights & Biases Artifacts | Which dataset version fed which training run | Links data to a model version; no rights record | Free (5 GB); Pro from $60/month; Enterprise via sales | Fine as the run link, weak as the record |
| Hugging Face dataset cards | YAML license and source_datasets plus free text |
Partly, and unreliably | Free; Team $20/user/month; Enterprise $50/user/month | The loser as evidence of record |
What Customers Can Make You Show, Framework by Framework
The questionnaire you receive is usually stitched together from four sources, and each asks for something different. Knowing which one a question came from tells you how much detail you owe.
| Framework | Who it binds | What you must be able to show |
|---|---|---|
| EU AI Act Article 53(1)(d) | Providers of general-purpose AI models, including anyone who modifies one | A public summary on the Commission's template; a modifier covers only the modification data |
| California AB 2013 | Anyone who designs or substantially modifies a generative AI system offered to Californians | A high-level summary of training datasets covering 12 items, from sources and owners to synthetic data use |
| ISO/IEC 42001 Annex A.7.3 and A.7.5 | Organisations certified, or seeking certification | Documented acquisition and a recoverable lineage for each dataset |
| NIST AI RMF MAP 4.1 | Voluntary; common in US enterprise questionnaires | Documented legal risk of third-party data, including IP infringement |
The EU template, published July 24, 2025, asks for data types and sizes, a list of sources (public datasets, private datasets, scraped data, user data, synthetic data) and how copyright opt-outs were handled. For scraped data it wants the top 10% of domains by volume (top 5% or 1,000 for SMEs). The Commission's FAQ says fines of up to 3% of global turnover or €15 million became enforceable on August 2, 2026, and that a downstream modifier discloses only the data used for the modification, cross-referencing the original model's summary. Most enterprise fine-tunes will not be GPAI models in their own right, but your customers' lawyers will borrow the template's structure anyway because it is the only official one.
AB 2013 took effect January 1, 2026 and covers systems released since January 1, 2022. Its 12 items are the most useful checklist in circulation: sources or owners, purpose, number of data points, data types, whether data is copyrighted or public domain, whether it was purchased or licensed, personal information, aggregate consumer information, cleaning and processing, collection period, first use date, and synthetic data. xAI sued to block it; Judge Jesus Bernal denied the preliminary injunction on March 4, 2026, finding xAI's trade secret argument relied on "abstraction and hypotheticals rather than specific facts." Plan as if the law stands.
ISO/IEC 42001 control A.7.5 requires that you record where each dataset came from and what happened to it (creation, updates, transformations, transfers) across the data's life cycle and the AI system's. And MAP 4.1 of the NIST AI RMF asks for documented approaches to the legal risks of third-party data, including infringement of third-party IP. None of the four makes you publish dataset names, but all four assume you can list them internally.
Why Acquisition Method Belongs in Every Manifest
The most expensive provenance fact in AI so far was how the data was obtained, and most manifests do not record it. In Bartz v. Anthropic, Judge William Alsup held that training on lawfully purchased books was fair use but that downloading and keeping pirated copies was not; the resulting class settlement of $1.5 billion over about 482,000 works, roughly $3,000 per work, received final approval on July 20, 2026. Courts are still splitting on the training question itself. On September 29, 2026 the Third Circuit affirmed that ROSS Intelligence's use of Westlaw headnotes to train a legal research tool was not fair use, citing harm to an emerging market for licensing training data, while leaving generative AI questions open. In the UK, Getty's secondary infringement claim against Stability AI failed in November 2025 because the model weights were not held to be infringing copies. A record of acquisition method, licence and date per dataset is what lets your counsel apply whichever of these ends up governing your case.
Croissant: Pick It as the Manifest Format
Croissant is MLCommons' open metadata vocabulary for ML datasets, and since version 1.1 it can carry machine-readable provenance and usage terms. Croissant 1.1, released February 12, 2026, recommends W3C PROV-O properties such as prov:wasDerivedFrom and prov:wasGeneratedBy, down to the field level, and supports ODRL permissions and constraints for usage restrictions. The RAI 1.0 extension adds rai:dataCollection and rai:dataCollectionRawData for the collection process and raw source. MLCommons says Croissant metadata makes more than 700,000 datasets across Hugging Face, Kaggle, OpenML and the rest of the web findable, so many public datasets you pull in already arrive with a starting manifest.
It wins because it is a file you own, versioned next to the data, readable by a script that generates the questionnaire answer. Map the AB 2013 items onto it and you can answer the next questionnaire from a query instead of a meeting.
Who should not pick it: a team that only does retrieval over a vendor model and trains nothing. You have no training data to describe; answer with the vendor's documents and a data-flow diagram for retrieval.
The Data & Trust Alliance Fields: Use Them as the Checklist
The Data Provenance Standards v1.0.0, announced July 9, 2024, define 22 metadata fields in three groups: Source, Provenance and Use. The Use group covers confidentiality, consent and licence to use, which is the part Croissant leaves to you. IBM tested them in its dataset clearance process for foundation model training, and the 19 co-developing members include Mastercard, Pfizer, Walmart and UPS.
It is a list of fields, not a system, so nobody should adopt it alone. Its value is that a manifest that fills all 22 fields plus the AB 2013 items will cover nearly any questionnaire line you will see.
OpenLineage or Unity Catalog: Use the Lineage You Already Have
Lineage tools record what happened to data automatically, and nothing about whether you had the right to use it. That makes them necessary for ISO 42001 A.7.5 and useless on their own for a copyright question.
OpenLineage is an Apache-2.0 standard, a graduate project of the LF AI & Data Foundation, that emits run, job and dataset events from Airflow, Spark, Flink, dbt and others, with Marquez as the reference server. Pick it if your pipelines run on those tools outside Databricks. Skip it if your training data is assembled in notebooks and ad hoc scripts: nothing emits events, so the graph will be empty where the questionnaire looks.
On Databricks, Unity Catalog lineage captures table and column lineage, notebooks, jobs and model serving, with no separate lineage charge in its documentation. Two limits matter for provenance. Column lineage fails when a job reads a path such as s3://bucket/path instead of a table name, which is how a lot of training data gets loaded. And the lineage system tables keep only a rolling one-year window, while Catalog Explorer retains lineage from September 1, 2024 onward. A model you shipped 14 months ago may already have lost its queryable lineage, so export it at each release. Do not pick it if your training data lives outside Databricks.
Weights & Biases Artifacts: The Run Link, Not the Record
Weights & Biases Artifacts answers one question well: which version of which dataset went into the model you shipped. On the current pricing page (W&B now sits under CoreWeave), registry and lineage tracking are in every tier: Free includes 5 GB a month, Pro starts at $60 a month with 100 GB and $0.03 per extra GB, and Enterprise is contact sales. Audit logs are Enterprise only, so a team that needs to prove who changed a dataset record should not stop at Pro. Use it to hold the pointer from model version to manifest version, and keep the rights information in the manifest.
Hugging Face Dataset Cards: Why They Lose
A dataset card is a README with YAML metadata, written by whoever uploaded the data. The card format has a license field, a source_datasets field and free text, all self-reported. When the Data Provenance Initiative audited more than 1,800 text datasets, it found licence omission rates above 70% and licence categorisation error rates above 50% on popular hosting sites. If you copy the licence from a card into your answer, you are repeating a field of the kind that audit found miscategorised more than half the time.
Hugging Face itself is fine to host data on. Team at $20 per user per month and Enterprise at $50 add audit logs, resource groups and storage regions, which help with access control. Keep using cards for discovery, and verify the licence at the original source before it goes in your manifest.
Third-Party and Synthetic Data: What to Demand in Writing
For licensed data, the attestation you need from the seller covers four things: the source list, the collection method, the licence chain with the right to use it for model training, and a warranty backed by indemnity. A seller who will not give you the collection method is asking you to carry the Bartz risk on its behalf. Also record the licence as it stood on the day you acquired the data; terms move, as Reddit's API shutdown shows.
Synthetic data inherits the terms of the model that generated it. The Llama 3.1 licence requires that any distributed model trained on Llama outputs carry "Llama" at the beginning of its name. For each synthetic dataset, record the generator model and version, the licence it was used under, the prompt set, and which real data seeded it. AB 2013 asks directly whether synthetic data was used, so "it's synthetic" does not end the question.
Indemnification: What It Covers and Where It Stops
Every major vendor indemnity covers the vendor's model and outputs, and none covers the data you add. Anthropic's commercial terms defend claims that authorized use of the services, "which includes data Anthropic has used to train a model," infringes IP, and exclude claims arising from "Inputs or other data provided by Customer" and from modifications to outputs. Google's two indemnities cover Google's own training data and generated output on named services. Microsoft's Customer Copyright Commitment covers Output Content, and for Azure OpenAI it requires a copyright metaprompt, protected material filters, and a testing and evaluation report you retain and hand over if you claim. AWS's uncapped indemnity under Service Terms section 50.10 covers output of listed Amazon Nova models, which leaves third-party models on Bedrock outside it.
The indemnity you can rely on is the one whose conditions you can prove you met on the day of the claim. Store the Microsoft evaluation report, your filter settings and your prompt templates with the model version they apply to. For more on how these carve-outs get voided in practice, see our analysis of the guardrail circumvention exclusion.
Answering Honestly When the Answer Is "We Do Not Know"
Say what you do not know, why, and what you did instead; the EU uses the same format. The Commission's template lets providers of models released before August 2, 2025 state and justify information gaps where the data is unavailable. Borrow that structure.
For the base model, quote the vendor's disclosure and stop there. OpenAI and Anthropic both published AB 2013 documents by January 1, 2026, and neither names specific datasets; they describe "publicly available information" and data from third-party partners. A January 2026 paper by Judy Hanwen Shen and colleagues calls this a specification gap: a mismatch between the stated goals of data transparency and the disclosures actually needed to achieve them. Do not paper over it with your own guesses. A sample line: "The base model is [vendor model, version]. Its training data is described only at the category level in [link]. We have not been given, and cannot verify, the specific datasets. Our fine-tuning data is listed below."
For your own legacy data with lost lineage, name the dataset, state that the acquisition record is missing, give the date range you can establish, and say whether you have retired it or are retraining without it. A buyer can price a documented gap into the contract.
The Decision: Criteria That Predict Regret
Four questions predict whether this program holds up when a real claim arrives.
- Can you rebuild the manifest for a model version you shipped 18 months ago? If your lineage lives only in Unity Catalog system tables, it cannot, because they keep one year.
- Does every dataset record its acquisition method and licence as of the acquisition date? This is the field Bartz turned on.
- Can you name every model version a given dataset fed, so you can respond if a rights holder objects? W&B or MLflow run links answer this; a spreadsheet does not.
- Does each indemnity you rely on cover the model you actually run, with its conditions evidenced? A Claude model on Bedrock is outside AWS's Nova indemnity.
What changes the answer: if you train a model that you place on the EU market as general-purpose, you need the full Commission template and probably counsel, not just a manifest. If you train nothing, this whole guide reduces to collecting vendor documents. For the vendor review itself, our six questions for AI vendor security reviews covers indemnity and subprocessors.
What to Do Next
This Week:
- List every dataset in your current production model, with owner, acquisition method and licence. Mark each one where any of the three is unknown.
- Download your base model vendor's AB 2013 or EU summary and the indemnity terms, and file them with the model version.
This Month:
- Write a Croissant 1.1 manifest for each dataset, filling the 12 AB 2013 items and the DTA Use fields. Store it next to the data in version control.
- Turn on OpenLineage or confirm Unity Catalog lineage covers the training jobs, and add an export of lineage to every model release.
Before Your Next Model Release:
- Generate the questionnaire answer from the manifests with a script, and have counsel review the "we do not know" lines once.
- Get a written attestation of collection method and training rights from every data seller, and record generator model and licence for each synthetic set.
Start with the dataset list you can produce this week.
Continue Reading
- Sony Sued Anthropic. Re-Prompting Voids Your Indemnity.
- AI Vendor Security Review: 6 Questions That Change the Answer
- CoCounsel's New Model Runs on Qwen. Go Read the Card.
- EU AI Act Governance Tools: Buy Inventory, Not Policy Packs
- Best MLOps for Regulated AI: Domino, Then a Sign-Off Layer
- Reddit's API Shutdown Spares Sprinklr, Not Your Python Scripts
