Training data

Intermediate

Where training examples come from, how they are processed, and what a useful dataset description must disclose.

Last updated: Sep 13, 2026

Sources with inspectable provenance

Selected examples, checked on 2026-09-13. Sizes belong to the named release. Bytes, documents and tokens measure different things; token counts depend on the encoding.

Common Crawl

Common Crawl Foundation · Web

Crawl-ID / snapshot

An archive of web crawls. The crawl ID and filtering choices define the actual training input. Access to the archive does not establish one licence for all underlying pages.

Original source / dataset card

The Pile

EleutherAI · Mixed

2020 paper · 825 GiB · 22 subsets

A historical English mixture of 22 components. Inspect provenance and rights per component; the original mixture includes Books3. A whole-mixture label such as legitimate obscures these differences.

Original source / dataset card

FineWeb

Hugging Face · Web

2024 release · 15T tokens

Filtered English Common Crawl text. The release cited here is a historical snapshot, not a claim about the current total. The dataset card documents processing and ODC-By terms; source content rights remain a separate consideration.

Original source / dataset card

The Stack v2

BigCode · Code

v2 · 67.5 TB (full corpus)

Code sourced through Software Heritage. The full corpus, deduplicated corpus and training subsets have different sizes. Follow the dataset terms and original repository licences; pin a revision because opt-outs change releases.

Original source / dataset card

Cosmopedia

Hugging Face · Synthetic

v0.1 · 2024

Synthetic educational text generated using Mixtral. Record generator, prompts, seed sources and validation. Synthetic describes how data was produced, not its quality or legal status.

Original source / dataset card

Quality is a pipeline decision

Deduplication, language coverage, filtering and sampling change what a model learns. Keep a held-out evaluation set, check overlap with training data and evaluate target tasks. A larger corpus is not automatically a better corpus. Aggressive filters can also remove useful dialects, languages or minority viewpoints.

Synthetic data needs its own provenance

Model-generated data can expand instruction, code and reasoning examples. Preserve the generator version, prompts, source material and acceptance criteria. Validate answers with task-specific checks where possible and evaluate diversity and contamination. Repeatedly training on unchecked model outputs can amplify errors or lose coverage; the outcome depends on selection, fresh data and the training setup.

Access, licences and personal data are separate fields

Do not treat public, open, licensed and lawful as interchangeable labels. A dataset can combine components with different conditions. Document acquisition, original licences, opt-outs and personal-data handling. A lawsuit or allegation needs a source and procedural date; it is not itself a final ruling on every training use.

EU general-purpose model documentation

The AI Act includes obligations for providers of general-purpose AI models, including technical documentation, a copyright-compliance policy and a sufficiently detailed public summary of training content. Scope, exceptions and transition rules depend on the provider and model. This is not limited to high-risk downstream systems. Use the Commission guidelines for the current requirements.

European Commission: guidelines for GPAI providers

What to record before comparing datasets

  • Publisher and exact release or revision
  • Source domains, acquisition and transformations
  • Size, unit, tokenizer and filtering stage
  • Licence conditions and known provenance gaps
  • Split construction, benchmark overlap and evaluation results