Skip to main content

Envisioning is an emerging technology research institute and advisory.

LinkedInInstagramGitHub

2011 — 2026

research
  • Observatory
  • Newsletter
  • Methodology
  • Origins
  • Vocab
  • RSS Feeds
services
  • Signals Session
  • Bespoke Projects
  • Build Sessions
  • Pricing
  • Use Cases
  • Signals
  • Signal Scan↗free
impact
  • ANBIMAFuture of Brazilian Capital Markets
  • IEEECharting the Energy Transition
  • Horizon 2045Future of Human and Planetary Security
  • WKOTechnology Scanning for Austria
solutions
  • Innovation
  • Strategy
  • Consultants
  • Foresight
  • Associations
  • Governments
  • L&D
resources
  • Partners
  • Coding for Non-Coders
  • How We Work
  • Data Visualization
  • Multi-Model Method
  • FAQ
  • Security & Privacy
  • Public Sector
about
  • Manifesto
  • Community
  • Events
  • Support
  • Contact
ResearchServicesSignalsAbout
ResearchServicesSignalsAbout
  1. Home
  2. Vocab
  3. Data Mixture

Data Mixture

The composition of training data — the proportions of different sources, domains, languages, and document types — used to train a machine learning model.

Year: 2020Generality: 720Added: Aug 11, 2026
Back to Vocab

A data mixture is the composition of a model's training corpus: the relative proportions of different data sources, domains (web, books, code, scientific papers, dialogue), languages, document types, and quality strata that are mixed together before or during pretraining. Data mixture is one of the most consequential design choices in modern pretraining pipelines — small changes in mixture proportions can shift a model's downstream capabilities substantially, and frontier model labs treat the recipe as proprietary. Mixture design involves tradeoffs between coverage (diversity of topics and languages), quality (weighting cleaner sources like curated textbooks over noisy web scrapes), and capability transfer (some data sources disproportionately improve downstream reasoning). Data mixture inference is the related but distinct practice of reverse-engineering a closed model's mixture from its external behavior. Public mixture specifications (e.g., the GPT-3, LLaMA, and DeepSeek papers) have become a reference grammar for the field even when exact ratios are not reproduced.

Sources

  1. Exploring Claude/GPT Knowledge Cutoffs & Pre-training Timelines

    Shrivu's Substack · Aug 10, 2026

  2. Language Models are Few-Shot Learners

    arXiv (NeurIPS 2020) · May 28, 2020

  3. LLaMA: Open and Efficient Foundation Language Models

    arXiv · Feb 27, 2023

Research this in Signals

Scan Data Mixture for yourself.

Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Data Mixture already loaded, so edit it or scan as is.