What is a foundation model, and why its data requirements invert the usual ones

A narrow model needs more examples of one task. A foundation model needs coverage of everything you have not thought of yet, plus an evaluation slice you cannot buy.

A narrow model is built for one job: transcribe this audio, classify this message, detect this object. Its data requirement is a matter of quantity within a task that is already known. A foundation model is built to be a starting point for many jobs, some of which are not known when the corpus is assembled.

That inversion is the whole story from a purchasing perspective. The requirement stops being "more examples of the task" and becomes "coverage of everything downstream work might touch", which is harder to buy and much harder to verify after delivery.

Coverage beats volume, and coverage is a list

When a narrow model fails, it usually fails on the case missing from its training set. When a foundation model fails, it usually fails on a domain, register, or language missing from the corpus — a much larger unit of absence.

So the useful procurement question is not "how much text" but "which slices, in what shares". A slice inventory lists the axes that matter for your downstream work: language and script, register, format, time period, source type, and whether the material is human-written or machine-generated. Each axis gets a share, and those shares are decisions you make rather than facts the supplier discovers.

Insisting on the inventory instead of a total matters because one aggregate number hides everything that decides whether the corpus is usable. A very large corpus in a single language is not a foundation-model corpus for a product that has to work in five, no matter how impressive the total looks on a slide.

What the scaling literature says to a buyer

Published scaling work describes model loss as a joint function of parameters, data, and compute. The commercial reading is narrower and more useful: the three have to move together, and past a point, adding compute without adding data spends the compute for nothing.

Two implications follow. Data volume alone has diminishing returns, so the marginal value of the last slice of a corpus is lower than the first. And filtering plus deduplication shifts the curve itself rather than moving along it: a smaller corpus with the duplicates removed trains better than the larger raw version. That is a purchase decision, and it is usually the cheapest improvement available at any fixed compute budget.

Deduplication deserves its own question. Was it run across the whole corpus, or within each shard before the shards were combined? A corpus assembled from sources that were each deduplicated internally can still be full of near-duplicates across sources, and repeated text is the most common reason a model recites rather than reasons.

The slice you cannot buy

The corpus is purchasable. The thing that tells you whether the resulting model is any good is not. An evaluation set that reflects your deployment — your languages, your formats, your users' phrasing — has to be built from your own material, and it has to exist before training starts.

The ordering is not discipline for its own sake. Without a fixed evaluation set built in advance, every later comparison is against a moving reference, and it becomes impossible to attribute a change to the data, the recipe, or the measurement itself.

Budget for it as its own line. It takes longer than expected, because it needs the same care as training data and none of the volume, and teams that skip it spend the savings later on arguments they cannot settle.

Adapting the model needs a different dataset

Once the foundation model exists, the work moves to adaptation and the data changes shape completely: instructions or preference rows in the thousands, not a corpus. Buying more pretraining-scale text will not fix an adaptation problem, and buying adaptation data will not repair a model that lacks the underlying coverage.

The two purchases fail in different ways. Coverage failures look like a model that is competent in general and lost on your domain. Adaptation failures look like a model that knows the material and answers in the wrong form. Working out which one you have, before buying either, is worth more than any discount on the wrong one.

What to ask before buying a pretraining corpus

The last item on this list is what separates suppliers. Anyone can hand over files. A supplier who can also say what is not in the corpus, and why, has done the work that makes it usable in a project that will run for months.

  • The slice inventory with counts per slice, rather than a single total.
  • How deduplication was run — across the corpus, or within shards before combining.
  • Provenance and license per source, not per delivery, since one unusable source affects the whole set.
  • Whether the corpus contains material from the public benchmarks your evaluation uses, because that would compromise the evaluation before training starts.
  • Whether a composition document ships with the data, describing what is in it, how it was filtered, and what is known to be missing.

More insights

Submit a sourcing request

Tell us the language, the hours, and what the data needs to look like. You will get a real number and a real timeline — not a range. If we cannot source it well, we will tell you that instead.

  • Pilot batch before the full run, so problems surface early.
  • Consent documentation delivered with the data.
  • No medical or clinical data. No recorded telephone calls.

We reply within two business days. Your details are used only to answer this request. See our privacy policy.

Contact

Talk to a human

Send a specification and we will come back with a real number and timeline.

Submit a sourcing request

Or email hello@linguacorpus.com