Pretraining data curation

We select, clean, and remove duplicates from text collections for pretraining. The mix of sources reflects your model, your permitted uses, and your compute budget.

ScopeA single domain corpus or ongoing pretraining data work
Delivery timingThe pilot’s processing speed sets the full corpus schedule
How we check itA corpus audit reports how much text was kept, duplicates found, the mix of sources, token counts, and known coverage gaps. It makes no claim of model gains before training.

A pretraining corpus you can account for

We agree with you on which sources belong in the corpus and how much each should contribute. Filtering reports show what we removed and why. A corpus for continued pretraining, which extends an existing model’s training on new text, differs from the instruction-and-response examples used for . We specify the collection for the training stage you plan to run.

Good fit

  • Teams preparing a foundation-model training run or domain adaptation
  • Research groups that need repeatable corpus selection and filtering

Not the right fit

  • Teams treating public web content as permission to train
  • Projects that estimate useful training volume from file size alone

What you get

Corpus and source register

Approved sources with collection versions, permitted-use notes, language and domain coverage, and traceable source identifiers.

Curation and mixture recipe

Repeatable extraction, filtering, deduplication, and sampling rules, with counts of rejected text and adjustable proportions for each domain.

Training-ready corpus export

Corpus files in shards, estimated token counts for your tokenizer, checks that supplied benchmark material has not leaked into the corpus, and a release manifest.

How it works

Set the corpus objective

We match domain coverage, language mix, tokenizer, training budget, and permitted sources to the model run you plan.

Tune filters on samples

We inspect samples of accepted and rejected text so the filters neither remove useful domain material nor keep duplicated, low-quality text.

Build and audit the corpus

We run the agreed curation recipe, report how much text is kept from each source, and verify the shards and benchmark exclusions.

Scope, cost, and ownership

What we need from you

  • Model stage, tokenizer, and target corpus composition
  • A source inventory and a named owner for rights review
  • Evaluation benchmarks or exclusion lists to protect

What affects cost

  • Raw corpus size and parsing or requirements
  • Deduplication scale, mixture experiments, and compute environment
  • Number of corpus revisions and source mixture comparisons

Technical scope

  • Exact and near-duplicate detection sized to the corpus
  • Resumable processing, shard manifests, and tokenizer version records

Support and maintenance

We version source inventories and recipes, and recheck permissions, corpus composition, and exclusions whenever new collections are added.

Common questions

Do you train the foundation model too?

Not as part of this service, which delivers the corpus and the curation workflow. We can scope training infrastructure, model training, and evaluation separately once compute and engineering requirements are clear.

Can you curate a smaller domain corpus?

Yes. Continued pretraining can use a focused domain collection. We first review whether that training stage is justified, or whether retrieval or instruction tuning would fit your goal better.

Guides and resources

See also

Have a dataset project in mind?

Discuss your dataset