Corpus and source register
Approved sources with collection versions, permitted-use notes, language and domain coverage, and traceable source identifiers.
We select, clean, and remove duplicates from text collections for pretraining. The mix of sources reflects your model, your permitted uses, and your compute budget.
We agree with you on which sources belong in the corpus and how much each should contribute. Filtering reports show what we removed and why. A corpus for continued pretraining, which extends an existing model’s training on new text, differs from the instruction-and-response examples used for . We specify the collection for the training stage you plan to run.
Approved sources with collection versions, permitted-use notes, language and domain coverage, and traceable source identifiers.
Repeatable extraction, filtering, deduplication, and sampling rules, with counts of rejected text and adjustable proportions for each domain.
Corpus files in shards, estimated token counts for your tokenizer, checks that supplied benchmark material has not leaked into the corpus, and a release manifest.
We match domain coverage, language mix, tokenizer, training budget, and permitted sources to the model run you plan.
We inspect samples of accepted and rejected text so the filters neither remove useful domain material nor keep duplicated, low-quality text.
We run the agreed curation recipe, report how much text is kept from each source, and verify the shards and benchmark exclusions.
We version source inventories and recipes, and recheck permissions, corpus composition, and exclusions whenever new collections are added.
Not as part of this service, which delivers the corpus and the curation workflow. We can scope training infrastructure, model training, and evaluation separately once compute and engineering requirements are clear.
Yes. Continued pretraining can use a focused domain collection. We first review whether that training stage is justified, or whether retrieval or instruction tuning would fit your goal better.