AI training data services

Training datasets
and synthetic data.

We prepare training examples from your source material and generate data for the scenarios you lack. Each release includes the files, a schema, a record of where each example came from, and our quality findings.

PretrainingFine-tuningEvaluationData operations

Where we can help

Dataset services

We build the examples your model is missing, fix problems in the data you already have, and make future releases repeatable. Each project starts with what your model needs to learn.

Custom training dataset development

We prepare and label training examples for your model from approved sources. Each release records where every example came from, which cases it covers, and what changed since the last version.

View service

Pretraining data curation

We select, clean, and remove duplicates from text collections for pretraining. The mix of sources reflects your model, your permitted uses, and your compute budget.

View service

Synthetic data generation services

We generate examples for the scenarios your existing data does not cover. We check each batch against the required schema and review it for the task it is meant to support.

View service

Synthetic test data services

We create fictional records for application tests, staging environments, and demos. The files include relationships, expected outcomes, and deliberate error cases.

View service

Fine-tuning and preference datasets

We prepare instruction, conversation, tool-use, and preference examples for fine-tuning. We review answers and formats, and keep evaluation cases separate from training data.

View service

Data annotation and quality review

We define labeling rules, review annotated examples, and investigate disagreements. The release includes the label definitions and our quality findings.

View service

Evaluation datasets and benchmarks

We build evaluation sets with reference answers or scoring rules for your application. Cases cover routine tasks, known failures, and exceptions that need separate review.

View service

Dataset pipelines and governance

We automate dataset preparation and release, including quality checks, source history, and versioning. The pipeline comes with instructions for corrections and failed runs.

View service

A release you can inspect

Review a sample before the full batch.

We prepare a small batch first, and your team reviews it before you pay to scale up. Problems with examples, labels, or missing cases are easiest to fix at this stage.

Specify

We agree with you on the model task, the sources you may use, the cases to cover, the format, and acceptance criteria.

Prepare

We select and clean sources, generate missing scenarios, and label a pilot batch. Your reviewers check it, and we refine the labeling rules before doing the rest.

Evaluate

We check structure, duplicates, coverage, and labels, then test whether the data helps with your task, using references you approve.

Release

We deliver the version you accept, with its documentation and a repeatable way to refresh it.

What each release includes

A dataset specification, versioned files and a schema, a dataset card describing contents and limits, a record of where each example came from or how it was generated, and our quality findings. When data is split for training and testing, a manifest lists the records in each split. We also agree on who handles corrections and refreshes. Formats can include JSONL, CSV, Parquet, or media with linked manifests.

Why synthetic data

Targeted examples.
Fewer blind spots.

Synthetic data can add the examples your model rarely sees. The value comes from covering the right cases with correct labels, not from a larger file.

Compare what you add

Add the examples you are missing.

This means generating varied examples for known gaps, checking their labels, and keeping trusted source data in the mix. Broader coverage can help a model handle more than the common cases.

Model output quality

Illustrative curves. Not measured results or a performance forecast.

How training data quantity and quality can affect a modelThree schematic curves, not experimental data. Reviewed, targeted examples can expand coverage; repetitive examples can plateau; unreviewed examples can reduce quality. Horizontal axis: 100 to 100,000 training examples on a logarithmic scale. Vertical axis: lower to higher conceptual model output quality, with no numerical scores. Highlighted: Targeted synthetic data. The illustrated curve continues improving as reviewed examples cover additional cases, then levels off. Selected dataset size: 10,000 illustrative examples.HigherLower1001k10k100kTraining examples (log scale)
  • Targeted & reviewed
  • Repeated patterns
  • Unreviewed labels
10,000 examples

Cover the missing cases.

We use errors found during development to pick the rare inputs, edge cases, and variations worth generating.

Check what you teach.

We check answers against rules, source documents, or qualified reviewers, and remove duplicates and contradictions.

Measure on unseen data.

We keep the final test examples out of generation and training. Before scaling up, we compare task quality, failure cases, and cost.

About the illustration and research

The curves are drawn by hand and not fitted to research results, Permadyn benchmarks, or the downloadable sample. The dataset sizes are not recommended targets. Real results depend on the task, model, data mix, training setup, and evaluation.

DataEnvGym studies data generation guided by a model’s errors and weak skills, and reports improvements on the tasks it evaluated. That supports testing targeted generation, not assuming it will improve every model.

Research on recursive model training shows how indiscriminate reuse of generated data can lose information from the original distribution. It does not establish that all synthetic data is harmful.

Domain-specific data

Dataset examples by industry

These are examples of datasets we can scope. Your domain experts decide what is correct, relevant, and appropriate for the intended use.

Private equity

Diligence and portfolio workflows

Labeled document fields, source-linked research examples, and synthetic portfolio-company records for application testing.

Healthcare

Administrative AI

Fictional intake and scheduling scenarios, document classifications, and staff-assistant evaluation sets. Any clinical use needs its own scope with qualified clinical reviewers.

Finance

Research and document review

Reviewed extraction examples, research questions with source-backed answers, and synthetic financial records with explicit calculation and relationship rules.

Travel and hospitality

Guest and service workflows

Guest-request conversations, reservation scenarios, multilingual review sets, and edge cases for service-routing applications.

For healthcare reporting and analytics, visit Permadyn Analytics.

Synthetic data sample

Download a documented sample to see what a release looks like. We generated the records from templates and rules we wrote, with no customer data. The sample is for inspection, not an independently validated training dataset.

Dataset details

Includes the record schema, expected outputs, how each record was generated, a quality report, and SHA-256 checksums.

Splits: 8,400 training, 1,800 validation, and 1,800 test records.

Tested engineering reference

An order-workflow test dataset

This test dataset links fictional accounts, orders, line items, and payments. Automated checks confirm that the records connect and the amounts add up. Deliberately broken records show that those checks catch errors.

2,271fictional records4connected tables6deliberate errors caughtVersionedreproducible release

Relationships

Record IDs are unique, every reference to an account or order points to a record that exists, and the exported tables load into a relational database for validation.

Business rules

Line items add up to order totals. Discounts, illustrative taxes, payment balances, and dates follow written rules.

Broken cases

Separate test cases add a duplicate ID, a missing account, a wrong subtotal, a wrong total, an incorrect payment, and a payment dated before its order. Each one fails its expected check.

Handoff

The full handoff includes JSONL and CSV exports, the schema, checksums, a quality report, and a recipe that regenerates the same data.

We built this reference ourselves from fictional rules and a fixed random seed. It shows that the stated software checks work. It is not a customer result, a model-quality benchmark, or a claim that the data represents a real business.

Questions about AI data

Where should we begin?

With the task your model needs to do, a description of the data you have, and the examples you are missing. We agree on a dataset specification and a pilot sample before setting the full volume. If you already have a dataset, we can assess it first.

Do we need pretraining, fine-tuning, or retrieval data?

It depends on what the model gets wrong, because each solves a different problem. Pretraining and continued pretraining teach broad or domain-specific patterns. Fine-tuning teaches task behavior. Retrieval looks up source material when the model answers, without retraining. We review the application and the evidence before recommending any training-data investment.

Can you work across industries?

Yes. We scope data work for private equity, healthcare administration, finance, and travel and hospitality. For each project, we agree on domain requirements, reviewer qualifications, and permitted uses.

How do you price dataset work?

We price each project after the data specification and pilot. Cost depends on source preparation, record length and media types, generation or compute costs, reviewer expertise, quality targets, and how often the data is refreshed. We agree on deliverables and commercial terms before production.

Can you deliver in our infrastructure?

Yes, after an access and architecture review. We can build pipelines for your own storage and compute environment. Before we start, we agree on any model services, data movement, and who operates what.

Does a clean dataset guarantee a better model?

No. A dataset can be well structured and cover the right cases without improving the model. We set pass-or-fail checks for the dataset and run a separate evaluation of how the model behaves on your task. We claim only what those measurements show.

Start with one data requirement

What does your AI need to learn?

Tell us the task, available sources, data types, and examples you are missing. If you know the model, approximate volume, or delivery format, include those too.

We review what you need, define a first batch, and agree on deliverables and cost before starting.

See all AI development and integration services

Tell us about your dataset

We do not sell your contact details or add inquiries to marketing lists. Please leave out passwords and sensitive information.