What is synthetic data?

Synthetic data is information that is generated, rather than collected from real events, to represent selected properties, relationships, or scenarios. It can take the form of tables, documents, conversations, images, audio, or simulated events. How useful it is depends on the task it was designed for and the checks applied to it.

Synthetic data examples: training, testing, and evaluation

The same file format can do very different jobs. Training data, such as an instruction dataset, teaches a model how to behave. Test data checks that an application works. An evaluation set measures how well an AI system performs. Decide which job the data must do before choosing a generation method or a record count.

Illustrative synthetic data requirements by use: proposed specifications, not customer results
UseExampleAcceptance check
AI trainingDocument questions paired with reviewed answersAnswer correctness, coverage, and separate model evaluation
Software testingFictional orders and invoices with linked line itemsValid keys, reconciling totals, and expected workflow behavior
AI evaluationRequests with missing or conflicting informationIndependent references, scoring rules, and test cases kept out of training
Product demonstrationA fictional hotel with reservations and service requestsA coherent scenario and clear labeling that records are fictional

In an invoice workflow, a useful test pack might contain a valid invoice, a duplicate, a missing purchase order, and a mismatched total. Each needs an expected result, so the test can catch a software change that breaks existing behavior.

See our synthetic test data services

How synthetic data generation works

Rules and templates

Records are generated from explicit schemas and rules. For example, a script creates fictional customers and orders, then calculates invoice totals from the line items. This gives a team repeatable test data and controlled exceptions. It does not show that the data represents a real population.

Statistical and learned tabular models

A synthesizer, which is a statistical or machine learning model, can learn patterns from approved real data and generate new records with similar patterns. Relationships, rare groups, constraints, and usefulness for the intended task still need to be evaluated. DataCebo’s Synthetic Data Vault documentation describes tabular synthesis for single-table, sequential, and relational data, along with quality measurement and configurable constraints.

LLM-generated text and conversations

A large language model (LLM) can create task examples or turn approved source material into questions and answers. NVIDIA’s NeMo Curator documentation distinguishes generation from transformation and includes multilingual questions, paraphrasing, and knowledge extraction. Generated responses still need checks for correctness, relevance, duplication, and claims the source does not support.

Simulation

A simulation produces examples from a defined system or environment, such as inventory events under a proposed set of rules. Its assumptions determine what it can represent. A simplified model of the process may leave out real-world behavior, so the outputs need review by domain experts and validation against the intended use.

How do you evaluate synthetic data?

Evaluate structure, coverage, task performance, and privacy separately. A file that loads correctly may still leave out a required category or produce poor results in the target application.

  • Structure: validate types, required fields, keys, ranges, and relationships.
  • Coverage: examine ordinary cases, difficult cases, languages, categories, and missing groups.
  • Task performance: measure the application or model that uses the data against an agreed baseline and independent references.
  • Source and privacy review: record which inputs were approved, and check whether sensitive source content appears in the generated output.

Synthetic data is not automatically anonymous. A generator can reproduce supplied examples, and realistic output is not proof of privacy. Privacy requirements need explicit controls and their own evaluation. Report those findings separately from claims about statistical similarity or model performance.

Evaluation data must also stay separate from training data. If related source documents or near-duplicate examples appear in both, a test score can overstate how well the model handles new cases. Keep related examples together in one set, and document which overlap checks are possible with the data you can access.

What should a synthetic dataset delivery include?

Ask for the specification, the files and schema, the generation recipe or a record of where the data came from, a quality report, known limitations, and who owns future corrections. If your team must be able to regenerate the data, agree on access to the code, configuration, model dependencies, and permitted inputs before work begins.

A dataset card is a short document that describes the data, its intended use, and its limits. Hugging Face’s dataset-card guidance explains how a card supplies context and records metadata such as license, language, and size. The card should also record license restrictions inherited from the sources and the checks used to assess accuracy.

Synthetic data sample

Download 12,000 synthetic records with a schema, dataset card, quality report, and checksums. We generated them from templates and rules we wrote, with no customer data. The sample is for inspection, not an independently validated training dataset.

View and download the documented sample

When to commission custom synthetic data

Custom synthetic data is worth commissioning when your team can name the scenarios that are missing, or when producing the data you need keeps holding up the work. Start with a representative pilot and clear acceptance criteria. If approved data you already have, or a small hand-built test file, solves the problem, that may be all you need.

Permadyn’s synthetic data generation services cover task design, generation, validation, and release. For work that also involves preparing real source data, labeling, or evaluation sets, see our AI training data services.

Discuss your data requirement