Fine-tuning and preference datasets

We prepare instruction, conversation, tool-use, and preference examples for . We review answers and formats, and keep evaluation cases separate from training data.

ScopeA dataset for one defined task or behavior
Delivery timingReviewer capacity and model format set the release plan
How we check itThe release includes format checks and review results. A separately scoped training comparison measures behavior, cost, and regressions.

Teach the behavior your application needs.

We define and review the responses your application should produce, including formatting, tool inputs, and when to ask a person for help. The collection can contain example conversations with approved answers, or pairs of responses marked better and worse for preference training.

Good fit

  • Teams adapting an existing model to a repeatable task
  • Teams building AI that need reviewed tool-use and conversation examples

Not the right fit

  • Teams expecting to replace up-to-date, permission-controlled document search
  • Projects that treat unreviewed model output as reliable training examples

What you get

Behavior and review rubric

Definitions of correct responses, tool use, escalation, and response style, plus the criteria reviewers use to resolve ambiguity.

Instruction or preference release

Validated conversations, prompt-response examples, or chosen and rejected response pairs in the format your training tool expects.

Held-out tests and comparison plan

Evaluation examples kept out of training, and a plan for measuring whether the new data improves behavior compared with the existing model.

How it works

Inspect current model failures

We run representative application tasks to see whether failures come from missing knowledge, unclear instructions, or behavior that further training could fix.

Create and review examples

We write or generate candidate examples, reviewers check them against the rubric, and we record disagreements and final decisions.

Check the training format

We check message roles, tool schemas, sequence lengths, and split boundaries before delivering a versioned training collection.

Scope, cost, and ownership

What we need from you

  • Target base model and interface
  • Application examples and explicit desired behavior
  • Reviewers qualified to judge correctness and preferences

What affects cost

  • Conversation length, tools, and domain complexity
  • Number of review passes and preference disagreements
  • Trainer format conversions and planned evaluation comparisons

Technical scope

  • Supervised fine-tuning (SFT) conversations and direct preference optimization (DPO) pairs where appropriate
  • Tool input validation, with related conversation variants kept within one split

Support and maintenance

We track rubric and base-model versions, retire stale examples, and update held-out tests as product behavior changes.

Common questions

Is this the same as pretraining data?

No. Pretraining corpora mainly supply broad text or material. Instruction datasets teach how to respond to tasks, and preference pairs show which response is better under a defined rubric.

Can this support RLHF or DPO?

Yes. We can scope human preference labels and paired comparisons. Reinforcement learning from human feedback (RLHF) with a reward model and direct preference optimization (DPO) have different training requirements. We agree on the dataset format and review process with your training team.

Guides and resources

See also

Have a dataset project in mind?

Discuss your dataset