Evaluation coverage map
A table of core tasks, user groups, failure types, and high-stakes exceptions, with an agreed plan for sampling cases.
We build evaluation sets with reference answers or scoring rules for your application. Cases cover routine tasks, known failures, and exceptions that need separate review.
We start with the tasks your application must handle and the mistakes it has already made. Each case gets a reference answer or a scoring rubric. Training examples, development tests, and a protected final test set stay separate, so repeated tuning does not inflate the reported score.
A table of core tasks, user groups, failure types, and high-stakes exceptions, with an agreed plan for sampling cases.
Versioned examples with expected outputs or reviewer rubrics, the source of each case, difficulty tags, and scoring instructions.
Access rules for the held-out test set, checks that test cases have not leaked into training data, baseline comparison instructions, and a report by task and failure type.
We define which comparisons the evaluation must support and which error types need their own results before deployment.
We select representative cases, add targeted difficult examples, and check reference answers independently of any model-generated answers.
We check for overlap with available training material, document scoring uncertainty, and agree on when a new held-out set is needed.
We keep held-out tests out of training, investigate score drift, and version new cases as real application failures appear.
Yes, especially for targeted edge cases. We label them as synthetic and keep them distinct from representative real-world examples. Synthetic cases alone do not establish performance in production.
No. We check whether test cases overlap with the datasets and benchmark materials available to the project. We cannot know what an external model saw in pretraining data its developer has not disclosed.