Mira

DPO curation

DPO curation turns preference data into alignment-ready pairs. A model produces two responses to the same prompt; trained annotators choose the better one and write down why. The chosen and rejected responses are paired with the rationale and the annotator's confidence, so a trainer can weight disagreements instead of throwing them away. The quality of that material decides whether alignment training converges or stalls.

The workflow begins with scope. You hand us a prompt distribution and we build a prompt set that matches it, rather than scraping a generic pool. For each prompt the model generates two candidate responses, sampled at a pinned temperature and seed. The annotator never sees which response came from which checkpoint; the order is randomised so presentation bias cannot creep in. Annotators are calibrated against a rubric that encodes your alignment policy before the batch begins, and each judgement carries a rationale and a confidence flag.

Where two annotators disagree, the pair goes to adjudication: a senior reviewer reads both rationales and settles the label with a written note that travels with the pair. What you receive is a DPO-ready JSONL export — prompt, chosen, rejected, rationale, confidence, annotator id, adjudication note — with the seed, sampling parameters and rubric version pinned to the run. Every preference carries its evidence, so the alignment team can audit why the dataset prefers one turn over another.