Model evaluation
Model evaluation runs a candidate against task suites tuned to where it will actually deploy. A voice assistant shipping in four languages is not graded on a single English benchmark; it is graded on the slice that matters — noisy calls, accented speakers, code-switched turns. Evaluation that does not match the deployment is a number that looks clean and tells you nothing.
The first step is building the suite, not running it. We map the failure modes that would actually hurt the product: a missed wake word, a hallucinated slot, a code-switched turn the model drops, a noisy street call it cannot recover from. Each failure mode becomes a slice with a gold answer set and a scorer. Each run is frozen — prompts, seed, gold answers, scorer and model checkpoint versioned together under one run id — so a number from this week means the same thing as one from six months ago.
Reporting is slice-level, and regressions are surfaced as the slice that moved, not a flattened average. A model that holds its overall score but drops twelve points on accented English is reported as regressed on that slice, with the delta shown. What you receive is a versioned evaluation report: suite definition, run id, per-slice scores with confidence intervals, the regression map and the full evidence trail behind every scored example. The point of evaluation is not a leaderboard position; it is a decision you can make with the numbers in front of you.