Mira

Benchmarking

Benchmarking begins with a gold reference set — examples labelled, checked and frozen against a versioned style guide. You hand us a candidate model and a target language, and we run it head to head against that gold set under identical conditions. The gold set is the fixed point; it does not move between runs, so a model you ship in March can be compared with the same one rebuilt in July without ambiguity about drift.

Before a single model touches the set, we scope the benchmark to where the product will actually live. A voice assistant shipping in four languages is not graded on a generic English slice alone; the gold set matches the languages, accents, noise levels and domains it will meet in production. Slices that matter are reported separately so a problem in one corner is never averaged into invisibility.

Every run is frozen at the moment it starts: prompts, seed, gold answers, scorer version and model checkpoint are pinned together under one run id. The output is a versioned benchmark report — headline number, per-slice breakdown, regression map and the full evidence trail behind every score. The point of benchmarking is not a leaderboard; it is a measurement you can defend.