ASR · TTS · SLM
Speech annotation covers the whole model stack — ASR, TTS and spoken language models — under one workflow. The same audio can be transcribed, diarised and event-tagged for ASR, marked up for prosody and pronunciation for TTS, and labelled with instructions and dialogue turns for an SLM, without being re-handed between three teams. One clip, one source of truth, three label sets that all agree on where the speaker boundaries are.
The workflow begins with a shared schema, not three separate ones. We build a single annotation pass that produces all three label sets from the same audio against the same boundaries, so a disagreement in one set cannot quietly contradict another. For ASR the audio is transcribed, diarised and event-tagged; for TTS the same turns are marked for prosody, stress and phoneme-level detail; for an SLM the turns are labelled with instructions, dialogue acts and turn boundaries. Each label carries the annotator, the pass and a note.
A disagreement is not quietly averaged — when two annotators split on a boundary, a phoneme or a dialogue act, the clip is escalated, resolved by a senior reviewer and recorded against the span. The export matches your schema, the style guide is yours, and the raw evidence stays attached to every judgement. The point of speech annotation across the stack is one consistent ground truth feeding three models that finally agree on what happened in the clip.