A robot that takes instructions has to hear them in a real room — with echo, machinery, overlapping speakers and an accent it was not trained on. Clean speech data is not found. It is made.
Speech models fail in a particular way: not by producing nonsense, but by producing something fluent and wrong. A mistranscription that reads perfectly is invisible to an automatic quality check and invisible to a reviewer who does not speak the language. This is why transcription remains stubbornly human work at the quality end — and why the languages with the least data are exactly the ones where the errors are hardest to catch.
What this needs from people
The research
Talks & demonstrations
Where HSV fits
HSV runs German ASR transcription and audio editing as standing delivery lines. Both are staffed by trained reviewers rather than a single pass of automatic output, because the failure mode that breaks a speech model is the plausible mistranscription nobody flagged.
See the delivery line