Speak in your own voice — in a language you don't speak
Here is the part of real-time translation that almost everyone gets wrong, and that almost no one talks about: the voice you hear.
You can have excellent speech recognition and excellent translation, and still end up with a meeting that feels like a machine reading a list. Because the last step — turning the translated text back into sound — is where most tools quietly substitute you with a single generic synthetic narrator. Eight people in the room, one robot voice for all of them. You lose who is speaking, the emphasis, the personality. Intelligible, but not a conversation.
InterMIND does the last step differently. When you speak, the other participants hear the translation in a voice that's recognizably yours — carrying your timbre and your way of speaking — now saying the words in their language. It isn't a flawless impression yet; the point is that it's you rather than a stock narrator, and it's getting better. This works for every participant, in both directions, at the same time.
This post is the missing chapter of Inside the four translation pipelines that run InterMIND: that piece explained how audio becomes translated audio. This one is about whose voice comes out the other end.
The default everyone ships, and why it's flat
If you've used live translation in any of the big meeting platforms, you know the sound. A neutral, evenly-paced voice reads the translation. It's the same voice whether the speaker is your CEO opening a town hall or a colleague cracking a joke. The technology underneath is text-to-speech with one fixed voice model, and the design assumption is that intelligibility is enough.
In a real meeting it isn't. Half of what a meeting communicates is who is saying it and how. Strip the voice and you've turned a discussion into a transcript that happens to be spoken aloud. People stop reacting to each other and start waiting their turn.