Inside the four translation pipelines that run InterMIND
The old /product/overview/how-it-works page on mind.com is several major releases out of date. It describes a single "translation engine" the way most vendor pages do — one big arrow from "you speak" to "they hear." That picture was already a simplification two years ago. Today it is wrong.
The truth is that InterMIND runs four separate translation pipelines, each solving a different problem with a different engine, a different latency budget, and a different quality envelope. They share a language picker. They do not share an engine.
This is the updated answer to "how does it work."
A companion piece: "How many languages do you support?" covers what each pipeline covers (23 / 23 / 30 / 17). This post covers what each pipeline does — and why it is its own thing.
Why "one engine for everything" is a lie
A live meeting platform has at least four jobs to do at once, and they pull in incompatible directions:
- Real-time voice — audio in, translated audio out, under one second, every viewer in their own language. The hard constraint is latency.
- Real-time chat text — short messages, fast, with edits and quotes and HTML structure preserved.
- Real-time shared notes — character-by-character collaborative typing, with structural hierarchy (lists, headings, checkboxes) that has to survive translation.
- Asynchronous document files — a 40-page PDF dropped into chat. No latency budget. The hard constraint is fidelity — formatting, tables, page numbers, font.
You can build one giant LLM call that tries to do all four. We tried. It is bad at all four. The latency budget for voice means the model can't think; the fidelity budget for documents means the model has to. A chat edit needs a diff in the viewer's language; a 40-page PDF needs format preservation that no token-streaming model gives you.
So we run four. Here is each one.