Guide

How live translation, captions, and a searchable transcript work together in one InterMIND meeting

Voice translation, live captions and a searchable transcript are turning up together as a single marketed bundle across meeting and workplace platforms. What that bundle actually means differs a lot from vendor to vendor — most product pages state the three features without saying how each behaves: who has to turn it on, what a guest gets, or what "searchable" covers. Here is exactly how each of the three works in an InterMIND meeting, what is on by default, and where the honest limits are.

The Mind.com Team

How live translation, captions, and a searchable transcript work together in one InterMIND meeting

How live translation, captions, and a searchable transcript work together in one InterMIND meeting

A team spread across four time zones does not have a language problem on the calendar invite — it has one the moment somebody starts talking. Live translation, live captions and a transcript you can go back to are the three things that actually remove that friction, and increasingly they show up together on a single product page, described as one bundle rather than three separate add-ons.

One recent example: a September 2026 post from workplace platform NexGen Virtual Workplace opens with "Live voice translation, meeting captions, and searchable transcripts, built into the workday rather than bolted onto the meeting," and states that "these features are available within the NexGen Virtual Workplace platform for subscribers to use." That is the whole specification on offer — the post does not name a supported language count, does not say which subscription tier unlocks the three features, and does not describe what "searchable" means beyond "people can also look up past transcriptions... for a reference" (checked September 2026, source linked at the end). None of that is a criticism of the vendor; it is simply what a reader has to go find out elsewhere, or take on faith, before knowing whether the bundle fits their team.

Since the same three words — translation, captions, transcript — get used for meaningfully different mechanics across products, here is the specific, checkable version of each as it works in InterMIND today.

Translation: you turn it on, and it is your own voice

Translation is not ambient — you enable it deliberately, per meeting, from the More menu in the control bar (or the Alt+T shortcut), available once the room has at least two languages. From there each participant sets their own translation language independently; in a five-person call with five languages, everyone hears everyone else translated into their own, simultaneously.

The output is not a generic narrator voice. Speech recognition, sentence segmentation, translation and voice synthesis run as a pipeline, and the last stage resynthesizes the translated sentence in a version of the original speaker's own voice — recognizably them, not a flat text-to-speech read-out. The technical breakdown of how that pipeline works is in a separate post; the summary here is that once translation is on, "who said it" survives translation, not just "what was said." Coverage is 23 languages for voice, chat and shared notes, and 30 for whole-document file translation — the split and the reasons for it are in the language breakdown.

Captions: on by default, not a separate toggle to remember

Live captions in InterMIND are the translated text of the same stream, shown as a subtitle banner during the meeting. Two details matter more than the feature existing at all:

  • The default is on. A participant who changes nothing already sees the translated speech on screen once translation is enabled — there is no second switch to discover mid-meeting.
  • The banner remembers your last choice, per browser, so turning captions off in one meeting keeps them off in the next until you turn them back on, and it renders right-to-left automatically when the target language calls for it.

Captions are a read channel on top of the same translated audio, not a separate product with its own settings page to configure in advance — which is the opposite of the "who gets to see which caption language" question that dominates on platforms where caption languages are an organizer-side setting picked before the event.

Transcription: automatic, and specific about what "searchable" covers

Transcription has no Start/Stop button. From the moment a meeting is live, every word is captured, attributed to the speaker who said it, and stored on the server in the language it was actually spoken in — nobody has to remember to turn it on, and nobody can forget to.

After the meeting, a recap message appears in the meeting chat; opening it and switching to the Transcript tab shows the full, speaker-attributed, timestamped transcript, translatable on demand into the reader's own language even though it was captured in the speakers' original ones.

"Searchable" is the word that needs the most precision, so here is exactly what it means today, without rounding up:

  • Within the meeting's chat, it is real text search. Each turn of speech lands as a transcription message in the channel as the meeting runs, and those messages are covered by the same chat search every other message is — Ctrl/⌘+F opens it, you type a keyword, and matching messages jump into view with context.
  • The transcript page itself is plain rendered text, so your browser's own page search finds a word or a name on it exactly as it would on any web page.
  • What is not yet full-text search is the AI-written summary that sits above the transcript in the recap — today that summary surfaces by its title, not by searching its body. If what you are hunting for is a specific sentence someone said, the transcript tab is where it lives and where search reaches it; the short summary is a separate object with a narrower search surface, and closing that gap is on our roadmap.

That distinction — a spoken-word archive that is genuinely searchable, next to a short AI summary that currently is not — is the kind of specific behavior a one-line "searchable transcripts" bullet on a product page tends to skip past.

One seat, one plan limit — not three licenses

None of the three is gated behind its own add-on or its own license tier. What scales with the plan is the size of the room: 50 participants on Basic, 100 on Pro, 300 on Business, 1500 on Enterprise. A host on any plan gets translation, captions and transcription together; upgrading changes how many people can be in the room, not which of the three features they get.

What a guest gets, and what they do not keep

A participant who joins through a meeting link without ever creating an account gets the full in-meeting experience — their own translation language (defaulting to their interface language, which itself follows the browser), captions in it, and their words captured into the shared transcript like anyone else's. What they do not get is anything afterward: there is no account to deliver a recap to, and the anonymous identity is not retained once the call ends. An account is free, and it is specifically what makes the after-the-call half — the recap, the searchable transcript, the channel history — reachable later. That is a real limit, worth stating plainly rather than glossing over, since it is the opposite of the parts of this feature set that sound unconditional.

See it instead of reading about it


FAQ

Do I need to turn on translation, captions and transcription separately? Transcription is always on, automatically, from the moment a meeting starts. Translation is a deliberate choice — enabled from the More menu or Alt+T, once the room has at least two participant languages. Captions are the translated text of that same stream and default to on the moment translation is enabled, so in practice there is one switch to think about, not three.

Can I actually search a meeting transcript, or just scroll it? Both. The transcript is captured turn by turn as chat messages during the meeting, and those are covered by the same chat search (Ctrl/⌘+F) as every other message — type a keyword and matching turns jump into view. The full transcript page is also plain text, so a browser's own page search works on it directly. The one thing not yet covered by full-text search is the body of the short AI-written meeting summary, which today surfaces by its title.

Does everyone hear the same translated voice, or does it keep the speaker's own voice? Each participant's translated speech is resynthesized in a voice modeled on their own, not a shared generic narrator — the effect is a translated conversation that still sounds like the people having it, in both directions and for every participant independently.

What happens to the transcript and recap if someone joins as a guest, with no account? A guest gets full in-meeting translation, captions and transcription like any other participant. What does not carry over is anything afterward: there is no account for the recap or transcript to go to, and the guest's anonymous identity is not kept once the meeting ends. Signing up (free) is what makes the after-the-call record reachable.

Is there a separate plan or add-on for translation, captions or transcription? No — all three ship together on every plan. What changes between plans is the number of participants a host can host at once: 50 on Basic, 100 on Pro, 300 on Business, 1500 on Enterprise.


Sources: NexGen Virtual Workplace — Speaking the Same Language: How Translation and Transcription Work Inside NexGen Virtual Workplace, checked September 2026. InterMIND facts: real-time translation docs, transcription docs, language breakdown, own-voice translation, live benchmark.

Get new posts and product updates by email

One email a month with new posts and product updates. Unsubscribe anytime.