# Signing In ## Sign-In Methods InterMIND supports the following sign-in methods: | Method | How | | ---------------- | --------------------------------------------------------------------- | | **Google** | Click **Continue with Google** → select account | | **Microsoft** | Click **Continue with Microsoft** → sign in | | **Email (code)** | Enter email → click **Continue with Email** → enter the 6-digit code | | **SSO** | Click **Sign in with SSO** → enter work email → continue via your IdP | All methods lead to the same account if you use the same email address. There is no password-based sign-in — every email login uses a one-time code. ## Staying Signed In Once signed in, your session persists across browser tabs and restarts. You don't need to sign in every time you visit InterMIND. ## Signing Out 1. Open the sidebar (click the hamburger menu on mobile, or see the sidebar on desktop) 2. Click your profile avatar at the bottom of the sidebar 3. Select **Sign Out** ## Troubleshooting | Issue | Solution | | ------------------------------- | --------------------------------------------------------------------------------------- | | "Account not found" | Make sure you're using the same email and sign-in method you used to create the account | | Google/Microsoft popup blocked | Allow popups for the InterMIND domain in your browser settings | | Verification code didn't arrive | Check your spam folder, then click **Resend code** | | Code rejected as invalid | Codes are 6 digits and expire — request a fresh code with **Resend code** | | Session keeps expiring | Check that cookies are enabled and not blocked by browser extensions | ## What's Next - [Start your first meeting](https://intermind.com/first-meeting) — Create a meeting and invite participants - [Configure your profile](https://intermind.com/../settings/profile) — Set your name, avatar, and language # Creating an Account You can create an InterMIND account using any of the following methods. ## Sign Up with Google ::steps ### Go to the login page Open the InterMIND login page. ### Click Continue with Google A popup will appear for Google account selection. ### Select your Google account Choose the account you want to use. ### Done You're signed in — your profile name and avatar are imported automatically. :: ## Sign Up with Microsoft ::steps ### Go to the login page Open the InterMIND login page. ### Click Continue with Microsoft A popup will appear for Microsoft sign-in. ### Sign in with your Microsoft account You can use a personal, work, or school account. ### Grant permissions Approve the requested permissions. ### Done You're signed in — your profile name is imported automatically. :: ## Sign Up with Email (Verification Code) InterMIND email sign-in is passwordless — you receive a 6-digit code in your inbox each time. ::steps ### Go to the login page Open the InterMIND login page. ### Enter your work email Type your email address in the **Work email** field. ### Click Continue with Email A 6-digit verification code is sent to that address. ### Check your inbox Open the email from InterMIND and copy the 6-digit code (or follow the confirmation link to enter the code on a second device). ### Enter the code Paste the code into the **Verification code** field and click **Verify Code**. ### Done You're signed in. If it's your first time, an account is created automatically. :: ## Sign In with SSO If your organization has configured SSO, click **Sign in with SSO** on the login page and enter your work email. You'll be redirected to your identity provider. ## What Happens After Sign-Up After creating your account: - You start on the free **Basic** plan — no card required - A team is automatically created for you - You can immediately [start your first meeting](https://intermind.com/first-meeting) - A **14-day free trial** of a paid plan is available when you subscribe through the [Pricing](https://intermind.com/pricing) page (one trial per account — see [Billing & Plans](https://intermind.com/../billing)) ## Joining via Team Invitation If someone invited you to their team: ::steps ### Click the invitation link Open the link in the email you received. ### Sign in or create an account Use any of the methods described above. ### You're added to the team You are automatically added to the team that invited you. The team admin's subscription covers your access. :: ::callout{type="info"} Each user belongs to **at most one team** at a time. Accepting a new team's invitation leaves your current team (an owner with an active subscription must cancel and wait for it to end first). :: # Your First Meeting This guide walks you through creating and running your first meeting with real-time translation. ::steps{level="2"} ## Start a Meeting After signing in, open the **Meetings** page from the sidebar: - **New meeting** — Click **New meeting**. A meeting is created and its link is copied. Click the **Join** button next to the meeting in the list (or open the copied link) to enter. - **Join or invite** — Paste a meeting ID or URL into the search field to join — or type a contact's name or email to send them an invite. - **Schedule for Later** — Click the **calendar button** on the meeting's row to open Google Calendar or Outlook with the meeting link pre-filled. ## Pre-Join Setup Before entering the meeting room, you'll see the **Pre-Join Setup** screen: - **Camera preview** — See how you look before joining - **Microphone selection** — Choose which microphone to use - **Camera selection** — Choose which camera to use - **Mute/unmute** — Start with camera or microphone off Click **Join** when you're ready. :::callout{type="tip"} You can disable the pre-join screen in [Settings → Meeting Defaults](https://intermind.com/../settings/meeting-defaults) if you prefer to join instantly. ::: ## Enable Translation Once in the meeting: 1. Open the **More** menu in the control bar at the bottom of the screen 2. Click **Enable translation** (language icon) to turn on real-time translation — this item appears only when the meeting has more than one language in play 3. Set your **preferred language** in the cockpit **Settings** panel — this is the language you want to hear 4. When other participants speak, their words are translated into your chosen language in real-time You'll hear the translated audio automatically. You can also enable **subtitles** to see the translation as text. ## Invite Others Share the meeting with others: - **Copy the meeting link** from the address bar or the share button - **Send it to participants** via email, chat, or any messaging app - Participants with accounts can **click the link and join directly** - Participants without accounts can **join as guests** (the meeting host must approve them) ## Use Meeting Controls The control bar at the bottom provides quick access to: | Control | Function | | -------------- | ------------------------------------ | | 🎤 Microphone | Mute/unmute your microphone | | 📹 Camera | Turn camera on/off | | 🖥️ Screen | Share your screen | | 🌐 Translation | Enable/disable real-time translation | | 💬 Subtitles | Show/hide live subtitles | | 👋 Reactions | Send emoji reactions or raise hand | | 💬 Chat | Open the in-meeting chat sidebar | | 🚪 Leave | Leave the meeting | For details on each control, see [Meeting Controls](https://intermind.com/../meetings/controls). :: ## What's Next - [Learn about meeting controls](https://intermind.com/../meetings/controls) — Camera, microphone, screen sharing, and more - [Set up translation](https://intermind.com/../translation) — Choose languages and configure subtitles - [Use the chat](https://intermind.com/../chat) — Send messages during and outside meetings - [Invite your team](https://intermind.com/../users/inviting) — Add team members to your organization # System Requirements InterMIND runs fully in the browser — no installation, no plugins. Prefer a native app? Optional [desktop and mobile apps](https://intermind.com/apps) are available for macOS, Windows, and Android. ## Supported Browsers | Browser | Minimum version | Notes | | --------------------- | --------------------- | --------------------------------------------- | | **Chrome** | 90+ | Best support | | **Edge** | 90+ | Same engine as Chrome | | **Safari** | 15+ | Screen sharing requires Safari 17+ | | **Firefox** | 100+ | Works; some codec fallbacks on older hardware | | **Opera, Brave, Arc** | Chrome 90+ equivalent | Works (Chromium-based) | Internet Explorer and legacy (pre-Chromium) Edge are not supported. ::callout{type="tip"} For the best experience, use Google Chrome or Microsoft Edge with a stable internet connection. :: ## Bandwidth | Activity | Minimum | Recommended | | ----------------------------------- | -------------- | -------------- | | Audio-only meeting with translation | 1 Mbps up/down | 2 Mbps up/down | | HD video meeting with translation | 2 Mbps up/down | 5 Mbps up/down | | Screen sharing | +1 Mbps up | +2 Mbps up | Run a [speed test](https://fast.com){rel=""nofollow""} to check your real bandwidth — not the rated speed of your plan. ## Devices | Device | Required | | -------------------------- | ---------------------------------------------------------------------------------- | | **Microphone** | Required for speaking | | **Camera** | Optional — audio-only mode is supported | | **Speakers or headphones** | Required to hear translation | | **Headset** | Recommended in noisy environments — reduces echo and improves translation accuracy | ## Mobile | Feature | iOS Safari | Android Chrome | | ---------------------- | ----------- | -------------- | | Audio + video meetings | ✓ | ✓ | | Real-time translation | ✓ | ✓ | | Screen sharing | ✗ | ✗ | | Picture-in-picture | ✓ (iOS 15+) | ✓ | For meetings longer than 30 minutes or with screen sharing, use desktop. On Android, the [InterMIND app](https://intermind.com/apps) keeps the call alive in the background and adds push notifications. ## Verify Before an Important Meeting Run the [WebRTC Browser Test](https://intermind.com/browser-test) — it checks microphone, camera, network reachability, and codec support in 30 seconds. If any check fails, see [Troubleshooting](https://intermind.com/../troubleshooting). ## Related - [Desktop & Mobile Apps](https://intermind.com/apps) — Optional native apps for macOS, Windows, and Android - [Troubleshooting → Browsers & Devices](https://intermind.com/../troubleshooting/browsers) — When the requirements are met but something still doesn't work - [Troubleshooting → Network & Connection](https://intermind.com/../troubleshooting/network) — Firewall and VPN settings # Desktop & Mobile Apps InterMIND works fully in the browser — the apps are optional. Install one when you want the call to survive a minimized window, native message and call alerts, or a home-screen icon on your phone. Get them from the [Downloads page](https://intermind.com/downloads): | Platform | Package | Requirements | | ----------- | ------------------------------------- | ------------------- | | **macOS** | Universal DMG (Apple silicon & Intel) | macOS 11 or later | | **Windows** | x64 installer | Windows 10 or later | | **Android** | Google Play (early access) | Android 8 or later | The desktop builds are signed, so they open without security warnings. An iOS app is not yet available — use Safari on iPhone or iPad (see [System Requirements](https://intermind.com/system-requirements)). ## What's Different from the Browser Inside a meeting, the apps and the browser are identical — same translation engine, same controls. The differences are around the meeting window: | Capability | Web | Desktop | Mobile (Android) | | ------------------------------- | ------------------------- | ------------------------------------------- | ---------------------- | | Join & translate live meetings | Full | Full | Full | | Screen sharing | Full | Full | Not yet | | Call stays alive when minimized | Tab is throttled | Runs unthrottled | Foreground service | | Native notifications | In-tab only | System + dock badge | Push notifications | | Updates | Always latest, no restart | Downloads in background, applies on restart | Applies on next launch | ## When to Use Which - **Web** — no install, always the latest version. The default. - **Desktop** — you keep InterMIND open all day: background calls aren't throttled, and alerts arrive even when the window is hidden. - **Mobile** — meetings on the go with push notifications; the foreground service keeps the call running when you switch apps. ## Related - [System Requirements](https://intermind.com/system-requirements) — Browsers, bandwidth, and devices for the web version - [Joining a Meeting](https://intermind.com/../meetings/join) — Works the same in browser and apps # Welcome to InterMIND InterMIND is a video conferencing platform with **built-in AI-powered real-time translation**. Every participant speaks their own language and hears others translated into theirs, live during the call. ## What You Can Do - **Video Meetings** — Start or join HD video conferences with up to 1,500 participants - **Real-Time Translation** — Hear live spoken translation in 24 languages - **Live Subtitles** — Read translated subtitles during calls - **Transcription** — Save full transcripts of meetings - **Team Chat** — Send messages, files, and images with auto-translation (document translation is a paid feature, Pro and above) - **Screen Sharing & Recording** — Share your screen and record meetings - **Guest Access** — Invite anyone to join a meeting without an account ## Quick Start 1. **[Create an account](https://intermind.com/getting-started/account)** — Sign up with Google, Microsoft, or email 2. **[Sign in](https://intermind.com/getting-started/sign-in)** — Access the platform 3. **[Start your first meeting](https://intermind.com/getting-started/first-meeting)** — Create a meeting and invite participants ## Supported Languages InterMIND supports **24 languages** for real-time translation (voice, chat, and shared notes): | Language | Language | Language | | -------- | --------- | ---------- | | English | Spanish | French | | German | Italian | Portuguese | | Dutch | Russian | Polish | | Czech | Hungarian | Japanese | | Korean | Chinese | Danish | | Finnish | Icelandic | Norwegian | | Romanian | Swedish | Ukrainian | | Turkish | Arabic | Hindi | On-demand file translation in chat — a paid feature available on Pro and above — covers a wider list of 30 languages via DeepL. The marketing site itself is localized to 16 (two of which, Indonesian and Vietnamese, are site-only for now — in-meeting voice doesn't cover them yet). See [Choosing Languages](https://intermind.com/translation/languages) for the full per-surface breakdown, or read [How many languages do we support — six numbers, not one](https://intermind.com/blog/how-many-languages-do-you-support) for the methodology. ## System Requirements InterMIND runs fully in the browser — no installation needed. Optional [desktop and mobile apps](https://intermind.com/getting-started/apps) add background calls and native notifications. See [System Requirements](https://intermind.com/getting-started/system-requirements) for the supported browsers, bandwidth, and devices. # Choosing Languages ## Selecting Your Language Your translation language determines what language you hear when translation is active. ### During a Meeting 1. Click the **Language** button in the control bar 2. A language picker opens with 24 available languages 3. Search for your language by typing its name 4. Click to select the language 5. Translation immediately switches to the new language ### In Settings By default your translation language follows your **interface language** (set from the language switcher in the app header). To use a different language for translation only: 1. Go to **Settings** → **User** card 2. Toggle **Use a different language for translation** on 3. Pick your translation language from the dropdown 4. The hint changes to **"Different from interface language"** — your choice is saved to your profile and syncs across devices This single setting is used for: - Real-time voice translation in meetings - Automatic chat message translation - Subtitles language ## Supported Languages InterMIND currently supports **24 languages** for real-time translation (voice, chat messages, and shared notes — one set, three surfaces): | Language | Native Name | Region | | ----------------------- | ----------- | -------- | | English | English | Global | | Russian | Русский | Europe | | Spanish (Latin America) | Español | Americas | | French | Français | Europe | | German | Deutsch | Europe | | Italian | Italiano | Europe | | Portuguese (Brazilian) | Português | Americas | | Polish | Polski | Europe | | Dutch | Nederlands | Europe | | Czech | Čeština | Europe | | Hungarian | Magyar | Europe | | Danish | Dansk | Europe | | Finnish | Suomi | Europe | | Icelandic | Íslenska | Europe | | Norwegian | Norsk | Europe | | Romanian | Română | Europe | | Swedish | Svenska | Europe | | Ukrainian | Українська | Europe | | Japanese | 日本語 | Asia | | Korean | 한국어 | Asia | | Turkish | Türkçe | Europe | | Chinese (Simplified) | 中文 | Asia | | Arabic | العربية | MENA | | Hindi | हिन्दी | Asia | ::callout{type="info"} **File translation in chat uses a wider list of 30 languages** (DeepL Document API) — including Bulgarian, Greek, Estonian, Indonesian, and several other European languages not available for live voice. The platform API accepts all 24 real-time codes plus the file-only ones. Per-pair quality — including newer pairs like Arabic and Hindi — is published openly on [`/benchmark`](https://intermind.com/benchmark). Full breakdown: [How many languages do we support — six numbers, not one](https://intermind.com/blog/how-many-languages-do-you-support). :: ::callout{type="info"} All 24 languages are available on every plan — pair availability is not limited by tier. Monthly translation time and text-translation word budget depend on your plan; see [Usage & Limits](https://intermind.com/../billing/usage) for details. :: ## Tips - **Set your language before the meeting** in Settings to avoid configuring it every time - **Browser language fallback** — If you haven't set a language, InterMIND uses your browser's default language - The language setting is **per-user**, not per-meeting. Each participant in the same meeting can use a different language # Live Subtitles Live subtitles show the translated text of what participants are saying, displayed at the bottom of the meeting screen in real-time. ## Enabling Subtitles 1. Join a meeting with translation enabled 2. Click the **Subtitles** toggle in the meeting control bar 3. Translated text appears at the bottom of the screen as participants speak 4. Click again to hide subtitles ## How Subtitles Work - Subtitles show the **translated** text, not the original speech - They appear in the language you selected as your translation language - Subtitles update in real-time as words are recognized and translated - Each subtitle line shows what the current speaker is saying - When no one is speaking, subtitles fade away ## RTL Language Support The subtitle layout automatically flips to right-to-left when an RTL language is detected, so future RTL languages will render correctly. The current translation target list is left-to-right (see [Choosing Languages](https://intermind.com/languages) for the full list). ## Subtitle Appearance - Subtitles appear as a **banner at the bottom** of the meeting screen - They have a semi-transparent dark background for readability - Font size is optimized for screen reading at typical viewing distances - Subtitles don't overlap with the meeting control bar ## Default Setting Subtitle on/off state is **remembered between meetings** based on whatever you last set in the cockpit (per browser). There is no separate Settings toggle — the last state of the **Subtitles** button is what you'll get when you join the next meeting. ## Tips for Using Subtitles - **Presentations** — Enable subtitles during presentations for participants who prefer reading - **Noisy environments** — Subtitles help when audio quality is poor - **Language learning** — Read the translation while hearing the original to learn new languages - **Accessibility** — Subtitles make meetings more accessible for participants with hearing difficulties # Transcription Transcription captures all spoken words during a meeting and makes them available as a transcript after the meeting, with speaker attribution. ## How Transcription Works Transcription is **automatic** — there is no Start/Stop button. Whenever a meeting is live, speech is captured in real time: - Every word is transcribed and attributed to the speaker who said it - Words are collected on the server as the meeting runs - The transcript is in each speaker's original language ## Viewing the Transcript After the meeting ends, a **recap** message appears in the meeting chat. Open it and switch to the **Transcript** tab to read the full transcript on its own page — see [Meeting Recap](https://intermind.com/../meetings/recap). The transcript page can show the transcript **translated into your language** — it's generated in the speakers' original languages and translated on demand when you open it in a different language. ## Use Cases - **Meeting minutes** — Capture everything said for reference - **Compliance** — Keep records of discussions for regulatory purposes - **Absent participants** — Share the transcript with people who couldn't attend - **Review** — Go back and review specific parts of the discussion ## Important Notes - The quality of transcription depends on audio quality — use a good microphone and minimize background noise - The transcript is captured in each speaker's original language; open the transcript page in your own language to read it translated - The transcript is produced once the meeting ends and reached from the recap in the meeting chat ## Related - [Meeting Recap](https://intermind.com/../meetings/recap) — Summary, transcript, and recording after the call - [Meeting Controls](https://intermind.com/../meetings/controls) — In-meeting control bar - [Chat](https://intermind.com/../chat) — Where the post-meeting recap lives # Your Own Voice When InterMIND translates your speech for another participant, they don't hear a robotic text-to-speech narrator. They hear a voice that's **recognizably yours** — carrying your timbre and your way of speaking — now saying the words in their language. This works in both directions and for every participant independently. In a meeting where five people speak five languages, each person hears the other four in their own language, and each of those four still sounds like themselves. ## What It Sounds Like Most live-translation tools replace the speaker with a single generic synthetic voice. The result is intelligible but flat — you lose who is talking, the emphasis, the personality. InterMIND keeps the speaker's voice, so a translated meeting feels like a conversation between the people who are actually in it, not a queue of announcements read by a machine. ## How It Works InterMIND uses a **cascaded pipeline**, and the voice step is the last stage: 1. **Speech recognition** — your words are transcribed in your own language, as you speak. 2. **Segmentation** — the transcript is grouped into stable sentence fragments (clauses) so translation can begin before you finish the sentence. 3. **Translation** — each fragment is translated progressively into the listener's language. 4. **Voice synthesis** — each translated fragment is spoken back using a sample of **your own voice**, and sent to the listener. While the meeting is still gathering enough of your speech to model your voice (roughly the first **5–10 seconds**), the synthesis uses the audio fragment that matches what you just said in your source language. Once there's a long enough sample, it switches to using that sample for everything after. In practice you don't notice a switch — the translation sounds more like you as the call goes on. It won't be a flawless impression of your voice, but it's recognizably you rather than a generic narrator — and it keeps improving as the model hears more of you. ## Languages Your-own-voice translation is available for **all 24 voice languages** — the same set listed in [Choosing Languages](https://intermind.com/languages). There's nothing to enable separately: when [translation is on](https://intermind.com/../translation), participants automatically hear you in your own voice. ## Privacy The voice sample used for synthesis is **ephemeral**. It exists only for the duration of the live meeting and is **not stored anywhere** — the Mind API and SDK that power the real-time session keep no data once the conference session ends. This voice sample is unrelated to InterMIND's video-and-voice **recording** features, which are separate, explicit recordings you start on purpose. ## On the Roadmap: Lip-Sync Hearing the translation in your own voice is the first half of a larger goal. The next step we're working toward is **lip-sync** — re-timing the speaker's mouth on camera to match the translated audio, so each participant appears to be speaking the other's language. Combined with own-voice translation, the aim is a call where people who share no common language see and hear each other as if each spoke the other's language natively. This is a roadmap item, not a shipped feature yet — own-voice translation above is live today. ::callout{type="info"} Want the full technical picture? See [Speak in your own voice — in a language you don't speak](https://intermind.com/blog/own-voice-translation) on the blog. :: # Real-Time Translation InterMIND's core feature is **AI-powered real-time speech translation**. During video meetings, every participant speaks their own language and hears others translated into theirs — live, with minimal delay. ## How It Works 1. You speak in your own language 2. InterMIND's AI captures your speech 3. The speech is translated into the target language of each listener 4. Each participant hears the translated audio in real-time — [in the original speaker's own voice](https://intermind.com/translation/own-voice), not a synthetic narrator 5. Optionally, translated subtitles appear on screen Each participant independently chooses their own language. In a meeting with 5 people speaking 5 different languages, everyone hears the translation in their own language simultaneously. ## Key Features | Feature | Description | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **24 languages** | Voice, chat, and shared notes (file translation: 30, a paid feature on Pro and above) — [why these counts vary](https://intermind.com/blog/how-many-languages-do-you-support) | | **Real-time** | Minimal delay between speech and translation | | **Per-participant** | Each person chooses their own target language | | **Bidirectional** | Works in both directions — you hear others and they hear you | | **Own voice** | Each participant hears the translation in the [speaker's own voice](https://intermind.com/translation/own-voice), not a robotic narrator | | **Subtitles** | Optional text subtitles for the translation | | **Transcription** | Save the full transcript of a meeting | ## Enabling Translation 1. Join a meeting 2. Open the **More** menu in the control bar and click **Enable translation** (or press Alt+T) — only available when the room has at least two languages 3. Set your **Translation Language** in the cockpit settings panel (gear icon in the More menu) 4. Start speaking — other participants will hear you in their chosen language When translation is enabled: - You hear a **translated audio stream** instead of the original - The translation replaces the original audio (the original is not played simultaneously) - You can toggle it off at any time to hear the original speech ## In This Section - [Choosing Languages](https://intermind.com/translation/languages) — Select and change your translation language - [Live Subtitles](https://intermind.com/translation/subtitles) — Display translated text during meetings - [Transcription](https://intermind.com/translation/transcription) — Record and save meeting transcripts - [Usage & Limits](https://intermind.com/billing/usage) — Translation time limits by plan, tracked on the Billing page # Starting a Meeting ## Quick Start ![Meetings page](https://intermind.com/img/docs/meetings/meetings-page.png) ::steps ### Navigate to Meetings Go to the **Meetings** page from the sidebar. ### Click New meeting A meeting is created and its link is copied to your clipboard; the meeting appears at the top of the list. The list names the meeting after its participants; use the row's rename action to set your own title. ### Share the link Send it via email, messaging app, or any other channel — or type a contact's name or email into the search field and send them an invite directly (they get an email, and a push notification if they use the mobile app). ### Join the meeting Click the **Join** button next to the meeting in the list (or open the copied link). You're taken to the pre-join setup screen. ### Configure your devices Set up your camera and microphone, then click **Join**. :: ## Scheduling a Meeting Any meeting can be added to your calendar — including one that has already met (the link is permanent, so a one-off room can become a recurring one without losing its chat and recaps): ::steps ### Go to the Meetings page Open the **Meetings** page from the sidebar. ### Click the calendar button on the meeting's row Your configured calendar provider opens in a new tab with the event pre-filled (title, join link, and recap link). Pick the date, invitees, and — for a regular meeting — the recurrence right there. ### Done The calendar event includes the meeting link for participants. The meeting row picks up the date and any later changes from your calendar; once the event is observed, the same button opens it for editing. :: ::callout{type="tip"} The calendar provider (Google Calendar or Outlook) is connected in **Settings → Calendar** — there is no provider picker on the Meetings page itself. :: ## Meeting Host The person who creates a meeting becomes the **host**. The host has additional privileges: - Approve or deny guest access requests - Mute other participants' microphones remotely - All files and recordings produced in the meeting count toward the host's storage quota ## Meeting Capacity The maximum number of participants depends on your plan: | Plan | Max Participants | | -------------- | ---------------- | | 🆓 Basic | 50 | | ⭐ Pro | 100 | | 🏢 Business | 300 | | 🏛️ Enterprise | 1,500 | # Joining a Meeting ## Join via Link The most common way to join a meeting: 1. Click the meeting link shared with you (e.g., `https://intermind.com/abc-def-hij`) 2. If you're signed in, you go directly to the pre-join setup 3. Configure your camera and microphone 4. Click **Join** ## Join via Meeting ID If you have a meeting ID instead of a link: 1. Go to the **Meetings** page 2. Click **Join Meeting** 3. Enter the meeting ID or paste the full URL 4. Click **Join** 5. Configure your camera and microphone in the pre-join setup 6. Click **Join** to enter the meeting ## Join as a Guest If you don't have an InterMIND account, you can join ad-hoc meetings as a guest: 1. Click the meeting link — you're sent to the **Join Meeting** screen 2. Enter your **display name** 3. Optionally pick a translation language (defaults to your interface language) 4. Click **Join** 5. The **Pre-Join Setup** screen lets you preview camera and mic 6. Click **Join** again — a host-approval request is sent and you enter the **waiting room** ("Waiting for approval") 7. Once the host clicks **Admit**, you enter the meeting If the host clicks **Deny**, or does not respond within **5 minutes**, the request is declined and you can leave. ::callout{type="info"} Guest access is only available for ad-hoc meetings. Channels require a team account. :: ## Join Channel Meetings For team channels: 1. Navigate to the channel from the **Channels** page 2. Click to join the channel meeting 3. You must be a member of the team that owns the channel — if not, the meeting URL returns a 404 4. No host approval needed — team members have direct access ## Pre-Join Setup Before entering any meeting, you see the pre-join setup screen where you can: - **Preview your camera** — See how you look - **Select microphone** — Choose from available input devices - **Select camera** — Choose from available cameras - **Select speakers** — Choose audio output device - **Start muted** — Toggle microphone and camera off before joining You can skip this screen by disabling it in [Settings → Meeting Defaults](https://intermind.com/../settings/meeting-defaults). ## Troubleshooting Joining | Issue | Solution | | ---------------------- | ------------------------------------------------------------------------------------------- | | "Meeting not found" | Check that the meeting ID or link is correct | | Camera/mic not working | Make sure you've granted browser permissions for camera and microphone | | Stuck in waiting room | The host hasn't approved your request yet — wait or contact them | | "Access denied" | The host denied your join request, or you're not a member of the team that owns the channel | # Meeting Controls During a meeting, a control bar at the bottom of the screen gives you access to all meeting functions. ## Audio & Video ### Microphone - Click the **microphone icon** to mute/unmute - Click the **arrow next to the microphone** to select a different microphone device - When muted, a red indicator shows on your video tile - An audio level indicator shows when your microphone is picking up sound ### Camera - Click the **camera icon** to turn on/off - Click the **arrow next to the camera** to select a different camera - When off, your video tile shows your avatar instead ### Speaker - Select your preferred audio output device from the settings - This controls where you hear other participants and translations ## Translation Controls ### Real-Time Translation - Open the **More** menu in the control bar and click **Enable translation** (Alt+T) — the menu uses the language icon - When enabled, you hear other participants' speech translated into your chosen language - The original audio is replaced with the translated version - Only available when the meeting has at least two languages in play (the menu item is hidden in single-language rooms) ### Language Selector - Open the cockpit **Settings** panel (gear icon in the More menu) and pick your **Translation Language** - Choose from 24 supported languages (see [Choosing Languages](https://intermind.com/../translation/languages)) - Your choice is saved on your account and persists across sessions - The picker is disabled when your translation usage limit is reached ### Subtitles - Open the **More** menu and click **Show subtitles** to toggle live subtitles - Subtitles appear at the bottom of the meeting screen - They show the translated text in real-time - Layout adapts automatically for right-to-left (RTL) scripts ### Transcription - Transcription is **automatic** — there is no button to start or stop it - Speech is captured as text throughout the meeting, attributed to each speaker - After the meeting, the transcript is reachable from the recap in the meeting chat — see [Meeting Recap](https://intermind.com/recap) - For more details, see [Transcription](https://intermind.com/../translation/transcription) ## Screen & Video Controls ### Self-View - Toggle to show or hide your own video feed - Useful to reduce distraction while presenting ### Resolution - Choose the maximum video resolution - Lower resolution saves bandwidth on slow connections - Options are Auto, 720p, 360p, 180p ## Presentation Tools ### Screen Sharing See [Screen Sharing & Recording](https://intermind.com/screen-sharing) for details. ### Laser Pointer - Available during screen sharing presentations - Point at specific areas on the shared screen - Other participants see your pointer in real-time ## Participant Management ### Participants Sidebar - Click to open the participants panel on the right side - See all participants in the meeting - View each participant's audio/video status ### Host Controls If you are the meeting host, you have additional controls: - **Mute participant** — Remotely mute any participant's microphone - **Admit / Deny guests** — Accept or reject guest access requests from the toast notification When you mute a participant, they see a notification: *"Your microphone was disabled by the host."* ## Reactions & Hand Raise - Click the **Reactions** button to send emoji reactions visible to all participants - Click the **Hand Raise** button to signal you want to speak - For more details, see [Reactions & Hand Raise](https://intermind.com/reactions) ## Meeting Chat - Click the **chat icon** to open the in-meeting chat sidebar - Send messages to all participants during the meeting - The chat sidebar is resizable by dragging its edge - Chat messages are persisted and available after the meeting For full chat features, see [Chat](https://intermind.com/../chat). ## Leaving a Meeting - Click the **Leave** button (door icon) to exit the meeting - The meeting continues for other participants - Your chat messages and transcription are preserved ## Idle Away & Auto-Mute If you're inactive (no mouse or keyboard input) for **30 minutes**, InterMIND marks you as **away**: 1. Your microphone is **automatically muted** 2. An **"You're away"** notification appears, explaining that your mic was muted 3. Move your mouse or press any key to clear the away state You are **not** removed from the meeting — staying in a call as a passive listener of a translated conversation is a core use case, so being marked away only costs one click to un-mute. Coming back does not auto-unmute you: you un-mute when you choose to speak. Auto-mute-on-idle is on by default. See [Meeting Defaults](https://intermind.com/../settings/meeting-defaults) for details on how the toggle is persisted per browser. ## Keyboard Shortcuts All in-meeting shortcuts use **Alt** (or **Option ⌥** on Mac) plus a letter: | Shortcut | Action | | --------- | ------------------------------------------- | | `Alt + M` | Toggle microphone | | `Alt + C` | Toggle camera | | `Alt + T` | Toggle translation | | `Alt + H` | Toggle chat sidebar | | `Alt + P` | Toggle participants sidebar | | `Alt + L` | Reset layout to defaults | | `Alt + R` | Resize the open chat / participants sidebar | | `Esc` | Close cockpit / chat / participants panel | The same list is available in-call via the **Help** modal in the cockpit menu. # Screen Sharing & Recording ## Screen Sharing Share your screen to present documents, slides, or applications during a meeting. ### How to Share Your Screen 1. Click the **Screen** button in the meeting control bar 2. Your browser will show a screen sharing dialog 3. Choose what to share: - **Entire Screen** — Share everything visible on your screen - **Application Window** — Share a specific application - **Browser Tab** — Share a single browser tab (audio included) 4. Click **Share** ### During Screen Sharing - Other participants see your shared screen as the main display - Your camera video moves to a smaller tile - You can use the **laser pointer** to highlight areas on your shared screen - Click **Stop Sharing** to end the screen share ### Tips for Screen Sharing - Close notifications and private content before sharing your entire screen - Sharing a browser tab includes tab audio, which is useful for sharing videos - The shared content quality depends on the video resolution setting ## Recording Record meetings for later review or to share with people who couldn't attend. ### How to Record 1. Click the **Record** button in the meeting controls 2. Your browser may ask for permission to capture your screen 3. Select your screen or tab to record 4. The recording captures: - **Your microphone audio** - **Tab audio** (other participants' voices and translations) - **Screen content** 5. A recording indicator shows during the active recording ### Stopping the Recording - Click the **Stop Recording** button to end - The recording file is saved to your team's storage - You can also stop by clicking the browser's "Stop sharing" button ### Important Notes About Recording - **Storage check** — Recording won't start if your team's storage is at capacity. See [Team Storage](https://intermind.com/../users/storage) for limits - **Browser protection** — If you try to close the browser tab during recording, a warning dialog appears to prevent accidental data loss - **Recording status** — Other participants can see that you are recording - **Host ownership** — Recording files count toward the meeting host's storage quota, regardless of who started the recording ### Recording Storage Recordings are stored in your team's shared cloud storage. Storage limits depend on your plan: | Plan | Storage Limit | | ---------- | ------------- | | Basic | 1 GB | | Pro | 30 GB | | Business | 1 TB | | Enterprise | 5 TB | To manage recordings and storage, see [Team Storage](https://intermind.com/../users/storage). # Reactions & Hand Raise Express yourself during meetings without interrupting the speaker. ## Emoji Reactions Send quick emoji reactions that appear floating on screen for all participants. ### How to Send a Reaction 1. Click the **Reactions** button in the meeting control bar 2. Choose an emoji from the reaction picker 3. The emoji appears as a floating animation on everyone's screen 4. Reactions automatically disappear after a few seconds ### When to Use Reactions - 👍 **Thumbs up** — Agree with something or acknowledge - 👏 **Clapping** — Applaud a presentation or idea - 😂 **Laugh** — React to something funny - ❤️ **Heart** — Show appreciation - 🎉 **Celebration** — Celebrate an achievement Reactions are a great way to provide feedback during presentations without unmuting your microphone. ## Hand Raise Raise your hand to signal that you want to speak, without interrupting the current speaker. ### How to Raise Your Hand 1. Click the **Hand Raise** button in the meeting control bar (or in the reactions menu) 2. A hand icon appears on your video tile, visible to all participants 3. The host and other participants can see who has their hand raised 4. Click the button again to **lower your hand** ### Hand Raise Best Practices - Raise your hand during presentations to queue up for questions - The host can see a list of raised hands and call on participants in order - Lower your hand after you've spoken - Hand raise status persists until you manually lower it or leave the meeting ## Visibility - Both reactions and hand raises are **visible to all participants** in the meeting - Reactions appear as floating animations across the meeting screen - Raised hands show as an indicator on the participant's video tile and in the participants sidebar # Guest Access Guest access allows people without an InterMIND account to join ad-hoc meetings. This is perfect for external participants like clients, interviewees, or collaborators. ## How Guest Access Works ### For Guests 1. **Click the meeting link** shared by the host 2. You're sent to the **Join Meeting** screen 3. Enter your **display name** and click **Join** 4. The **Pre-Join Setup** screen opens — preview your camera and mic 5. Click **Join** again. An anonymous account is created and the host-approval request is sent 6. You enter the **Waiting Room** ("Waiting for approval") 7. Once approved, you enter the meeting with full access to: - Video and audio - Real-time translation - Subtitles - In-meeting chat - Reactions and hand raise ### For the Host When a guest requests to join your meeting: 1. You receive a toast notification: **"{name} wants to join"** 2. You can: - **Admit** — The guest enters the meeting - **Deny** — The guest is informed their request was declined Both the guest and host see updates instantly. ## Guest Limitations Guests have some limitations compared to registered users: | Feature | Registered User | Guest | | ----------------------- | --------------- | -------------------------- | | Join ad-hoc meetings | ✅ | ✅ (host approval required) | | Join team channels | ✅ | ❌ | | Use translation | ✅ | ✅ | | Send chat messages | ✅ | ✅ (in-meeting only) | | Access persistent chats | ✅ | ❌ | | Create meetings | ✅ | ❌ | | Manage settings | ✅ | ❌ | | Team management | ✅ | ❌ | ## Important Notes - Guest access is **only available for ad-hoc meetings**, not channels - Each guest gets a **temporary anonymous account** that is automatically cleaned up - Guests don't need to install anything — it all works in the browser - The host must be present in the meeting to approve guest requests - Multiple guests can request to join simultaneously — the host sees all pending requests - **Guest requests expire after 5 minutes** — if the host doesn't respond within 5 minutes, the request is automatically declined # Sharing Documents & Notes In-Call During a meeting you can put a document or a rich-text note on everyone's screen, next to the video tiles. Any participant can share. ## Sharing a Document Documents linked from Google Docs or OneDrive in the meeting chat (see [File Sharing](https://intermind.com/../chat/files)) can be opened for the whole conference: 1. During a live meeting, click the linked document in chat 2. It opens in a shared panel for every participant, labeled with who shared it 3. Each participant scrolls and zooms independently 4. The sharer closes the panel for everyone with the close button ## Sharing a Note Rich-text notes from chat can be shared the same way: 1. Open the message menu on a note in the meeting chat 2. Choose **Share to conference** 3. The note opens next to the video tiles for all participants While shared: - **The sharer** edits the note live — everyone sees the changes and the sharer's cursor in real time - **Everyone else** reads it in their own language and can toggle between **Translated**, **Original**, and **Original + translation** views ## Sharing a Meeting Recap A [meeting recap](https://intermind.com/recap) message in chat can also be shared into the conference with the same **Share to conference** action — useful for reviewing the previous session's summary together. Each viewer reads it in their own language. ## Related - [File Sharing](https://intermind.com/../chat/files) — Upload files and link Google Docs / OneDrive documents in chat - [Document Translation](https://intermind.com/../chat/document-translation) — Read shared PDF and Office files in your language - [Meeting Recap](https://intermind.com/recap) — Summary, transcript, and recording after the meeting # Meeting Recap After a meeting ends, InterMIND produces a **recap** — a single page with everything the session left behind: - **Summary** — an AI-generated summary of the discussion - **Transcript** — the full speaker-by-speaker transcript - **Recording** — the video recording, when the session was recorded The recap is displayed in your language: whatever language the meeting was held in, the summary is translated for you on demand. ## Where to Find It - **Email** — every participant receives the recap link by email once processing completes - **Chat** — in channels, a recap document message is posted to the channel chat; click it to open the recap - **Calendar** — the calendar event created for a meeting includes the recap link, so it works from day one The recap link (`…/r/…`) is permanent — the same link serves every session of the meeting. ## Who Can Open It Recaps are **members-only**: opening the link requires signing in, and only meeting members — the host and people who participated — get access. Forwarding the link to someone who wasn't in the meeting does not grant access. ## Editing the Summary Meeting members can correct the AI summary directly on the recap page: 1. Open the recap and click **Edit** under the summary 2. Fix the text and save Every edit creates a new revision. **Version history** shows the AI original and each edit as a word-level diff, with author and timestamp. The recap message title in chat updates to match the edited summary. ## Rating the Summary Under the summary you'll see **"Was this summary useful?"** with thumbs up/down. A thumbs-down offers a quick reason — wrong numbers, missed key points, wrong focus, or made things up. Ratings help improve summary quality. ## Recap Archive - **Channels** keep the recap of **every past session** — open the recap page and switch between sessions - **Ad-hoc meetings** keep only the **latest session's** recap, mirroring how their chat history is purged after each call ## Deleting a Recap The meeting **host** can delete a session's recap artifacts — summary, transcript, recording, and their translations. Deletion is permanent; the recap is not regenerated. ## Requirements A recap is produced from the meeting's transcription — see [Transcription](https://intermind.com/../translation/transcription) for how transcripts are captured. ## Related - [Transcription](https://intermind.com/../translation/transcription) — Where transcripts come from - [Screen Sharing & Recording](https://intermind.com/screen-sharing) — Recordings included in the recap - [Sharing Documents & Notes In-Call](https://intermind.com/shared-documents) — Review a recap together in the next session # Meetings & Conferences InterMIND provides two types of video meetings for different use cases. ## Meeting Types ### Ad-hoc Meetings Ad-hoc meetings are **one-time meetings** for quick calls with anyone: - Created instantly — no scheduling required - Shareable link — anyone with the link can request to join - **Guest access** — people without accounts can join (host must approve) - Chat history is purged after the call — the meeting transcript and any attached agenda note/files are preserved (the agenda re-appears in chat at the start of the next conference) - Perfect for client calls, interviews, and external meetings ### Channels Channels are **persistent meeting rooms** for your team: - Always available — join anytime from the **Channels** page - **Members only** — requires team membership; non-members get a 404 on the meeting URL - Chat history is preserved between sessions - Perfect for daily standups, team discussions, and ongoing collaboration ## Creating Meetings The Meetings page in the sidebar is one row of controls plus your meetings: | Control | What Happens | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **New meeting** | Creates a meeting and copies the link to your clipboard; the meeting appears at the top of the list. Click **Join** next to it to enter. | | **Search field** | Paste a meeting ID or link to join — or type a contact's name or email and send them an invite (a new meeting is created and the invite goes out by email, plus an in-app notification for registered users). | To schedule a meeting, use the **calendar button** on its row in the list — it opens your configured calendar provider (Google Calendar or Outlook) with the meeting link pre-filled. Dates, invitees, and recurrence are set in the calendar itself; the list picks them up automatically. You can attach a **Note** or a **File** to any of your meetings — expand its row in the list. Attached content is sent to the meeting chat at start. ## In This Section - [Starting a Meeting](https://intermind.com/meetings/start) — Create and launch meetings - [Joining a Meeting](https://intermind.com/meetings/join) — Join via link, ID, or invitation - [Meeting Controls](https://intermind.com/meetings/controls) — Camera, microphone, and all meeting tools - [Screen Sharing & Recording](https://intermind.com/meetings/screen-sharing) — Present your screen and record calls - [Guest Access](https://intermind.com/meetings/guest-access) — How guests join without an account - [Reactions & Hand Raise](https://intermind.com/meetings/reactions) — Emoji reactions and hand raising - [Sharing Documents & Notes In-Call](https://intermind.com/meetings/shared-documents) — Put a document or live note on everyone's screen - [Meeting Recap](https://intermind.com/meetings/recap) — AI summary, transcript, and recording after the call # Messages Learn how to use the InterMIND chat to communicate with your team and meeting participants. ![Channels page](https://intermind.com/img/docs/chat/chats-page.png) ## Sending Messages ### Text Messages 1. Open the **Chat** panel from the meeting sidebar, or pick a channel from the **Channels** page 2. Type your message in the input field at the bottom 3. Press **Enter** or click the **Send** button Messages are delivered instantly to all participants. ### Message Formatting Messages support basic formatting: - Standard text input - Emoji support (use your system emoji picker or type emoji codes) - URLs are automatically converted to clickable links ## Editing Messages You can edit your own messages after sending: 1. Hover over the message you want to edit 2. Click the **Edit** (pencil) icon 3. A **fullscreen editor** opens with a rich-text toolbar (bold, italic, strikethrough, headings, lists, blockquotes, links, code blocks) 4. Modify the text 5. Click **Save** to confirm or **Cancel** to discard **Important notes:** - Only the message author can edit their messages - The "(edited)" indicator appears only when the actual text content changes — formatting-only changes (e.g., adding bold) do not mark the message as edited - Translation is re-triggered when text changes, but preserved on format-only edits ## Deleting Messages To delete a message you sent: 1. Hover over the message 2. Click the **Delete** (trash) icon 3. Confirm the deletion Deleted messages are removed for all participants. **Host delete:** The meeting host can delete any participant's message — the same **Delete** option in the message menu is available on others' messages when you're the host. ## Notes Send rich-text notes in the chat for longer, formatted content: 1. Click the **"+"** button next to the message input 2. Select **Add Note** 3. A fullscreen editor opens with a toolbar: headings (H1–H3), bold, italic, strikethrough, lists, blockquotes, links, code blocks 4. Write your note and click **Send** Notes are stored as HTML files and rendered in a sandboxed preview within the chat. You can edit a note later via the **Edit Note** option in the message menu. ## Reactions React to any message with emoji: 1. Hover over a message 2. Click the **React** (smiley face) icon 3. Select an emoji from the picker 4. Your reaction appears below the message Multiple participants can react to the same message, and reactions are grouped by emoji with a count. ## Pinning Messages Pin important messages so they're easy to find: 1. Hover over the message 2. Click the **Pin** icon 3. The message is highlighted and accessible from the pinned messages section To view pinned messages: - Look for the **Pinned** indicator or filter in the chat panel - Pinned messages remain at the top or in a separate section for easy access To unpin: - Click the **Unpin** icon on a pinned message ## Message History ### Viewing History All messages are stored permanently: - Scroll up in the chat to load older messages - Messages load in batches as you scroll - The full history is preserved even after meetings end ### Searching Messages Use the search functionality to find specific messages: - Type keywords in the search bar - Results show matching messages with context - Click a result to jump to that message in the chat ## Chat Notifications You'll receive notifications for new messages: - **In-app badge** — Unread message count appears on the chat icon - **Browser tab title** — The tab title updates to reflect unread chats ## Tips - Use **Enter** to send messages quickly - **Shift + Enter** creates a new line without sending - **Alt + Arrow Up/Down** navigates through your sent message history - **Ctrl/⌘ + F** opens chat search - **Ctrl/⌘ + /** opens the Ask AI help drawer - Messages are auto-translated if you have a different language set (see [Chat Translation](https://intermind.com/translation)) - React to messages instead of typing short responses like "OK" or "👍" # Chat Translation InterMIND automatically translates chat messages so team members can communicate in their preferred languages without any extra effort. ## How It Works When a message is sent in the chat: 1. The message is stored in its original language 2. The system identifies all participants who use a different language — including those currently offline 3. The message is translated into each required language 4. Online recipients see the translation in real-time; offline participants receive translations when they reconnect 5. The original text is always available if needed Translation happens transparently — you simply type in your language, and everyone reads in theirs. ## Setting Your Chat Language Your chat translation language follows your account language settings: 1. Open the **Settings** page from the sidebar 2. In the **User** card, either keep "Same as interface language" (the default) or toggle on **Use a different language for translation** and pick one 3. All incoming chat messages are translated to that language The same setting also drives in-meeting voice translation. ## Viewing Original Messages The chat panel has a **Show Original with Translation** display toggle. When on, each translated message is shown in a two-column layout with the original text alongside the translation. Notes have their own per-message **Translated / Original / Diff** switch in the preview header. ## Translation Accuracy ### What Works Well - Common business phrases and terminology - Short, clear sentences - Standard greetings and responses ### Best Practices for Clearer Translations - Write in clear, simple sentences - Avoid slang, idioms, or culture-specific references - Use proper punctuation - Break complex ideas into shorter messages - Spell out abbreviations on first use ## Translation and File Messages - Text messages are translated automatically - File names and captions are not translated - Image content is not translated (only accompanying text) ## Chat Translation Limits Chat translation is unlimited on every paid plan. Only the free Basic plan carries a monthly word allowance. Usage is tracked per **meeting host** (the person who created the meeting) using a **rolling 30-day window**. | Plan | Chat Translation | | ---------- | ---------------- | | Basic | 10K words/month | | Pro | Unlimited | | Business | Unlimited | | Enterprise | Unlimited | As older usage expires, your available words increase. When the limit is reached, new messages are still delivered but not translated. You can: - Continue chatting without translation - Upgrade your plan for a higher limit - Wait for older usage to expire ## Supported Languages Chat translation supports the same 24 languages as meeting translation, on every plan. See [Choosing Languages](https://intermind.com/../translation/languages) for the full list. ## Tips - Set your preferred language once in settings — all chats auto-translate - If a translation looks off, check the original message for context - For critical communications, consider sending both the original and a note about key terms - Translation quality is best for commonly spoken language pairs ## Long Messages Messages longer than **1,000 characters** are automatically split by paragraph breaks (`\n\n`) and each chunk is translated separately. A single paragraph that still exceeds 1,000 characters is broken further at sentence, clause, or word boundaries, so long text is always translated in full rather than skipped. # File Sharing Share files, images, and documents directly in InterMIND channels and meeting conversations. Each upload can be up to **500 MB**. ## Uploading Files ### Drag and Drop The easiest way to share a file: 1. Drag a file from your computer 2. Drop it onto the chat area 3. The file uploads automatically and appears in the chat ### Upload Button 1. Click the **Attach** (paperclip) icon in the chat input area 2. Select a file from your computer 3. The file uploads and appears as a message in the chat ### Pasting Images You can paste images directly from your clipboard: 1. Copy an image (screenshot, from a webpage, etc.) 2. Click in the chat input 3. Press **Ctrl+V** (or **Cmd+V** on Mac) 4. The image is uploaded and shared ### Linking Cloud Documents The **Attach** menu can also link documents from cloud storage: 1. Click the **Attach** (paperclip) icon 2. Choose **Google Docs** or **OneDrive** 3. Connect your account (first time only) and pick a document The document appears in chat as a link — recipients open it with their own access rights in the cloud. During a live meeting, a linked document can be opened for the whole conference — see [Sharing Documents & Notes In-Call](https://intermind.com/../meetings/shared-documents). ## Supported File Types ### Images Supported image formats: - JPEG / JPG - PNG - GIF - WebP - SVG Images display as inline previews in the chat. Click an image to view it full-size. ### Documents You can share various document types: - PDF files - Office documents (Word, Excel, PowerPoint) - Text files - Other common file formats Documents appear as download links with file name and size. Documents shared by others open translated into your language — see [Document Translation](https://intermind.com/document-translation). ## Storage Limits File uploads count toward your team's storage quota: | Plan | Storage Limit | | ---------- | ------------- | | Basic | 1 GB | | Pro | 30 GB | | Business | 1 TB | | Enterprise | 5 TB | ### Checking Storage Usage View your team's storage usage: 1. Open the **Users** page from the sidebar (owners) or the **Billing** page 2. The **Storage Usage** banner / card shows used vs. available space 3. A progress bar indicates how close you are to the limit ### When Storage is Full If your team reaches the storage limit: - You cannot upload new files - Existing files remain accessible - Upgrade your plan or delete old files to free up space ## File Management ### Downloading Files - Click on any shared file to download it - Images can be saved by right-clicking and choosing "Save Image" ### File Visibility - Files shared in a meeting chat are visible to all meeting participants - Files shared in a channel are visible to all channel members - Files in channels persist in chat history indefinitely; files in ad-hoc meeting chats are deleted when the meeting chat is purged after the call ## Recording Files Screen recordings made during meetings are also stored as files: - Recordings are saved to the team's storage - They appear in the chat as well as in the team's file browser - See [Screen Sharing & Recording](https://intermind.com/../meetings/screen-sharing) for details ## Tips - Compress large images before sharing to save storage space - Use the team's file browser (see [Team Storage](https://intermind.com/../users/storage)) to manage all shared files - Storage usage is shared across the entire team — coordinate with team members if space is limited ## Related - [Document Translation](https://intermind.com/document-translation) — Read shared PDF and Office files in your language - [Team Storage](https://intermind.com/../users/storage) — Detailed storage management and file browsing - [Screen Sharing & Recording](https://intermind.com/../meetings/screen-sharing) — Recordings also count toward storage - [Billing & Plans](https://intermind.com/../billing) — Upgrade your plan for more storage # Document Translation Documents shared in chat can be read in your own language. When you open a file someone else uploaded, InterMIND translates it into your language first — there is no button to press. ## How It Works 1. A participant shares a document in chat 2. You click it to open 3. If it's in another language, the tile shows **Translating document…** 4. The translated copy opens, marked **Translated** The sender always sees the original, and if the document is already in your language nothing is translated. Once translated, the copy is cached — reopening it is instant, and so is opening it for teammates with the same language. ## Supported Formats PDF, DOCX, DOC, PPTX, and XLSX, translated into 30 languages via DeepL. Other file types open as-is. ## Plans & Limits Document translation is a **paid feature**, metered per distinct document over a rolling 30-day window: | Plan | Documents / 30 days | | -------------- | ------------------- | | 🆓 Basic | Not available | | ⭐ Pro | 10 | | 🏢 Business | 30 | | 🏛️ Enterprise | 100 | - Usage counts against the **meeting host's** plan, whoever opens the document - Re-translating the same file into additional languages is free — the meter counts distinct files - Track usage on the **Billing** page under the **Document Translation** meter — see [Usage & Invoices](https://intermind.com/../billing/usage) When the quota is reached — or on the Basic plan — the original file still opens, just untranslated. ## Related - [File Sharing](https://intermind.com/files) — Upload and link documents in chat - [Sharing Documents & Notes In-Call](https://intermind.com/../meetings/shared-documents) — Open a document for the whole conference - [Usage & Invoices](https://intermind.com/../billing/usage) — All usage meters in one place # Chat InterMIND includes a built-in chat system designed for multilingual teams. Messages are synchronized in real-time and can be automatically translated into each participant's preferred language. ## How Chat Works - Messages appear instantly for all participants - All messages are stored on the server — if you join a meeting late, you see the full chat history - Chat is available both during meetings (sidebar panel) and as standalone **Channels** on the **Channels** page ## Key Features | Feature | Description | | ------------------- | -------------------------------------------------- | | Real-time messaging | Messages appear instantly for all participants | | Auto-translation | Messages automatically translated to your language | | File sharing | Share images and documents in chat | | Reactions | React to messages with emoji | | Pin messages | Pin important messages for easy reference | | Rich-text notes | Add formatted notes with headings, lists, and code | | Edit & delete | Modify or remove your messages | | History | Full message history preserved | ## Chat Types ### Meeting Chat (ad-hoc) Every ad-hoc meeting has a chat panel in the sidebar: - Messages sent during the meeting are linked to that meeting - For ad-hoc meetings, the chat history is deleted after the call (except any agenda note or file the host attached at meeting creation, which is preserved) ### Channels Channels are persistent team chats that live on the **Channels** page, independent of any single meeting: - Use them for ongoing team communication - Start a meeting directly from a channel - All members see the full message history, which is kept indefinitely ## Getting Started - [Sending Messages](https://intermind.com/chat/messages) — How to send, edit, and manage messages - [Sharing Files](https://intermind.com/chat/files) — Upload and share files in chat - [Chat Translation](https://intermind.com/chat/translation) — Automatic message translation - [Document Translation](https://intermind.com/chat/document-translation) — Read shared files in your language ## Related - [Meeting Controls](https://intermind.com/meetings/controls) — Access chat during meetings - [Meeting Defaults](https://intermind.com/settings/meeting-defaults) — Configure chat display settings # Profile Settings Your profile settings control how you appear to other participants and which language is used for translation. ![Settings page with the User section at the top](https://intermind.com/img/docs/settings/settings-page.png) ## Accessing Profile Settings 1. Click **Settings** in the sidebar 2. The **User** section is at the top of the page ## Avatar ### Uploading an Avatar 1. Click the **Upload** button in the User section 2. Select an image from your computer 3. The image is automatically cropped to a square and resized to 256×256 pixels **Supported formats:** JPEG, PNG, GIF, WebP :br**Maximum file size:** 10 MB Your avatar is automatically optimized (compressed to JPEG quality 0.85) for fast loading. ### Default Avatar If you don't upload a photo, a default avatar is generated from your initials with a colored background. This appears everywhere your profile picture is shown. ### Removing Your Avatar Click **Remove** next to your avatar to revert to the auto-generated initials avatar. This button only appears if you've uploaded a custom photo. ## Display Name Your display name is shown to other participants in meetings, chat, and team views. 1. Type your name in the **Name** field 2. Your name is saved automatically (with a brief delay) 3. If you leave the field empty, your email address is used as a fallback ## Email Your email address is displayed as read-only. It's set during account creation and cannot be changed from the settings page. ## Translation Language Your translation language determines what you hear when translation is active in a meeting and what language incoming chat messages are translated to. By default, your translation language **follows your interface language** — the User card shows the hint **"Same as interface language"**. ### Using a Different Translation Language 1. In the **User** card, toggle **"Use a different language for translation"** on 2. A **Translation Language** dropdown appears 3. Search or scroll through the list and select your language 4. The hint changes to **"Different from interface language"** To go back to the default, toggle the switch off. Selecting the same language as your interface language also clears the override automatically. Your interface language itself is selected from the language switcher in the app header (it determines page UI strings). ### Supported Languages InterMIND supports 24 languages for real-time translation (voice, chat, and shared notes). On-demand file translation in chat — a paid feature, available on Pro and above — covers a wider 30. Full breakdown: [How many languages do we support — six numbers, not one](https://intermind.com/blog/how-many-languages-do-you-support). | Language | Native Name | | ----------------------- | ----------- | | English | English | | Russian | Русский | | Spanish (Latin America) | Español | | French | Français | | German | Deutsch | | Italian | Italiano | | Portuguese (Brazilian) | Português | | Polish | Polski | | Dutch | Nederlands | | Czech | Čeština | | Hungarian | Magyar | | Danish | Dansk | | Finnish | Suomi | | Icelandic | Íslenska | | Norwegian | Norsk | | Romanian | Română | | Swedish | Svenska | | Ukrainian | Українська | | Japanese | 日本語 | | Korean | 한국어 | | Turkish | Türkçe | | Chinese (Simplified) | 中文 | | Arabic | العربية | | Hindi | हिन्दी | ### Default Language If you haven't set a language yet, InterMIND uses your browser's language setting. If your browser language isn't in the supported list, it defaults to English. ## Syncing Across Devices Your profile settings (name, avatar, language) are stored in the cloud and sync across all devices where you're signed in. Changes take effect immediately on all sessions. # Appearance InterMIND lets you toggle between light and dark mode from the sidebar, and customize the primary and secondary accent colors used in each mode from the Settings page. ## Color Mode (Light / Dark) Color mode is toggled from the sidebar, not the Settings page: 1. Look for the **Toggle theme** button in the app sidebar (or the sun/moon icon in the header on landing pages) 2. Click to cycle between the available modes The mode follows your system preference by default. Switching dark/light during a meeting is handled separately — the meeting room is always rendered in dark, and sidebars (chat, participants) follow your preference inside it. ## Theme Colors You can customize the primary and secondary accent colors for both light and dark themes independently. This is done from the **Application preferences** card on the **Settings** page. ### What Can Be Customized | Setting | Affects | | ----------------------- | --------------------------------------------- | | Light theme — Primary | Buttons, links, active elements in light mode | | Light theme — Secondary | Accents, badges, secondary UI in light mode | | Dark theme — Primary | Buttons, links, active elements in dark mode | | Dark theme — Secondary | Accents, badges, secondary UI in dark mode | ### Changing Theme Colors 1. Go to **Settings** → **Application preferences** 2. Each theme (Light / Dark) shows a **Primary** and **Secondary** swatch 3. Click a swatch to open the color picker 4. Pick a color — the change applies immediately ### Resetting Colors Click the **Reset** (circular arrow) button at the top of the Application preferences card to restore all four color settings to their defaults. ### Syncing Colors Across Devices Theme colors are stored on your account and follow you across devices when you sign in. # Meeting Defaults InterMIND remembers most of your in-meeting preferences between sessions. Defaults are **stored locally in your browser**, so they're per-browser/per-device. The only meeting-related preference synced to your account is your translation **language** (set in the User card on the Settings page). ## Pre-Join Setup There is no toggle for the Pre-Join Setup screen. InterMIND shows it automatically the **first time** you join on a device, and again whenever your camera / microphone / speaker configuration has changed since your last join. Once you've joined with a given device setup, later joins skip straight into the meeting until that setup changes. ## Device Defaults The selected camera, microphone, and speaker are remembered the next time you join a meeting: 1. Open the **Pre-Join Setup** screen before entering, or open the cockpit's device dropdowns during a meeting 2. Pick a camera / microphone / speaker 3. Your last selection is reused on the next join Camera and microphone **on/off state** at join time follows whatever you last set in the Pre-Join Setup or cockpit. ## Translation & Subtitles The on/off state of translation and subtitles is remembered between meetings: - Open the **More** menu in the cockpit and click **Enable translation** (Alt+T) to toggle real-time translation - Open the **More** menu and click **Show subtitles** to show/hide live subtitles Transcription is **automatic** — there is no toggle. Speech is captured throughout the meeting and the transcript is reachable from the recap card afterward (see [Transcription](https://intermind.com/../translation/transcription)). Your translation **language** is set in the User card on the Settings page (see [Profile Settings](https://intermind.com/profile)). ## Idle Away & Auto-Mute If you're idle (no mouse or keyboard input) for **30 minutes**, InterMIND marks you as away and **mutes your microphone**, showing a "You're away" notification. You are not removed from the meeting — moving your mouse or pressing a key clears the away state. See [Meeting Controls](https://intermind.com/../meetings/controls) for details. This is on by default and its on/off state is part of the device settings persisted locally in your browser. ## Sidebar Widths The widths of in-meeting sidebars (chat panel, participants panel) are remembered between sessions. Resize a panel by dragging its edge, or cycle through preset widths with **Alt + R** while the panel is open. ## Resetting Defaults There's no global "reset" button. To reset most defaults, clear your browser's local storage for the InterMIND site — all `inter-mind:`-prefixed entries belong to these device defaults. # Keyboard Shortcuts InterMIND uses the physical **Alt** key (called **Option** on macOS) for most in-meeting shortcuts, so they work the same on every keyboard layout. ## Global These work anywhere in the app. | Shortcut | Action | | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | **Cmd/Ctrl + /** | Open the Ask AI help drawer | | **Esc** | Close the topmost panel — closes drawers and sidebars in order: Cockpit settings → chat sidebar → participants sidebar. Respects open modals. | ## In a Meeting Active only inside a meeting room. Disabled while a modal dialog is open. | Shortcut | Action | | ------------------ | --------------------------------------- | | **Alt/Option + M** | Toggle your microphone | | **Alt/Option + C** | Toggle your camera | | **Alt/Option + T** | Toggle real-time translation | | **Alt/Option + H** | Toggle the chat sidebar | | **Alt/Option + P** | Toggle the participants sidebar | | **Alt/Option + L** | Reset the stage layout to defaults | | **Alt/Option + R** | Cycle chat / participants sidebar width | ## In Chat When the chat surface is focused (in-meeting sidebar or a standalone channel on the Channels page). | Shortcut | Action | | ---------------------------------- | ---------------------------------- | | **Cmd/Ctrl + F** | Open chat search | | **Enter** (in message box) | Send the message | | **Shift + Enter** (in message box) | Insert a newline | | **Alt/Option + ↑** | Recall your previous message | | **Alt/Option + ↓** | Recall the next message in history | ### Inside Chat Search | Shortcut | Action | | ----------------------- | -------------------------- | | **Enter** | Jump to the next match | | **Shift + Enter** | Jump to the previous match | | **↑** / **↓** | Navigate search history | | **Esc** | Close search | ## Notes - **Alt vs Option** — same physical key. On macOS it's labelled **Option**; on Windows / Linux it's **Alt**. Shortcuts use the key's physical position, so they work regardless of language layout. - **Modals take priority** — when a dialog is open (sign-in, settings, file picker), meeting shortcuts are suspended so you can type freely. - **No conflict with the browser** — InterMIND avoids Cmd/Ctrl combos that overlap with browser shortcuts (Ctrl+T, Ctrl+W, etc.), except for Cmd/Ctrl+F in chat (intentional) and Cmd/Ctrl+/ for help. ## Related - [Meeting Controls](https://intermind.com/../meetings/controls) — Visual equivalents in the control bar - [Messages](https://intermind.com/../chat/messages) — Chat message basics # Calendar The **Calendar Integration** card on the **Settings** page keeps meetings created in InterMIND synchronized with your calendar. ## Connecting a Calendar 1. Open the **Settings** page from the sidebar 2. In the **Calendar Integration** card, click **Connect** next to **Google Calendar** or **Microsoft Outlook Calendar** 3. Sign in with your account in the OAuth popup 4. The card shows the connected account's email Click **Disconnect** next to a connected calendar to remove the connection. ## Default Calendar With both calendars connected, click **Make default** to pick which one is used for scheduled meetings. The current default carries a **Default for scheduled meetings** badge. ## How It Works - Every meeting you host has a calendar button on its row in the Meetings list. Clicking it opens your default calendar with the event pre-filled: title, join link, and the meeting recap link. - Scheduled meetings pick up their dates and changes from your calendar. If the calendar is disconnected, meeting rows keep showing the date from the last sync until you reconnect. # Backup & Restore The **Backup & Restore** card on the **Settings** page lets you save all your chats as a ZIP and bring them back later. The card appears when your plan includes data export. ## What's Included | Data Type | Description | | -------------- | ------------------------------------ | | Chat messages | All text messages from your channels | | Transcriptions | Meeting transcription records | | Files | Uploaded files and documents | | Images | Shared images | | Recordings | Conference screen recordings | ## Creating a Backup 1. Open the **Settings** page from the sidebar 2. In the **Backup & Restore** card, click **Create Backup** 3. A real-time progress bar tracks each step: scanning meetings → downloading files → compressing 4. When complete, click **Download** to save the ZIP file ## Restoring from a Backup Restore selected chats from a previously created backup. | Data Type | Description | | ------------- | ------------------------------------------------------- | | Chat messages | All message types (chat, transcription, image, file) | | Files | Uploaded back to your team's Storage | | Recordings | Conference recordings uploaded to Storage | | Metadata | Pins, reactions, translations, file names, sizes, types | 1. Open the **Settings** page from the sidebar 2. In the **Backup & Restore** card, click **Restore from Backup** 3. Select a ZIP archive 4. Pick which chats to restore (others stay untouched) 5. Confirm — a full-screen overlay shows restore progress 6. After restore, a summary shows the results along with any errors or warnings ::callout{type="warning"} Selected chats are replaced with the backup snapshot — any newer messages, files, or recordings added after the backup was made will be lost. Chats you don't select stay as they are. :: # Settings The Settings page lets you personalize your InterMIND experience. Organization-level controls (SSO, domains, directory sync) live on a separate **Integrations** page. ## Settings Page The **Settings** page in the sidebar contains these cards, stacked top to bottom: | Card | What You Can Configure | | --------------------------- | ----------------------------------------------------------------------- | | **User** | Avatar, display name, email (read-only), translation-language override | | **Application preferences** | Theme color customization (primary and secondary, light and dark) | | **Calendar** | Connect Google or Microsoft calendar and pick your default | | **Data Export** | Export and import your data (shown when your plan includes data export) | | **Delete Account** | Permanently delete your account and associated data | Color **mode** (light / dark / system) is toggled via the theme button in the app sidebar, not from the Settings page. ::callout{type="info"} Directory Integration, Domain Management, and SSO are configured on the separate [Integrations](https://intermind.com/docs/integrations) page. :: ## Quick Links - [Profile Settings](https://intermind.com/settings/profile) — Avatar, display name, and translation language - [Appearance](https://intermind.com/settings/appearance) — Theme color customization - [Meeting Defaults](https://intermind.com/settings/meeting-defaults) — Device and feature defaults - [Calendar](https://intermind.com/settings/calendar) — Connect Google or Outlook calendar and pick your default - [Backup & Restore](https://intermind.com/settings/data-export) — Export your chats as a ZIP and restore them later # Inviting Members InterMIND offers multiple ways to add people to your team — from individual email invitations to bulk imports and directory syncing. ![Users page](https://intermind.com/img/docs/users/users-page.png) ## Invitation Methods Access the invitation options from the **Users** page by clicking the **Invite Users** dropdown button. ### 1. Send Email Invitations Send invitation emails to specific addresses: 1. Click **Invite Users** → **Send Email Invites** 2. Enter email addresses in the text area (one per line, or separated by commas, semicolons, or spaces) 3. Optionally add tags using bracket notation: `user@example.com [developer, frontend]` 4. Click **Send Invitations** Each recipient receives an email with a link to join your team. **Supported email formats:** - `user@example.com` — plain email - `"John Doe" ` — name and email - `user@example.com [tag1, tag2]` — email with tags ### 2. Invitation Link Generate a shareable link that anyone can use to join: 1. Click **Invite Users** → **Invitation Link** 2. Copy the generated link 3. Share it via any channel (Slack, email, etc.) The link can be regenerated if you want to invalidate the old one. ### 3. Bulk Import Import multiple users from a file: 1. Click **Invite Users** → **Bulk Import** 2. Upload a file (.txt, .csv, or .json) 3. Review the parsed list of emails 4. Send invitations **File format examples:** **TXT (one email per line):** ```text alice@company.com [design] bob@company.com [engineering] carol@company.com ``` **CSV:** ```csv email,tags alice@company.com,"design,ui" bob@company.com,"engineering" ``` **JSON:** ```json [ { "email": "alice@company.com", "tags": ["design"] }, { "email": "bob@company.com", "tags": ["engineering"] } ] ``` For Excel files (.xls/.xlsx), save them as CSV first. ### 4. From Directory (Google Workspace / Microsoft 365) Import members directly from your corporate directory: 1. Click **Invite Users** → **From Contacts** 2. If not connected, you'll be redirected to the **Integrations** page to connect a directory 3. Once connected, search for contacts by name, email, or department 4. Select users to add and click **Add Members** Contacts that are already team members show as "Member". Pending invitations show as "Invited". ### 5. Domain Auto-Join Users from your verified domain can join automatically: 1. Verify your domain in **Integrations** → **Domain Management** 2. When someone creates an account with an email matching your domain (e.g., `@yourcompany.com`), they are automatically added as a **Member** 3. The team owner receives a notification Email aliases (e.g., `user+tag@domain.com`) are excluded from auto-join. ## Invitation Details ### Expiration - Email invitations expire after **7 days** - The expiration is shown in the Pending Invitations table - Expired invitations can be resent ### Accepting an Invitation When someone receives an invitation: 1. They click the link in the email or see a banner in the app 2. If not signed in, they're prompted to create an account or sign in 3. The invitation banner shows with an **Accept** button 4. Clicking **Accept** adds them to the team as a **Member** ### Pending Invitations Track all outstanding invitations on the Team Users page: | Column | Description | | ----------- | ------------------------------- | | **Email** | Invited email address | | **Status** | Pending, Accepted, or Expired | | **Expires** | When the invitation expires | | **Tags** | Tags assigned during invitation | ## Tips - Use tags when inviting to organize members from the start (e.g., `[engineering]`, `[marketing]`) - The invitation link is great for large groups — share it in a team channel - Connect your directory for the easiest experience with large organizations - Verify your domain to enable automatic onboarding for new employees # Managing Members The Users page gives you a comprehensive view of all team members with tools for organizing, searching, and managing roles. ## Member Table The team member table shows: | Column | Description | | ----------- | --------------------------------------------------- | | **Select** | Checkbox for bulk actions (not shown for the owner) | | **User** | Avatar, display name, and email address | | **Tags** | Editable tags for organizing members | | **Role** | Color-coded role badge (sortable) | | **Storage** | Amount of storage used by this member | | **Actions** | Dropdown menu with management options | ### User Details Each user row displays: - **Avatar** — Profile picture or auto-generated initials - **Display name** — Their chosen name - **Email** — With the domain highlighted in color if it matches a verified domain - **Tooltip** — Shows the team name and when they joined on hover - **Expand button** — Click to see the user's files (if any) ## Tags Tags help you organize team members by department, project, or any custom category. ### Adding Tags 1. Click in the **Tags** column for any member 2. Type a tag name 3. Press **Enter** to add it 4. Add as many tags as needed Tags are saved automatically. ### Assigning Tags During Invitations You can assign tags when inviting users: ```text alice@company.com [design, frontend] bob@company.com [engineering, backend] ``` ### Searching by Tags Use the search bar at the top of the member table: 1. Type in the **Search by tags** field 2. Members are filtered in real-time 3. The search uses AND logic — all words must match at least one tag 4. Partial matches work (e.g., "eng" matches "engineering") ## Sorting Click the **Role** column header to sort members by role: | Sort Order | Role | | ---------- | ------ | | 1st | Owner | | 2nd | Admin | | 3rd | Member | | 4th | Guest | The default sort is by role in ascending order. Click again to toggle sort direction. You can also sort by **Storage** to see who uses the most space. ## Individual Actions Click the **Actions** dropdown (⋮) on any member row: | Action | Description | | ------------------ | ----------------------------------------------------------------------------- | | **Make Admin** | Promote to admin role (shown when current role is member) | | **Make Member** | Set role to standard member (shown when current role is admin) | | **Grant License** | Give this member a paid seat — required for full feature access on paid teams | | **Revoke License** | Revoke the member's paid seat (they fall back to Basic features) | | **Delete** | Remove from the team | The owner cannot be deleted or have their role changed, and the owner row also has no action menu for self-actions. ## Bulk Actions Select multiple members using the checkboxes, then use the bulk action buttons: 1. Check the boxes next to the members you want to manage 2. A toolbar appears with available actions: - **Make Admin (N)** / **Make Member (N)** — change role for all selected - **Grant License (N)** / **Revoke License (N)** — toggle paid seats for all selected (if not enough free seats, you'll see a "Grant License — X of Y" partial label) - **Delete (N)** — remove all selected from the team The owner row doesn't have a checkbox and is excluded from bulk actions. ### What Happens When a Member Is Removed When you delete a member from the team: - They're removed from the team's member list - They're removed from the team chat - The team ID is removed from their user profile - They can be re-invited later ## Verified Domains Email addresses matching your verified domains get special visual treatment: - The domain part is highlighted in the primary UI color - This makes it easy to identify which members are from your organization vs. external guests ## Tips - Use tags consistently across your team (e.g., always use "engineering" not "eng" or "developers") - Promote trusted team members to Admin for help with team management - Sort by Storage to identify which users consume the most space - Use the search function with tags to quickly find specific groups of people # Team Storage InterMIND provides a shared storage pool for your team. All uploaded files, images, and screen recordings count toward your team's storage limit. ## Storage Limits by Plan | Plan | Storage Limit | | -------------- | ------------- | | 🆓 Basic | 1 GB | | ⭐ Pro | 30 GB | | 🏢 Business | 1 TB | | 🏛️ Enterprise | 5 TB | Storage is a **shared team pool** — the total is shared by all team members, regardless of the number of users. ## Storage Warnings The Users page shows warning banners based on storage usage: | Usage | Banner | Impact | | ---------- | -------------------------------------------------------- | --------------------------------------- | | **≥ 80%** | Info — "Storage Usage: X / Y (Z%)" | No restrictions, awareness alert | | **≥ 95%** | Warning — "Storage Almost Full" + "Upgrade Plan" button | Consider upgrading soon | | **≥ 100%** | Error — "Storage Limit Exceeded" + "Upgrade Plan" button | **File uploads and recordings blocked** | You also receive email notifications at the 80%, 95%, and 100% thresholds. ## How Storage Is Counted ### What Counts - **Uploaded files** — Documents, spreadsheets, etc. shared in chat - **Shared images** — Photos and screenshots shared in chat - **Screen recordings** — Recordings made during meetings - **Meeting files** — Any files associated with meetings ### Who Owns What - **Uploaded files** — Counted against the person who uploaded them - **Screen recordings** — Counted against the meeting host (creator) ### Per-User Storage View On the Users page, the **Storage** column shows how much space each member uses. This helps identify who's consuming the most storage. ## Browsing Files ### Expanding User Files To see a team member's files: 1. Go to the **Team Users** page 2. Find the member in the table 3. Click the **expand** button (▸) next to their name 4. Their files appear in a list below their row Each file shows: - **Icon** — Based on file type (video, image, document) - **Name** — File name (or ID for recordings and images) - **Date** — When the file was uploaded or created - **Size** — File size in MB ### Previewing Files Click on any file to open a preview modal: | File Type | Preview | | --------------------------------------------- | --------------------------------------- | | **Images** (JPEG, PNG, GIF, WebP) | Full-size image viewer | | **Videos / Recordings** | Video player with controls and autoplay | | **Audio** | Audio player with musical note icon | | **PDF** | Embedded PDF viewer | | **Office docs** (.doc, .xlsx, .ppt) | Google Docs Viewer | | **Text files** (.txt, .md, .json, .csv, etc.) | Embedded text viewer | | **Other formats** | "Open in new tab" link | ## Managing Storage ### When Storage Is Full If your team exceeds the storage limit: 1. **Uploads are blocked** — No new files can be uploaded 2. **Recording is blocked** — Screen recording during meetings is disabled 3. **Existing files remain accessible** — You can still view and download files ### Freeing Up Space To reclaim storage: - Delete old or unnecessary files - Download important recordings locally and remove them from InterMIND - Ask team members to review their uploaded files ### Upgrading Storage To get more storage: 1. Click the **Upgrade Plan** button in the storage warning banner 2. Or go to [Pricing](https://intermind.com/pricing) and select a higher plan 3. Storage increases immediately after upgrading ## Real-Time Updates Storage information updates in real-time: - New file uploads and deletions are reflected immediately - Storage usage bars and numbers refresh automatically - Warning banners appear and disappear based on current usage ## Tips - Monitor the storage banner on the Users page regularly - Sort the member table by **Storage** column to find the biggest consumers - Screen recordings tend to be the largest files — download them locally if you need to free space - Encourage team members to avoid uploading unnecessarily large files - Choose a yearly plan for both cost savings and sufficient storage for your needs # Users & Organization This section is for **owners and admins**. Regular members can skip it — your role gives you everything you need by default. Member management lives on the **Users** page (invitations, roles, storage), while your organization name is set on the **Billing** page. ## How It Works - Every user automatically belongs to an **organization** (named "{Your Name}'s Team" by default) - You can invite others to join via email, link, bulk import, or directory sync - Members share a storage pool and can participate in shared channels - Each user belongs to **at most one organization** at a time — accepting another invitation replaces the current one ## Roles | Role | Badge Color | Description | | ---------- | ----------- | ------------------------------------------ | | **Owner** | Primary | The organization creator; full control | | **Admin** | Info | Elevated privileges, assigned by the owner | | **Member** | Neutral | Standard member (default role) | | **Guest** | Warning | Joined via meeting link; limited access | ### Role Assignment - **Owner** — Automatically set for the person who created the organization - **Admin** — Promoted by the owner from the Users page - **Member** — Default role when joining via invitation or domain auto-join - **Guest** — Automatically assigned when someone joins a meeting via link without being a member ## Organization Name You can customize your organization name on the **Billing** page (owners and admins only) — see [Usage & Invoices](https://intermind.com/billing/usage). The name shows up to teammates and on shared invitations. ## Quick Links - [Inviting Members](https://intermind.com/users/inviting) — All ways to add people - [Managing Members](https://intermind.com/users/managing) — Roles, tags, search, and bulk actions - [Storage](https://intermind.com/users/storage) — Shared storage pool and file management # SSO Setup This guide is for the IT admin connecting a company identity provider (IdP) to InterMIND. After setup, members sign in from the regular login page: **Sign in with SSO** → work email → your IdP → back in InterMIND. **Available on:** Business and Enterprise plans **Configured by:** team owner or admin **Protocol:** OpenID Connect (OIDC). SAML 2.0 sign-in is in development — SAML configuration is stored but cannot be used to sign in yet. ## Prerequisites 1. A **verified domain** — verify your email domain via DNS TXT record first (see [Domain Management](https://intermind.com/docs/integrations#domain-management)). SSO sign-in only accepts accounts whose email domain your team has verified; this is the tenant boundary. 2. An IdP that supports **OIDC with discovery** — it must serve `/.well-known/openid-configuration` under the Issuer URL. Okta, Microsoft Entra ID, and Google all do. ## What to register in your IdP Create an **OIDC Web Application** in your IdP with: | Setting | Value | | ----------------------- | ------------------------------------------------------------------------------------------------ | | Redirect URI (callback) | `https://intermind.com/api/auth/sso/callback` — also shown in the SSO card after you select OIDC | | Grant type | Authorization Code (PKCE S256 is used automatically) | | Scopes | `openid email profile` | The ID token your IdP issues must include the user's `email`, and the email's domain must be one of your verified domains — otherwise sign-in is refused. Then fill in the **SSO** card on the Integrations page: | Field | What to paste | | ------------------------- | --------------------------------------------------------------------------- | | Display Name | Any label your members will recognize | | Issuer URL | Your IdP's issuer — the URL that serves `/.well-known/openid-configuration` | | Authorization URL | The `authorization_endpoint` from that discovery document | | Client ID / Client Secret | From the app you registered | The client secret is encrypted at rest and never returned to the browser after saving. ## Okta 1. Admin console → **Applications → Create App Integration** → sign-in method **OIDC**, application type **Web Application** 2. **Sign-in redirect URI:** `https://intermind.com/api/auth/sso/callback` 3. Assign the users or groups who should have access 4. Copy the **Client ID** and **Client Secret** 5. In InterMIND: Issuer URL = your Okta org URL (e.g. `https://acme.okta.com`, or your authorization server's issuer such as `https://acme.okta.com/oauth2/default` if you use one); Authorization URL = the `authorization_endpoint` from `/.well-known/openid-configuration` ## Microsoft Entra ID (Azure AD) 1. Entra admin center → **App registrations → New registration** 2. Platform **Web**, redirect URI `https://intermind.com/api/auth/sso/callback` 3. **Certificates & secrets → New client secret** — copy the secret **Value** immediately 4. Client ID = the **Application (client) ID** on the Overview page 5. Make sure the ID token carries the user's email: **Token configuration → Add optional claim → ID → email** 6. In InterMIND: Issuer URL = `https://login.microsoftonline.com//v2.0`; Authorization URL = `https://login.microsoftonline.com//oauth2/v2.0/authorize` ## Google Workspace No app registration needed. In the SSO card choose the **Google Workspace** provider type and save — members on your verified domains sign in with their Google account and join your team automatically. (Google can also be connected as a generic OIDC provider with issuer `https://accounts.google.com` if you prefer explicit client credentials.) ## Test the connection 1. Open the login page in a private/incognito window 2. Click **Sign in with SSO** and enter a work email on your verified domain 3. You are redirected to your IdP; after authenticating, you land back in InterMIND signed in 4. The sign-in is recorded in the team audit log (exportable from the **Users** page) as `auth.login` with method `sso` ## Troubleshooting | Symptom | Cause | | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | | "SSO is not configured" after entering the email | No enabled SSO config matches that email domain — check the domain is verified and the SSO card is saved | | `SSO login is not available: plan` | The team's plan no longer includes SSO | | `SSO login is not available: domain-not-verified` | The domain is still pending DNS verification | | `SSO login is not available: config-incomplete` | Client ID or Client Secret missing — re-save the SSO card | | `SSO login is not available: type-unsupported` | The stored config is SAML — SAML sign-in is not available yet | | `SSO IdP discovery failed` | Issuer URL is wrong or doesn't serve `/.well-known/openid-configuration` | | "login session expired, start again" | More than 5 minutes passed between starting sign-in and the IdP callback | | Sign-in refused after the IdP redirects back | The IdP returned an email outside your verified domains, or no `email` claim at all (Entra: add the optional email claim) | ## Security properties For security questionnaires: the SSO flow is Authorization Code with **PKCE (S256)**, **state**, and **nonce**; the ID token's **signature is validated against the IdP's JWKS**, along with issuer and audience; the IdP is authoritative only for **domains verified via DNS** — an assertion for any other email never produces a session; the OIDC client secret is **encrypted at rest**; every SSO sign-in lands in the team **audit log**. Plan, domain, and configuration gates are enforced server-side on both the sign-in start and the callback. # Integrations InterMIND integrates with external services to enhance your workflow — enterprise directory sync, domain management, and SSO. ![Integrations page](https://intermind.com/img/docs/integrations/integrations-page.png) The **Integrations** page is reached from the sidebar. Which cards appear depends on your role and plan: | Card | Who sees it | | --------------------- | ------------------------------------------------------------ | | Directory Integration | Org admins (owner / admin role) | | Domain Management | Org admins on any paid plan; full functionality on Business+ | | SSO | Org admins on any paid plan; configuration on Business+ | ::callout{type="info"} Personal integrations — [Calendar](https://intermind.com/docs/settings/calendar) and [Backup & Restore](https://intermind.com/docs/settings/data-export) — live on the **Settings** page. :: []{#domain-management} ## Domain Management Verify your organization's domains to enable SSO and automatic team membership. ### Adding a Domain 1. Open the **Integrations** page from the sidebar 2. In the **Domain Management** card, type your domain (e.g., `yourcompany.com`) in the input field 3. Press **Enter**, comma, or space to add it 4. The domain appears with a "pending" status ### Verifying a Domain Domains are verified via DNS TXT records: 1. Copy the verification code shown for your domain 2. Add a DNS TXT record in your domain registrar (e.g., GoDaddy, Cloudflare) 3. Use the verification code as the TXT record value 4. Click **Verify All** or wait for the 60-second automatic check 5. Once DNS propagates, the domain status changes to "verified" ✓ ### Domain Statuses | Status | Icon | Meaning | | -------- | ---- | ------------------------------------- | | Verified | ✓ | Domain confirmed — SSO can be enabled | | Pending | ⏳ | Waiting for DNS verification | | Invalid | ✗ | DNS check failed — verify TXT record | ## SSO **Available on:** Business and Enterprise plans :br**Requires:** At least one verified domain Members signing in with an email on one of your verified domains join your team automatically (see [Domain Management](https://intermind.com/#domain-management) above). With an identity provider configured in the **SSO** card, they sign in through it: login page → **Sign in with SSO** → work email → your IdP. ### Identity Provider Configuration Follow the step-by-step [SSO Setup](https://intermind.com/docs/integrations/sso-setup) guide for Okta, Microsoft Entra ID, and Google Workspace. | Provider | Sign-in | Configuration | | ------------------------- | ------------------------------------------------ | ------------------------------------------------------- | | **OIDC / OpenID Connect** | Available | Issuer URL, Authorization URL, Client ID, Client Secret | | **Google Workspace** | Available — members use their Google account | Verified domain only — no credentials needed | | **SAML 2.0** | In development — stored config can't sign in yet | IdP Entity ID, SSO URL, X.509 Certificate | ## Directory Integration Sync your organization's user directory to simplify team management. ### Supported Providers | Provider | Permissions | | ---------------------------- | ---------------------------------- | | **Google Workspace** | Read-only access to user directory | | **Microsoft 365 (Entra ID)** | Read-only access to user directory | ### Connecting a Directory 1. Open the **Integrations** page from the sidebar 2. In the **Directory Integration** card, click **Connect** next to your provider 3. Sign in with an **admin account** in the OAuth popup 4. Grant read-only access to the user directory 5. The connection status updates to "Connected" **Important:** An admin account on the provider side is required. Only read access to the user directory is requested — InterMIND never modifies your directory. ### Disconnecting Click **Disconnect** next to a connected provider to remove the directory integration. # Managing Your Subscription This guide covers how to start, change, and cancel your InterMIND subscription. ![Billing page](https://intermind.com/img/docs/billing/billing-page.png) ## Subscribing to a Plan ### From the Pricing Page ::steps ### Go to the Pricing page Visit the [Pricing](https://intermind.com/pricing) page (accessible from the footer or navigation). ### Choose billing interval Toggle between **Monthly** and **Yearly** billing (yearly saves 25%). ### Select a plan Click the plan's CTA: - **Basic** — **Get started free** (no checkout, free forever) - **Pro** — **Try for free** the first time, or **Get Pro** if you've already used your trial - **Business** — **Try for free** the first time, or **Get Business** if you've already used your trial - **Enterprise** — **Contact Sales** to fill out the inquiry form ### Complete checkout You'll be redirected to the secure Stripe checkout page. Enter your payment details and confirm. ### Plan activated After successful payment, you're redirected back to InterMIND with your new plan active. :: ### For Pro and Business Plans During checkout, you can adjust the **number of licenses** (seats) for your team. This lets you pay for exactly the number of team members you need. ### If You're Not Signed In If you click a plan button while not logged in: ::steps ### Redirect to login You'll be redirected to the login page. ### Sign in After signing in, the checkout process starts automatically for your selected plan. :: ## Changing Your Plan ### Upgrading To upgrade to a higher plan: ::steps ### Go to Pricing or Billing Visit the [Pricing](https://intermind.com/pricing) page or [Billing](https://intermind.com/billing) page. ### Click Upgrade Select the desired plan. ### Confirm in Stripe Portal The portal shows a prorated price for the remainder of your billing period. ### Done Confirm the upgrade — your new plan takes effect immediately. :: ### Downgrading To switch to a lower-tier plan: ::steps ### Go to Pricing or Billing Visit the [Pricing](https://intermind.com/pricing) page or [Billing](https://intermind.com/billing) page. ### Select the lower plan Choose the plan you want to switch to. ### Confirm in Stripe Portal The downgrade typically takes effect at the end of your current billing period. :: ### Switching Billing Interval To switch between monthly and yearly billing: 1. Open the **Stripe Customer Portal** (via Billing page) 2. Change your billing interval 3. Yearly billing applies the 25% discount automatically ## Cancelling Your Subscription ### How to Cancel ::steps ### Open Billing Go to the **Billing** page from the sidebar. ### Click Cancel You'll be redirected to Stripe's Customer Portal. ### Confirm cancellation Confirm the cancellation in the portal. :: ### What Happens After Cancellation - **During trial**: The subscription ends **immediately** — no further access to paid features - **After trial**: Your subscription remains active until the end of the current billing period - The Billing page shows "Subscription Ending" with the end date - After the period ends, your account reverts to the **Basic** (free) plan ### Reactivating If you've cancelled but the billing period hasn't ended yet: 1. Go to the **Billing** page 2. Click **Reactivate** 3. Your subscription continues as normal — no new charge until the next billing period ### Resubscribing If your subscription has fully expired: 1. Go to the [Pricing](https://intermind.com/pricing) page 2. Choose a plan and go through the checkout process 3. Note: the free trial is **not available** if you've used it before ## Payment Methods Payment is handled entirely through Stripe: - **Credit/debit card** is the accepted payment method - Manage your payment methods through the Stripe Customer Portal - Update your card before it expires to avoid payment failures ## Payment Issues If a payment fails: - You'll receive a notification about the failed payment - Stripe will retry the payment automatically - If the issue persists, update your payment method in the Customer Portal - Continued payment failure may result in your subscription being cancelled ## FAQ ### Can I switch plans at any time? Yes, you can upgrade or downgrade at any time. Upgrades take effect immediately with prorated billing. Downgrades take effect at the end of your billing period. ### What happens to my data if I cancel? Your data is preserved after cancellation. If you resubscribe, everything will be as you left it. However, if your storage exceeds the Basic plan limit, you'll need to upgrade to access all files. ### Can I get a refund? Refunds follow the refund policy of the payment provider that processed your purchase. Purchases made through Paddle are governed by [Paddle's Refund Policy](https://www.paddle.com/legal/refund-policy){rel=""nofollow""} and refund requests are handled by Paddle. Except where required by applicable law, subscription fees are non-refundable. ### Is my payment information secure? Yes. All payment processing is handled by Stripe. InterMIND never stores your credit card details. # Usage & Invoices The Billing page provides a clear view of your current plan, resource usage, and payment history. ## Accessing the Billing Page 1. Sign in to InterMIND 2. Click **Billing** in the sidebar navigation 3. The page displays your plan status, usage meters, and invoices ## Current Plan Status At the top of the Billing page, you'll see: - **Plan name** — Your current subscription tier (Basic, Pro, Business, or Enterprise) - **Billing status** — Whether you're on a trial, monthly, or annual plan - **Next billing date** — When your next payment is due - **Subscription status** — Active, trialing, or ending (if cancelled) Owners and admins can also set the **organization name** here — it shows up to teammates and on shared invitations. ## Usage Monitoring ### Voice Translation The **Voice Translation** meter shows your meeting translation usage over a **rolling 30-day window**: - **Progress bar** — Visual indicator of used vs. available translation time - **Exact numbers** — Minutes used out of your monthly allowance - **Color coding**: - **Blue/Green** — Normal usage (under 80%) - **Yellow/Warning** — Approaching limit (80%–99%) - **Red/Error** — At or over limit (100%) The 30-day window means usage gradually expires — as sessions from 30+ days ago fall off, your available time increases. Each plan includes a monthly allowance, measured per **meeting host** — minutes spent translating in any meeting you created count against your plan, not the listener's: | Plan | Monthly Translation Time | | -------------- | ------------------------ | | 🆓 Basic | 2 hours | | ⭐ Pro | 25 hours | | 🏢 Business | Unlimited | | 🏛️ Enterprise | Unlimited | #### When You Reach the Limit 1. **Warning** — you see a notification that you're approaching the limit 2. **Language lock** — at the limit, you can no longer change your translation language 3. **Translation stops** — real-time translation is paused until usage falls below the limit The meeting itself continues — chat, video, and every other feature keep working (chat translation has its own limit, below). If you consistently run out, [upgrade your plan](https://intermind.com/pricing) or wait for older usage to roll off the 30-day window. ### Storage The **Storage** meter shows your team's shared storage pool: - **Progress bar** — Used storage vs. plan limit - **Exact numbers** — GB used out of your storage allowance - **Same color thresholds** — Blue → Yellow (≥80%) → Red (≥100%) Storage is shared across your entire team — all members' files, recordings, and uploads count toward the same limit. ### Text Translation The **Text Translation** meter shows your chat translation usage over the same **rolling 30-day window**: - **Progress bar** — Words translated vs. plan limit. Only the free Basic plan has a word cap (10K/month); on every paid plan text translation is unlimited and the bar is replaced by an **Unlimited** label - **Exact numbers** — Words used out of your monthly allowance, shown in thousands (e.g. `3K / 10K` on Basic) - **Same color thresholds** — Blue → Yellow (≥80%) → Red (≥100%) Text translation is tracked per **meeting host** — only messages translated for participants in meetings you created count toward your limit. When the Basic cap is reached, new messages are still delivered — just not translated. See [Chat Translation](https://intermind.com/../chat/translation). ### Document Translation The **Document Translation** meter shows uploaded-document translations used over the same **rolling 30-day window**: - **Progress bar** — Documents translated vs. plan limit - **Exact numbers** — Documents used out of your allowance (e.g. `4 / 10 documents`) - Metered in **distinct files** — Pro 10, Business 30, Enterprise 100 per 30 days. Re-translating the same file into more languages does not count again - The free Basic plan shows **Not on this plan** — document translation requires a paid plan ## Invoices ### Viewing Invoices The invoices section shows your complete payment history: | Column | Description | | ----------- | ------------------------------------------ | | **Date** | When the invoice was generated | | **Total** | The amount charged | | **Status** | Payment status badge (Paid, Open, or Void) | | **Actions** | Link to view the full invoice on Stripe | ### Invoice Details Click **View** on any invoice to open the full invoice on Stripe's hosted page, where you can: - See itemized charges - Download a PDF copy - View payment method used ### Pagination Invoices are loaded 10 at a time. Click **Load more** at the bottom to see older invoices. ## Understanding Your Usage ### Translation Time Calculation - Translation time is tracked **per meeting host** — usage in any meeting you created counts against your plan, regardless of which participant triggered the translation - Displayed in minutes on the UI - Only active translation time counts — pausing translation stops the meter - Usage rolls off after 30 days, so the available limit gradually refreshes ### Storage Calculation - Storage is calculated as the **total size of all files** across the team - Includes: uploaded files, shared images, screen recordings - Does not include: text messages, user profiles ## Tips - Check the Billing page regularly to avoid hitting usage limits unexpectedly - If translation usage is consistently high, consider upgrading to a plan with more hours - Storage is shared — coordinate with your team if storage is getting full - Download important recordings locally if you're approaching the storage limit - Set up a yearly billing cycle to save 25% on your plan # Billing & Plans InterMIND offers four subscription tiers designed for individuals, teams, and organizations. All plans include core features like real-time translation, video calls, and screen sharing. ## Plans Overview | Feature | 🆓 Basic | ⭐ Pro | 🏢 Business | 🏛️ Enterprise | | ---------------------------- | ------------- | -------------- | ----------------- | ----------------- | | **Best for** | Individuals | Professionals | Teams | Organizations | | **Translation time** | 2 hours/month | 25 hours/month | Unlimited | Unlimited | | **Text translation** | 10K words | Unlimited | Unlimited | Unlimited | | **Document translation** | — | 10 docs/month | 30 docs/month | 100 docs/month | | **Meeting participants** | 50 | 100 | 300 | 1,500 | | **Team members** | 1 | Unlimited | Unlimited | Unlimited | | **Storage** | 1 GB | 30 GB | 1 TB | 5 TB | | **Meeting duration** | 1 hour | Unlimited | Unlimited | Unlimited | | **SSO (OIDC with your IdP)** | — | — | ✓ | ✓ | | **Support** | Live chat | Live chat | Live chat & video | Live chat & video | ::callout{type="tip"} **Most popular**: The **Pro** plan is the best choice for most professionals — it includes 25 hours of translation, unlimited text translation, and unlimited team members. [See pricing →](https://intermind.com/pricing) :: ### Features Included in All Plans ::callout{type="info"} Every plan comes with all core features — no hidden limitations: - ✅ FHD video & audio calls - ✅ Real-time translation with live subtitles - ✅ Real-time transcription - ✅ Text translation (10K words/month on Basic, unlimited on all paid plans) - ✅ Screen sharing - ✅ Cloud & local recording - ✅ File sharing in chat - ✅ Full chat history - ✅ Private translation infrastructure \:: ## Pricing ### Billing Intervals Choose between two billing options: - **Monthly** — Pay month-to-month, cancel anytime - **Yearly** — **Save 25%** compared to monthly billing Yearly prices are displayed as monthly equivalents (annual price ÷ 12) for easy comparison. ### Enterprise Pricing Enterprise plans have custom pricing. Click **Contact Sales** on the [pricing page](https://intermind.com/pricing) to fill out an inquiry form with your requirements. ## Free Trial All plans (except Enterprise) include a **14-day free trial**: - Full access to all plan features during the trial - A credit card is required to start the trial - You can cancel anytime during the trial without being charged - If you don't cancel, billing starts automatically after 14 days - **One trial per account** — once used, the trial option is no longer available :::callout{type="info"} If you cancel during the trial period, the subscription ends immediately rather than continuing until the trial expires. ::: ## New Customer Offer New customers (who have never had a subscription) receive **50% off the first 3 months**: - The discount applies automatically at checkout - Works with both monthly and yearly plans - Applies to all users in your organization - After 3 months, standard pricing resumes - Can be combined with the 25% annual discount ## Getting Started - [Managing Your Subscription](https://intermind.com/billing/manage) — Upgrade, downgrade, or cancel your plan - [Usage & Invoices](https://intermind.com/billing/usage) — Monitor usage and view invoices :: # Audio Issues ## Microphone Not Working | Symptom | Likely cause | Fix | | ----------------------------------------------- | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | | "Microphone access denied" prompt | Browser permission blocked | Click the lock icon in the address bar → set **Microphone** to **Allow** → reload | | Mic icon shows a red slash, no audio level | Mic is muted or wrong device selected | Click the arrow next to the microphone button in the control bar → pick the right device | | Audio level indicator stays flat when you speak | OS-level mic permission blocked | macOS: System Settings → Privacy & Security → Microphone → enable browser. Windows: Settings → Privacy → Microphone | | Mic works in other apps but not InterMIND | Another browser tab is holding the device | Close other tabs that may use the mic (Zoom Web, Google Meet, etc.) and reload | | Bluetooth headset connects but no audio | Headset is in low-quality (HFP) mode | Disconnect and reconnect the headset, or switch to wired | ## No One Can Hear You | Symptom | Likely cause | Fix | | --------------------------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------ | | Your audio level moves but others hear silence | Wrong mic device routed | Open **Settings** → **Devices** → select the correct microphone | | Others hear muffled/distant audio | Laptop's far-field mic instead of headset | Switch device in the control bar; physically move closer to the laptop if no headset | | You speak but only one participant doesn't hear you | That participant has their speaker muted or wrong output device | Ask them to check their speaker selection — not your issue | ## You Can't Hear Others | Symptom | Likely cause | Fix | | ------------------------------------------------ | --------------------------------------------- | ------------------------------------------------------------------------------------ | | Total silence from everyone | Browser tab muted (right-click tab → Unmute) | Unmute the browser tab | | Silence from one participant only | They are muted, or their mic is broken | Check their mic indicator in the participant list | | Translation audio playing but no original voices | Translation replaces original audio by design | This is expected — see [Real-Time Translation](https://intermind.com/../translation) | | Volume too low even at 100% | OS output is low or wrong output device | Check OS sound settings; switch output device | ## Echo or Feedback | Symptom | Likely cause | Fix | | ----------------------------------------- | ----------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | | Everyone hears their own voice with delay | Someone in the room is using a speaker without a headset | The offender should mute, switch to a headset, or use the **Echo cancellation** toggle in browser audio settings | | Crackling, popping, or robotic voice | Bluetooth bandwidth saturated, or another app using the mic | Switch to wired audio; close other audio apps | ## Related - [Meeting Controls](https://intermind.com/../meetings/controls) — Mic and speaker selection during a call - [Browsers & Devices](https://intermind.com/browsers) — Required permissions and supported browsers # Video & Camera Issues ## Camera Not Detected | Symptom | Likely cause | Fix | | -------------------------------------------- | ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------- | | "Camera not found" or no devices in selector | Camera blocked at OS level | macOS: System Settings → Privacy & Security → Camera → enable browser. Windows: Settings → Privacy → Camera | | Browser asks for permission every join | Permission was denied or set to "Ask every time" | Address bar → lock icon → set **Camera** to **Allow** | | External webcam not listed | Driver issue or USB hub power | Replug directly into the laptop, not through a hub; reinstall driver | | Camera works in other apps but not InterMIND | Another tab is holding the device | Close other tabs using the camera (Google Meet, Zoom Web, etc.) and reload | ## Black or Frozen Video Tile | Symptom | Likely cause | Fix | | ----------------------------------------------------------- | ----------------------------------------------- | ------------------------------------------------------------------- | | Your tile shows your avatar instead of video | Camera is off (click the camera icon to toggle) | Toggle camera back on | | Your tile is black even with camera on | Camera privacy shutter is closed | Open the physical shutter on your webcam | | Video frozen on your end, others see you fine | Local rendering issue | Reload the page (Cmd/Ctrl+R) | | Video frozen on others' end | Network bandwidth too low | See [Network & Connection](https://intermind.com/network) | | Wrong camera being used (e.g. external instead of built-in) | Default device mismatch | Click the arrow next to the camera button → select the right device | ## Video Quality Issues | Symptom | Likely cause | Fix | | ------------------------------------ | ----------------------------------------- | --------------------------------------------------------------------------------------------- | | Pixelated, low-resolution video | Bandwidth-saving auto-adjust | Move closer to your router, switch to wired; InterMIND auto-restores HD when bandwidth allows | | Lighting is too dark | Your face is back-lit (window behind you) | Add a front-facing light or sit facing a window | | Video lag (you move, video stutters) | CPU bottleneck from other apps | Close heavy apps (browser tabs with video, IDEs running builds) | ## Screen Sharing | Symptom | Likely cause | Fix | | ----------------------------------- | ---------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- | | "Screen sharing not supported" | Safari ≤16 doesn't support `getDisplayMedia()` on some platforms | Use Chrome, Edge, or update to Safari 17+ | | Black screen when sharing on macOS | Screen Recording permission not granted | System Settings → Privacy & Security → Screen Recording → enable the browser, then quit and reopen the browser | | Audio not shared when sharing a tab | Tab audio sharing requires Chrome/Edge | Switch browsers, or share window without audio and unmute the source | ## Related - [Screen Sharing & Recording](https://intermind.com/../meetings/screen-sharing) — How sharing works during a call - [Meeting Controls](https://intermind.com/../meetings/controls) — Camera selection and toggle # Translation Quality This page covers issues unique to InterMIND's real-time translation. For general translation behavior, see [Choosing Languages](https://intermind.com/../translation/languages). ## Wrong Language Detected | Symptom | Likely cause | Fix | | --------------------------------------------------- | ------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | | Your speech is being detected as the wrong language | Speaking language not set or auto-detect picked the wrong one | Open the language selector in the control bar → set your **speaking language** explicitly | | Translation comes out in a language you don't want | Listening language is wrong | Click the globe icon → set your **listening language** | | Multilingual speakers cause language flipping | Auto-detect re-evaluates on each utterance | Pin your speaking language to one value; switch manually if you change languages | ## Translation Lag | Symptom | Likely cause | Fix | | ---------------------------------------------- | ---------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | | 3+ second delay between speech and translation | Normal latency window — translation waits for utterance boundary | This is expected; speak in complete sentences for best results | | Lag growing over time during a long meeting | Network buffer accumulating, packet loss | Check the connection indicator; see [Network & Connection](https://intermind.com/network) | | Translation stops mid-sentence | Speech recognition lost the audio stream | Pause, then speak a complete sentence — the system re-anchors at utterance boundaries | ## Missing Subtitles | Symptom | Likely cause | Fix | | ----------------------------------------------------- | ---------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ | | No subtitles appear at all | Subtitles toggle is off | Click the **CC** / subtitles button in the control bar | | Subtitles appear in original language, not translated | Translation is disabled, only transcription on | Enable translation via the globe icon | | Subtitles cut off after one line | Subtitles auto-clear after a few seconds | This is expected — full transcript is preserved, see [Transcription](https://intermind.com/../translation/transcription) | ## Accuracy & Word Errors | Symptom | Likely cause | Fix | | ------------------------------------------------- | ---------------------------------------------------- | ------------------------------------------------------------------------------ | | Brand names, acronyms, or jargon mistranslated | The system has no glossary context | Speak slowly when introducing technical terms; spell out acronyms on first use | | Numbers transcribed wrong (e.g. "fifteen" → "50") | Background noise or accent | Use a headset mic; reduce ambient noise | | Translation says the opposite of what was meant | Negation or context loss across utterance boundaries | Speak in complete sentences with explicit subject + verb + object | | Names of people appear translated/garbled | Names are treated as nouns by some language pairs | Repeat the name or write it in chat | ## Background Noise Breaks Translation | Symptom | Likely cause | Fix | | ------------------------------------------------------------------ | ----------------------------------------- | --------------------------------------------------------------- | | Translation works in quiet rooms but fails in cafes / open offices | Ambient noise confuses speech recognition | Use a directional headset mic; enable browser noise suppression | | Music or TV in background gets transcribed as speech | The recognizer hears it as utterances | Mute the source; use a noise-cancelling headset | | Multiple speakers overlapping → mixed translation | Two voices in the same utterance | Encourage one-speaker-at-a-time; use hand-raise | ## Related - [Choosing Languages](https://intermind.com/../translation/languages) — Supported languages and how detection works - [Live Subtitles](https://intermind.com/../translation/subtitles) — Subtitle behavior and controls - [Usage & Limits](https://intermind.com/../billing/usage) — Per-plan minute limits # Network & Connection InterMIND needs a stable connection in both directions. Most "InterMIND is broken" reports are really network issues. ## Bandwidth Requirements | Activity | Minimum | Recommended | | -------------------------------- | -------------- | -------------- | | Audio-only meeting + translation | 1 Mbps up/down | 2 Mbps up/down | | HD video meeting + translation | 2 Mbps up/down | 5 Mbps up/down | | Screen sharing | +1 Mbps up | +2 Mbps up | Run a [speed test](https://fast.com){rel=""nofollow""} to check your real bandwidth — not the rated speed of your plan. ## "Reconnecting…" Loop | Symptom | Likely cause | Fix | | ------------------------------------------- | ------------------------------------- | -------------------------------------------------------------- | | Banner keeps flickering "Reconnecting…" | Packet loss on the route | Switch from Wi-Fi to wired; move closer to the router | | Reconnects only when video is on | Bandwidth not sufficient for HD video | Turn camera off; InterMIND auto-falls-back to audio-only | | Reconnects when others share their screen | Total downstream bandwidth exceeded | Ask the sharer to lower share quality or close other downloads | | Reconnects in offices with strict firewalls | UDP / WebRTC ports blocked | See **Firewall & VPN** below | ## Choppy Audio or Video | Symptom | Likely cause | Fix | | ------------------------------- | ------------------------------------------------------ | ----------------------------------------------------------------------------------- | | Robotic / chopped voice | Packet loss > 5% | Switch to wired, close bandwidth-heavy apps (cloud sync, video calls in other tabs) | | Video freezes for a few seconds | Bandwidth spike consumed by another app | Pause Dropbox/Google Drive sync, OS updates, BitTorrent | | Audio cuts but video is fine | Audio uses different codec — packet loss on audio path | Toggle camera off and on; if it persists, switch network | | Issues only when on VPN | VPN adds latency and may rate-limit UDP | Disconnect VPN if your IT policy allows | ## Firewall, VPN, and Corporate Networks InterMIND uses WebRTC, which needs specific ports open. | Port / Protocol | Purpose | Required | | --------------- | ---------------------------- | --------------------------------------------- | | TCP 443 | Signaling, fallback media | Yes | | UDP 3478 | STUN (NAT traversal) | Recommended | | UDP 1024–65535 | Media (audio, video, screen) | Recommended; falls back to TCP 443 if blocked | If your network blocks UDP entirely, InterMIND falls back to TCP 443 — it works but quality is reduced. If you're on a managed corporate network, ask IT to allow `*.intermind.com` and these ports. ## Translation Cuts Out Mid-Sentence | Symptom | Likely cause | Fix | | ------------------------------------------------------------------ | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | | Translation stops, then resumes after a delay | Speech stream interrupted by packet loss | Check connection indicator; see [Translation Quality](https://intermind.com/translation-quality) | | Translation works at the start of the meeting, fails after 30+ min | Bandwidth degraded over time (Wi-Fi roaming, network congestion) | Reconnect on a stable network | ## Diagnostic Steps 1. Run the [WebRTC Browser Test](https://intermind.com/browser-test) — it reports latency, packet loss, and TURN reachability 2. Open browser DevTools → Network tab → check for failing requests to `*.intermind.com` 3. Try a different network (mobile hotspot) — if it works there, the office network is the issue ## Related - [Browsers & Devices](https://intermind.com/browsers) — Browser-level settings that affect connection - [Meeting Controls](https://intermind.com/../meetings/controls) — Connection indicator in the control bar # Browsers & Devices ## Supported Browsers | Browser | Minimum version | Notes | | ---------------------------------- | --------------------- | --------------------------------------------- | | **Chrome** | 90+ | Best support, full feature set | | **Edge** | 90+ | Same engine as Chrome — full feature set | | **Safari** | 15+ | Works; screen-sharing requires Safari 17+ | | **Firefox** | 100+ | Works; some codec fallbacks on older hardware | | **Opera, Brave, Arc** | Chrome 90+ equivalent | Works (built on Chromium) | | **Internet Explorer, legacy Edge** | — | Not supported | Run the [WebRTC Browser Test](https://intermind.com/browser-test) to verify your browser before joining an important meeting. ## Permission Prompts | Symptom | Likely cause | Fix | | ------------------------------------- | ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | | Permission prompt loops on every join | Browser is set to "Ask every time" | Click the lock icon in the address bar → set **Microphone** and **Camera** to **Allow** for the InterMIND domain | | "Permission denied" with no prompt | Permission was previously denied | Same as above — change from **Block** to **Allow**, then reload | | Permissions reset every session | Browsing in Incognito / Private mode | Use a regular window, or grant permission again each session | | OS-level prompt never appeared | Browser was previously denied at OS level | macOS: System Settings → Privacy & Security → Camera/Microphone → enable browser. Windows: Settings → Privacy → Camera/Microphone | ## Safari Quirks | Symptom | Likely cause | Fix | | ----------------------------------------- | ---------------------------------------------------------- | -------------------------------------------------------- | | Audio cuts out when tab is in background | Safari throttles inactive tabs aggressively | Keep the InterMIND tab in the foreground during meetings | | Camera switches to wrong device on rejoin | Safari doesn't always honor "default device" memory | Manually select the device after joining | | Screen sharing missing | Safari ≤16 doesn't support `getDisplayMedia()` reliably | Upgrade to Safari 17+ or use Chrome/Edge for sharing | | Echo cancellation weak | Safari's audio processing is less aggressive than Chrome's | Use a headset to avoid echo | ## Mobile Browsers Mobile is supported but has limits versus desktop. | Feature | iOS Safari | Android Chrome | | ---------------------------- | ----------------------------------------------- | ------------------------------------- | | Audio + Video meetings | ✓ | ✓ | | Translation (listen + speak) | ✓ | ✓ | | Screen sharing | ✗ (iOS doesn't expose `getDisplayMedia` to web) | ✗ | | Background mode | Limited — Safari pauses inactive tabs | Limited — Chrome pauses inactive tabs | | Picture-in-picture | ✓ on iOS 15+ | ✓ | For meetings longer than 30 minutes or with screen sharing, desktop is recommended. ## Incognito / Private Mode | Issue | Cause | | ------------------------------------------------ | --------------------------------------------- | | Permissions reset each session | Private mode wipes site data on close | | Sign-in fails with "third-party cookies blocked" | Private mode blocks cookies more aggressively | | Performance worse than normal window | Less caching, more overhead | For meetings, use a regular browser window. ## Browser Extensions Some extensions break WebRTC: | Extension type | What it breaks | | --------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | | Ad blockers (uBlock, AdBlock) | Usually fine; rarely block WebRTC signaling — disable for `*.intermind.com` if issues persist | | Privacy / VPN extensions (NordVPN, Ghostery) | Can block STUN/TURN — disable during meetings | | Tab managers / freezers (The Great Suspender, Auto Tab Discard) | Suspend the InterMIND tab and break the call — disable for InterMIND | | Anti-tracking extensions | Can block analytics endpoints, usually harmless | Test in an incognito window with extensions off — if it works there, an extension is the cause. ## Clearing Cache If something is stuck and nothing else works: 1. Open browser DevTools (Cmd/Ctrl+Shift+I) 2. Right-click the reload button → **Empty Cache and Hard Reload** 3. If still broken: Settings → clear site data for `intermind.com` ## Related - [Audio Issues](https://intermind.com/audio) — Mic and speaker problems - [Video & Camera Issues](https://intermind.com/video) — Camera permissions and screen sharing - [Network & Connection](https://intermind.com/network) — Bandwidth and firewall issues # Troubleshooting Most issues fall into one of five categories. Each page lists symptoms, the likely cause, and the fix. ## Where to Start | If you're seeing… | Open | | ------------------------------------------------------------- | -------------------------------------------------------------------------------- | | No one hears you, or you can't hear others | [Audio Issues](https://intermind.com/troubleshooting/audio) | | Your camera is black, frozen, or not detected | [Video & Camera Issues](https://intermind.com/troubleshooting/video) | | Wrong language, lag, missing subtitles, mistranslation | [Translation Quality](https://intermind.com/troubleshooting/translation-quality) | | "Reconnecting…", choppy stream, dropouts | [Network & Connection](https://intermind.com/troubleshooting/network) | | InterMIND won't start, permission prompt loops, mobile issues | [Browsers & Devices](https://intermind.com/troubleshooting/browsers) | ## First Step: Run the Browser Test Before deep-diving, the [WebRTC Browser Test](https://intermind.com/browser-test) checks your microphone, camera, network, and codec support in 30 seconds. If it fails, the issue is local — fix that first. ::callout{type="tip"} The browser test report tells you exactly which subsystem failed (mic permission, camera permission, STUN/TURN reachability, codec support). Most "InterMIND doesn't work" tickets resolve here. :: ## Still Stuck? If none of the pages here match what you're seeing, [contact support](mailto\:support@intermind.com) with: - The exact symptom (one sentence) - Browser name + version - Operating system - A screenshot of the [browser test](https://intermind.com/browser-test) result # Security & Privacy This page describes how InterMIND handles your data in plain terms. For legal language, see the [Privacy Policy](https://intermind.com/privacy) and [Terms of Service](https://intermind.com/terms). ## In Transit All connections to InterMIND go over **HTTPS / WSS** (TLS). HTTP requests are redirected to HTTPS automatically — there is no plaintext fallback. This includes the web app, the WebSocket server that carries chat and signaling, and any media routed through our infrastructure. ## Authentication | Method | Status | | --------------------------------------------------------------------------------------- | ------------------------------------------ | | **Email + verification code** | Available | | **Sign in with Google** (OAuth 2.0) | Available | | **Sign in with Microsoft** (OAuth 2.0) | Available | | **SSO** (Google / Microsoft, automatic team join by verified domain) | Available on Business and Enterprise plans | | **Enterprise SSO via your own IdP** (OIDC — Okta, Microsoft Entra ID, Google Workspace) | Available on Business and Enterprise plans | Passwords are never stored — sign-in is either a one-time verification code or an OAuth flow with Google / Microsoft. Sessions are kept in HTTP-only cookies; you can sign out from the profile menu at any time. Enterprise SSO uses OpenID Connect Authorization Code with **PKCE (S256)**, **state**, and **nonce**. ID tokens are validated against the IdP's published **JWKS** (signature, issuer, audience), and a sign-in is accepted only for email domains the team has **verified via DNS** — your IdP is authoritative only for domains you have proven to own. OIDC client secrets are encrypted at rest, and every SSO sign-in is recorded in the team audit log. SAML 2.0 sign-in is in development. Setup: [SSO Setup](https://intermind.com/docs/integrations/sso-setup). ## Real-Time Translation Real-time voice and subtitle translation runs on **InterMIND's own infrastructure** — speech audio is processed on our private WebSocket service, not sent to a public OpenAI / Google Translate / Azure endpoint. Document translation (PDF, DOCX, PPTX, XLSX) uses **DeepL** as the translation provider. Files are uploaded to DeepL over a TLS connection for translation and the result is returned to your meeting chat. DeepL's data handling is governed by [DeepL's own privacy terms](https://www.deepl.com/en/privacy/){rel=""nofollow""} — they do not retain content for training. ## Recordings and Transcripts | Data | Storage | Retention | | ------------------ | ------------------------------------------ | --------------------------------------------------------------------------- | | Meeting recordings | S3-compatible object storage | Stored until you delete them | | Transcripts | Linked to the recording / meeting | Same as the recording | | Chat messages | Database, linked to the meeting or channel | Until you delete them, or — for ad-hoc meetings — purged when the call ends | You control retention: your content stays available for as long as your account or team workspace is active, and you can delete any recording, channel, or message at any time. Deleting your account permanently erases all associated content across our database and storage. We do not impose an automatic expiry — you decide how long your content is kept. See our [Privacy Policy](https://intermind.com/privacy) for the full retention statement. ## Where Your Data Lives InterMIND's primary application region is **Paris (CDG, France)** on Fly.io; the database lives in **Frankfurt, Germany**. Recordings and files are kept on S3-compatible storage pinned to **EU regions (Frankfurt / Amsterdam)**. There is currently no per-customer region pinning. ## What Other Parties See | Party | What they see | | ------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | **InterMIND** | Meeting metadata, chat content, recordings (until you delete them), transcripts, your account info | | **DeepL** | Only the contents of documents you ask to translate, on a per-file basis | | **Mistral** (French AI provider) | The meeting transcript, once at meeting end, to generate the post-meeting summary — processed with zero data retention and never used for training | | **Google / Microsoft** (if you sign in with them) | Your name, email, profile photo — standard OAuth scopes only | | **Stripe** (if you pay) | Billing details (card, address) — InterMIND never sees the raw card number | | **Sentry, PostHog** (error & analytics) | Anonymous usage events and crash traces; meeting content is never sent | ## What InterMIND Does Not Do - **No model training on your data.** Conversations, documents, and recordings are not used to train any AI model. - **No selling or sharing.** Your data is not sold or shared with advertisers. - **No third-party tracking inside meetings.** No tracking pixels, no ad SDKs in the meeting room. ## Deleting Your Data | To delete | How | | ------------------------- | ------------------------------------------------------------------------------------------------------------ | | A single recording | Open the meeting → Recordings → delete | | A standalone chat history | Open the channel → Delete Channel Permanently (removes the channel and all its messages, files, and history) | | Your entire account | Profile → Settings → Delete account, or contact | Account deletion removes your profile, meetings, recordings, transcripts, and chat history. Backups may take up to 30 days to age out. ## Reporting a Security Issue Email . ## What the DPA covers The DPA governs the personal data we process to provide the Service on your behalf: - **Subject matter and duration:** provision of the InterMIND meeting and translation service for the term of your agreement with us. - **Nature and purpose:** hosting and transmission of meetings (audio and video), realtime speech transcription and translation, chat, document translation, an optional post-meeting AI summary, billing, and support. - **Categories of data subjects:** your users and their meeting guests, including external invitees. - **Categories of personal data:** account data (email, name, photo), meeting metadata, speech transcriptions (including speaker names), chat messages and attachments, recordings, and billing data. The Service does not require special categories of data; meeting content may incidentally contain anything participants choose to discuss. ## Our commitments as processor The DPA reflects the Article 28(3) GDPR obligations, including: - We process personal data only on your documented instructions. - Our personnel are bound by confidentiality. - We apply appropriate technical and organizational security measures. - We use subprocessors under a general written authorization and give at least 30 days' prior notice of any addition or replacement, so you can object on reasonable data-protection grounds. The current list is published at [/legal/subprocessors](https://intermind.com/legal/subprocessors). - We assist with data-subject requests. Deletion and export are self-service in the product; we assist with the rest. - We notify you of a personal data breach affecting your data without undue delay and in any event within 72 hours of becoming aware of it. - On termination we delete or return your personal data, subject to backup aging. We do not currently hold SOC 2 or ISO 27001 certification. On request we provide a security questionnaire and our compliance documentation instead. ## International transfers Processing happens in the EU by default (see the [subprocessor list](https://intermind.com/legal/subprocessors) for provider regions). Where a provider's corporate entity is outside the EU/EEA, transfers are covered by Standard Contractual Clauses (SCCs) and, where applicable, the EU–US Data Privacy Framework (DPF), together with supplementary technical measures such as encryption in transit and at rest. Because InterMIND is operated from the UAE, the specific Chapter V mechanism and SCC module are confirmed per customer when the DPA is signed. ## Requesting a signed copy The DPA is signed per customer, so we adapt the international-transfer terms to your situation before signature — it is not a self-serve download. To request a copy for review or signature, email [privacy@mind.com](mailto\:privacy@mind.com?subject=DPA%20request) with the subject "DPA request". For agreements signed by MindMeeting OÜ through the sales channel, the DPA is executed by MindMeeting OÜ as processor; Golden Fish CSP LLC (service operator) and the providers listed above act as subprocessors. # Privacy Policy **Effective date: 17 June 2026** ## 1. Who we are InterMIND ("the Service") — the meeting platform with real-time speech translation available at intermind.com and through the InterMIND mobile apps — is operated by: **Golden Fish Corporate Services Provider LLC** ("Golden Fish CSP LLC", "we", "us"), City Avenue Building, Office 405-070, Port Saeed, Dubai, United Arab Emirates. For personal data processed through the Service, Golden Fish CSP LLC acts as the data controller, except where your organization uses InterMIND under its own agreement with us — in that case your organization is the controller and we process data on its behalf (see our [Data Processing Addendum](https://intermind.com/legal/dpa)). **Publisher and Intellectual Property Owner:** MindMeeting OÜ (Estonia), Juhkentali 8, Tallinn 10132, Estonia. **Service Operator and Contracting Entity:** Golden Fish Corporate Services Provider LLC (United Arab Emirates). The InterMIND mobile apps are published by MindMeeting OÜ on behalf of, and under license from, Golden Fish CSP LLC, which operates the Service and is responsible for personal data processed through it. **Privacy contact:** **EU/EEA representative (Article 27 GDPR):** MindMeeting OÜ, Juhkentali 8, Tallinn 10132, Estonia, acts as our representative in the European Union under Article 27 GDPR. EU/EEA users and supervisory authorities may contact our representative on data-protection matters at (subject line "EU representative"). This policy covers intermind.com and the InterMIND apps only. The corporate site mind.com is operated separately and has its own policy. ## 2. What we collect and why | What | Details | Why (purpose) | Legal basis | | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | ----------------------------------------------------------- | | Account data | Email, display name, profile photo, language preferences | Creating and operating your account | Contract | | Meeting audio/video | Live streams during a call — **transient, never stored** (see §3) | Running the call: transmission, speech recognition, real-time translation | Contract | | Transcriptions | Recognized and translated speech with speaker names | Live captions/translation; available to participants after the meeting until deleted | Contract | | Recordings & files | Meeting recordings you start, chat attachments, documents | Making your content available to you and participants | Contract | | AI meeting summary | Your meeting transcript is sent **once, at meeting end** to generate a digest | Post-meeting summary in the meeting chat | Legitimate interest (disclosed here) | | Document translation | Contents of documents you submit for translation | Translating the document at your request | Contract | | Chat messages | Message text and edit history; ad-hoc meeting chats are purged when the call ends | Messaging | Contract | | Billing data | Name, email, billing address, subscription and usage records. **Card numbers never touch our systems** — payments are handled by Stripe | Charging for paid plans; tax/accounting obligations | Contract; legal obligation | | Usage analytics | Product events, session replays, error traces — **only after you consent via the cookie banner** (analytics is off by default) | Improving the product, fixing errors | Consent (analytics); legitimate interest (error monitoring) | | Transactional email | Your email address, one-time sign-in codes, notifications | Sign-in and service notifications | Contract | | Sales inquiries | Name, email, company, message from contact/partner forms | Responding to your inquiry | Legitimate interest | We do not sell personal data, and we do not use your meeting content to train AI models. ## 3. How meetings are processed - **Audio and video streams are not stored.** They pass through our media infrastructure (hosted in France) for transmission, speech recognition, and translation, and exist only for the duration of the call. Only what is listed above — transcriptions, recordings you explicitly start, chat — is persisted. - **AI summaries** are generated by Mistral AI (France) under a zero-data-retention agreement: the transcript is processed once and is not retained or used for training by Mistral. - **Recording is visible to participants.** You are responsible for complying with the laws that apply to you when recording or transcribing a conversation (some jurisdictions require the consent of all participants). ## 4. Who we share data with (subprocessors) We use a small set of infrastructure and service providers. The full, versioned list — including each provider's purpose, processing region, and safeguards — is published at [/legal/subprocessors](https://intermind.com/legal/subprocessors). Headlines, verified against our infrastructure: - Application hosting, database, realtime servers, object storage, analytics, and error monitoring all run in the **EU** (Frankfurt, Paris, EU multi-region storage). - Meeting media, speech recognition, and translation run on infrastructure in **France**. - Stripe (payments), Google/Microsoft (optional OAuth sign-in), and the corporate entities of some EU-hosted providers are US-based — covered by Standard Contractual Clauses and/or the EU–US Data Privacy Framework. ## 5. International transfers Processing happens in the EU by default (see §4). Where a provider's corporate entity is outside the EU/EEA, transfers are covered by Standard Contractual Clauses or an adequacy mechanism. The United Arab Emirates, where Golden Fish CSP LLC is established, has comprehensive data-protection legislation but is not the subject of an EU adequacy decision. Where personal data is accessed from, or transferred to, a country outside the EEA — including access by us for the administration of the Service — we rely on appropriate safeguards under Chapter V GDPR (Standard Contractual Clauses and, where applicable, the EU–US Data Privacy Framework) together with supplementary technical measures such as encryption in transit and at rest and EU-pinned storage and processing. A copy of the relevant safeguards is available on request at . ## 6. Data retention We retain your content — meeting recordings, transcripts, translations, and chat messages — for as long as your account or team workspace remains active, so that it stays available to you. You control retention: you can delete individual recordings, channels, or messages at any time, and deleting your account permanently erases all associated content (across our database and storage) together with the cancellation of any active subscription. We do not impose an automatic expiry on your content; you decide how long it is kept. Some data is held only transiently for operational reasons: account-data exports are available for 7 days before deletion, anonymous guest sessions are purged within 24 hours, and one-time email verification codes are swept on expiry. ## 7. Your rights Depending on your jurisdiction (including under the GDPR), you have the right to access, rectify, erase, and export your data, to object to or restrict certain processing, and to withdraw consent at any time. Two of these are self-service, effective immediately: - **Erasure** — delete your account in Settings; this permanently removes your data from our database and storage and cancels any active subscription. - **Portability** — export your full account data as a ZIP archive from Settings (download link valid for 7 days). For anything else, contact . If you are in the EU/EEA, you can also lodge a complaint with your local supervisory authority. If you are in the United Arab Emirates, you have equivalent rights under the UAE Personal Data Protection Law (Federal Decree-Law No. 45 of 2021), which you may exercise through the same contact. ## 8. Cookies and analytics We use a consent management platform (Usercentrics) to ask for your consent before any non-essential cookies or analytics run. Analytics (PostHog, EU cloud) and session replay are **off by default** and start only if you opt in. Essential cookies (session, security) do not require consent. You can change your choice at any time via the cookie settings link in the footer. ## 9. Security TLS for all traffic (HTTPS/WSS, no plaintext fallback); encryption at rest for the database and object storage; no passwords stored (one-time email codes or OAuth only); HTTP-only session cookies; server-side role enforcement; speech text scrubbed from client-side logs; isolated per-environment databases. ## 10. Children The Service is not directed at children. You must be at least 16 years old to use the Service. If you are under the age of majority in your country (18 in the United Arab Emirates), you may use the Service only with the consent and under the supervision of a parent or legal guardian. We do not knowingly collect personal data from children below the applicable age; if you believe a child has provided us with personal data, contact and we will delete it. ## 11. Changes to this policy We will post any changes on this page and update the effective date. For material changes we will notify you in the product or by email. ## 12. Contact Golden Fish Corporate Services Provider LLC — City Avenue Building, Office 405-070, Port Saeed, Dubai, United Arab Emirates. Email: # Service Level Agreement On every **paid plan** — Pro, Business, and Enterprise — we commit to a **Monthly Uptime Percentage of at least 99.5%**, backed by service credits. This page explains how uptime is measured, the credit tiers, the exclusions, and how to claim. For paid subscriptions this SLA forms part of our [Terms of Service](https://intermind.com/terms); for Enterprise customers it is also incorporated into the Enterprise agreement, and where the two differ, the signed agreement prevails. The free Basic plan is provided as described in the Terms of Service, without a service level commitment. ## The commitment - **Monthly Uptime Percentage: 99.5%**, measured per calendar month. - If we miss it, you are entitled to a service credit (see below). - Current and historical availability of every core service path is published live on [System Status](https://intermind.checkly-dashboards.com/){rel=""nofollow""} — you can verify the commitment at any time. ## How uptime is measured **Monthly Uptime Percentage** = the percentage of minutes in the calendar month during which the Service was not Unavailable. **"Unavailable"** means the core functions of the Service — signing in, joining a meeting, meeting audio and video, or real-time translation — do not work for your users due to causes within our control. Availability is measured by our continuous synthetic monitoring, which independently probes each core service path (application and database, realtime channel, translation engine, and authentication) around the clock; the Service counts as available only when all of them are. The results of these probes are public — see [System Status](https://intermind.checkly-dashboards.com/){rel=""nofollow""} for the current state and history of each path. Unavailability does not include degraded performance of individual features, or the accuracy of transcriptions and translations — machine translation quality is addressed in our [Terms of Service](https://intermind.com/terms), not this SLA. ## Service credits If the Monthly Uptime Percentage for a calendar month falls below the commitment, you can claim a credit against that month's fee for the affected subscription: | Monthly Uptime Percentage | Service credit | | ------------------------------ | ---------------------- | | Below 99.5%, at or above 99.0% | 10% of the monthly fee | | Below 99.0% | 25% of the monthly fee | Credits are applied to a future invoice, are capped at 30% of the affected month's fee in aggregate, cannot be exchanged for cash, and are your sole and exclusive remedy for our failure to meet the SLA. If you access the Service under a signed agreement with us that states a different credit basis — for example a partner agreement, where credits are calculated against the amounts payable to us for the affected period — that basis applies in place of the monthly fee, with the same tiers, cap, and claim process. ## Exclusions Downtime caused by the following does not count as Unavailability: - Scheduled maintenance that we announce in advance. - Factors outside our reasonable control: force majeure, internet or telecommunications failures beyond our network, or outages of your own equipment, network, or software. - Your violation of the Terms of Service, or suspension of your account under them (for example, non-payment or abuse). - Beta, preview, or trial features identified as such. - Months in which no fee was payable for the affected subscription — the free Basic plan and free trial periods (there is no fee to credit). ## How to claim a credit Email [support@intermind.com](mailto\:support@intermind.com?subject=SLA%20credit%20claim) with the subject "SLA credit claim" within 30 days after the end of the affected calendar month, including the dates and times of the incidents you observed. We verify claims against our monitoring data and confirm the credit within a reasonable time. ## Getting the SLA The SLA applies automatically to every paid subscription — no separate signature needed. Enterprise customers additionally get it incorporated into their Enterprise agreement; to set that up, start from the [pricing page](https://intermind.com/pricing) and contact sales. # Subprocessors To run InterMIND we use a small set of infrastructure and service providers ("subprocessors"). This page lists each one, what it processes, where, and under what safeguards. **Service Operator and Contracting Entity:** Golden Fish CSP LLC (UAE). **Publisher and Intellectual Property Owner:** MindMeeting OÜ (EE). Processing happens in the EU by default. Where a provider's corporate entity is outside the EU/EEA, transfers are covered by Standard Contractual Clauses (SCCs) and/or the EU–US Data Privacy Framework (DPF). We update this list whenever a subprocessor is added or removed. | Subprocessor | Entity domicile | Purpose | Personal data processed | Processing location | Safeguards | | ---------------------------------------- | --------------- | ------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | ------------------------------------------------------ | | Vercel Inc. | US | Web app hosting, serverless functions, AI Gateway (proxy) | All app traffic incl. meeting metadata, chat | Functions pinned to Frankfurt (fra1); AI Gateway has a zero-data-retention policy (no prompt retention) | DPA + SCCs; ZDR | | Fly.io Inc. | US | Realtime WebSocket server hosting | Chat fan-out, conference events, transcription words in transit | Paris (CDG) | DPA + SCCs | | Neon Inc. | US | Postgres (primary database) | Accounts, meetings, messages, transcripts, billing usage | Frankfurt (AWS eu-central-1) | DPA + SCCs; encryption at rest | | Tigris Data Inc. | US | Object storage | Meeting recordings, chat attachments, documents, export archives | EU multi-region (Frankfurt + Amsterdam) | DPA; TLS | | MindMeeting OÜ (Mind API) | EE | Meeting media (SFU), speech recognition, realtime translation | Audio/video streams, speech, participant names/roles | OVH, France | Intra-group agreement | | Mistral AI SAS | FR | Post-meeting AI summary | Meeting transcript (once, at meeting end) | EU | DPA; zero data retention; no training on paid API data | | DeepL SE | DE | Document translation | Contents of documents users submit for translation | EU (Cologne) | DPA; no retention for training | | Stripe Payments Europe Ltd / Stripe Inc. | IE / US | Billing | Name, email, billing address, payment method (card never touches InterMIND) | EU / US | DPA + SCCs; PCI-DSS | | PostHog | US (EU Cloud) | Product analytics, session replay | Usage events, console/replay (consent-gated, opt-out by default; meeting speech scrubbed from logs) | EU (eu.posthog.com) | DPA; EU residency | | Functional Software Inc. (Sentry) | US (EU region) | Error monitoring | Error traces, breadcrumbs (no meeting content) | EU (de.sentry.io) | DPA; EU residency | | Resend | US | Transactional email (sign-in codes, notifications) | Recipient email, message content | Ireland (eu-west-1) | DPA + SCCs | | Upstash Inc. | US | OIDC session cache (Redis) | Encrypted session/token cache | Frankfurt (fra1) | DPA + SCCs; encrypted payloads | | Pipedrive OÜ | EE | Sales CRM (contact / partner forms) | Lead name, email, company, message | EU | DPA | | Google LLC / Microsoft Corp. | US | OAuth sign-in only | Name, email, profile photo (standard OAuth scopes) | US | DPF / SCCs | For questions about this list or to request a copy of the relevant transfer safeguards, contact . # Terms of Service **Effective date: 17 June 2026** ## 1. Agreement These Terms of Service ("Terms") are a contract between you and **Golden Fish Corporate Services Provider LLC** ("Golden Fish CSP LLC", "we", "us"), City Avenue Building, Office 405-070, Port Saeed, Dubai, United Arab Emirates — the operator of InterMIND, the meeting platform with real-time speech translation available at intermind.com and through the InterMIND mobile apps (the "Service"). By creating an account or using the Service you accept these Terms. If you use the Service on behalf of an organization, you confirm you have authority to bind it, and "you" includes that organization. **Publisher and Intellectual Property Owner:** MindMeeting OÜ (Estonia), Juhkentali 8, Tallinn 10132, Estonia. **Service Operator and Contracting Entity:** Golden Fish Corporate Services Provider LLC (United Arab Emirates). The InterMIND mobile apps are published on the Apple App Store and Google Play by MindMeeting OÜ on behalf of, and under license from, Golden Fish CSP LLC; your contract for the Service is with Golden Fish CSP LLC. ## 2. The Service InterMIND provides online meetings with real-time speech transcription and translation, meeting recordings, persistent and in-meeting chat with message translation, document translation, and AI-generated meeting summaries. Feature availability and limits depend on your plan (see §4). **Translation accuracy.** Transcription and translation are produced by automated speech recognition and machine translation. They are provided for communication convenience and may contain errors. Do not rely on them as the sole basis for legal, medical, financial, or other consequential decisions, and verify accuracy before relying on any transcript or translation. ## 3. Accounts You sign in with a one-time email code or a third-party sign-in provider (e.g. Google or Microsoft). Keep access to your email account secure — anyone who controls it can access your InterMIND account. You are responsible for activity under your account. Team workspaces have administrators who can manage members, content, and billing for the workspace; if you join a team workspace, its administrators control that workspace's data. Guests may join meetings without an account; anonymous guest sessions are temporary and are purged automatically (see the Privacy Policy). ## 4. Plans, billing, trials - **Plans.** The Service offers a free plan and paid subscription plans. Current plans, limits (such as monthly translation minutes, participants, and storage), and prices are listed on the pricing page at intermind.com/pricing, which forms part of these Terms. - **Payment.** Paid plans are billed by subscription (monthly or yearly) through our payment provider, Stripe. Your card details are handled by Stripe and never touch our systems. - **Trials.** Paid plans may include a free trial; its length is shown at checkout. Unless you cancel before the trial ends, the subscription begins automatically. - **Renewal and cancellation.** Subscriptions renew automatically until cancelled. You can cancel at any time via the billing portal in Settings; cancellation takes effect at the end of the current billing period. Except as required by applicable law, fees are non-refundable and we do not provide refunds or credits for partial billing periods. Purchases made through our reseller Paddle are governed by [Paddle's Refund Policy](https://www.paddle.com/legal/refund-policy){rel=""nofollow""}; refund requests for such purchases are handled by Paddle. If you are a consumer in the EU/EEA or UK, you have a statutory 14-day right of withdrawal for digital services. Nothing in these Terms limits mandatory consumer rights under the law of your country of residence. - **Plan changes and prices.** We may change plan prices or limits with at least 30 days' prior notice; changes apply from your next billing period. - **Usage limits.** Plan limits (e.g. translation minutes per month) are enforced by the Service; when a limit is reached, the corresponding feature pauses until the next period or an upgrade. ## 5. Your content - **Ownership.** You retain all rights to the content you create or upload — meetings, recordings, transcripts, chat messages, documents ("Content"). We claim no ownership. - **License to operate.** You grant us the license needed to run the Service: to host, store, transmit, transcribe, translate, summarize, and display your Content to you and the people you share it with. This license ends when the Content is deleted. We do not use your Content to train AI models. - **Your responsibilities.** You are responsible for your Content and for your use of recording and transcription features in compliance with applicable law — some jurisdictions require the consent of all participants before a conversation is recorded or transcribed. The Service makes recording visible to participants, but obtaining any legally required consent is your responsibility. - **Deletion and export.** You can delete individual Content, or your entire account in Settings (permanent, immediate, includes storage and subscription cancellation), and export your account data as a ZIP archive. See the Privacy Policy for retention details. ## 6. Acceptable use You agree not to: use the Service for unlawful purposes; infringe others' rights (including recording people without required consent); upload malware or attempt to breach, probe, or overload the Service; resell or white-label the Service without an agreement with us; or interfere with other users. We may suspend or terminate accounts that violate these Terms. Where a violation is capable of being cured, we will give you notice and a reasonable opportunity (normally 14 days) to cure it before suspending or terminating your account. We may suspend access immediately, with notice as soon as reasonably practicable, where the violation is unlawful, poses a security risk, risks harm to others or to the Service, or where required by law. ## 7. Privacy and data protection Our Privacy Policy (/privacy) describes what we process and why; the subprocessor list is published at [/legal/subprocessors](https://intermind.com/legal/subprocessors). For organizations that need one, we offer a [Data Processing Addendum](https://intermind.com/legal/dpa) — request it at . ## 8. Intellectual property The Service, including its software, design, and branding, is protected by intellectual property rights owned by MindMeeting OÜ (Estonia) and/or its licensors and used by Golden Fish CSP LLC. These Terms grant you only the right to use the Service; no other rights are transferred. If you send us feedback or suggestions, we may use them without obligation to you. ## 9. Third-party services Sign-in via Google or Microsoft, and payments via Stripe, are subject to those providers' own terms. We are not responsible for third-party services. ## 10. Disclaimers The Service is provided "as is" and "as available". To the maximum extent permitted by law, we disclaim all warranties, express or implied, including fitness for a particular purpose and non-infringement. We do not warrant uninterrupted or error-free operation, or the accuracy of transcriptions and translations (§2). For paid subscriptions, our uptime commitment and its exclusive service-credit remedy are described in the [Service Level Agreement](https://intermind.com/legal/sla), which forms part of these Terms; otherwise the Service is provided without service level commitments. ## 11. Limitation of liability To the maximum extent permitted by applicable law: (a) neither party is liable for indirect, incidental, special, consequential, or punitive damages, or for lost profits, revenue, data, or goodwill, even if advised of the possibility; (b) our total aggregate liability arising out of or relating to the Service or these Terms is limited to the greater of the fees you paid us for the Service in the 12 months before the event giving rise to the claim, or USD 100; and (c) these limits do not apply to liability that cannot be excluded or limited under applicable law, including for death or personal injury caused by negligence, for fraud or fraudulent misrepresentation, or for willful misconduct or gross negligence. Where mandatory consumer-protection law gives you greater rights, those rights prevail. These limitations reflect a reasonable allocation of risk and are an essential basis of the bargain. ## 12. Termination You may stop using the Service and delete your account at any time. We may suspend or terminate your access for material breach of these Terms (with notice and a reasonable cure period where the breach is curable, as set out in §6), or discontinue the Service with reasonable prior notice — in which case prepaid fees for the unused period will be refunded pro-rata. Sections that by their nature survive termination (e.g. §§8, 10, 11, 13) do so. ## 13. Governing law and disputes These Terms are governed by the laws of the Dubai International Financial Centre (DIFC). The Courts of the DIFC have exclusive jurisdiction over any dispute arising out of or in connection with these Terms or the Service, and the parties submit to the jurisdiction of the DIFC Courts. This choice of governing law and forum does not displace mandatory data-protection law applicable to the processing described in the Privacy Policy. If you use the Service as a consumer, nothing in this section deprives you of the protection of mandatory provisions of the law of your country of residence, and you may bring proceedings in the courts of that country where the applicable law so requires. ## 14. Changes to these Terms We may update these Terms; we will post the new version on this page and update the effective date, and for material changes we will notify you in the product or by email before they take effect. Continued use after the effective date constitutes acceptance. ## 15. Contact Golden Fish Corporate Services Provider LLC — City Avenue Building, Office 405-070, Port Saeed, Dubai, United Arab Emirates. Email: # AirPods Live Translation: what it does well, and where a meeting needs more Apple's Live Translation on AirPods is the feature that made "translation in your ear" mainstream. It's genuinely good — for the job it was built for. This explainer covers what it does, what it needs, and the exact line where a real meeting asks for something Apple's feature isn't shaped to give. --- ## What it is You wear compatible AirPods, someone speaks to you in another language, and you hear a translation in your ear — while the feature lowers their original voice underneath. It's built for the traveler-and-local moment: a shop, a taxi, a conversation with one person in front of you. Point-to-point, low-friction, and it feels natural because there's no screen between you. ## What it needs By Apple's own documentation, as of August 2026: - **Hardware:** AirPods Pro 3, AirPods Pro 2, AirPods 4 with Active Noise Cancellation, or AirPods Max 2, on the latest firmware. - **A recent iPhone:** an Apple Intelligence-capable iPhone (iPhone 15 Pro and newer), running iOS 26 or later. - **Languages:** English, French, German, Spanish, Portuguese, Italian, Chinese (Mandarin, Simplified and Traditional), Japanese, and Korean. - **Region:** it launched in the US first; EU availability arrived with iOS 26.2 after extra engineering for the Digital Markets Act. For a phrase with one person, none of that is a problem. The limits only start to bite when the situation changes shape. --- ## Where a meeting is a different job Look at the model, not the polish. AirPods Live Translation is **one-to-one, on your device.** It translates the person in front of you, into your ear, using your phone and your earbuds. That's exactly right for face-to-face — and exactly the wrong shape for a group call: - **It's per-device, not per-room.** For the translation to be mutual, the other person needs their own compatible AirPods and their own recent iPhone. In a call with five people, that assumption breaks for almost everyone. - **It's one conversation partner at a time.** A meeting has several people talking, interrupting, thinking out loud. Earbuds translating "the person in front of you" don't map onto a room. - **No shared hardware in a video call.** You can't hand a Teams or Zoom call a pair of AirPods. Meeting translation has to live in the *call*, not in anyone's ears. None of this is a knock on Apple's feature. It's a category boundary: **AirPods translate a face-to-face exchange; a meeting needs the whole room translated for everyone.** --- ## What "the whole room" means A meeting translator turns the *call itself* multilingual: every participant speaks their own language, and every participant hears everyone else in **their** language, live, for the whole meeting — no matching earbuds, no shared phone, no "you both need the same gear." That's [real-time meeting translation](https://intermind.com/blog/real-time-meeting-translation), and it's an architecture, not a better pair of earbuds. Where Apple translates one person into your ear on your hardware, meeting translation translates the room for everyone, on whatever they're already using to join the call. --- ## Where InterMIND fits InterMIND does the group-call job that earbuds structurally can't: - **The whole call is multilingual.** Each participant hears the meeting in their own picked language, simultaneously — no shared device, no matching hardware. [24 languages live on voice](https://intermind.com/features/realtime-translation), chat and shared notes. - **Per-listener audio, sub-second.** Five people, five languages, one room. (Under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Webinars and conferences** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — a PDF or DOCX in the meeting, each viewer in their language. - **Quality you can audit** — per-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark). Keep the AirPods for the taxi and the shop; they're great at it. When it's a *meeting* — several people, several languages, no shared hardware — that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and hear per-listener translation across a whole room. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month quality on real traffic, methodology included. - The category, in full: [*Real-time meeting translation*](https://intermind.com/blog/real-time-meeting-translation). — The Mind.com Team --- *Sources: [Apple — Use Live Translation with your AirPods](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Apple — Live Translation on AirPods expands to the EU](https://www.apple.com/ie/newsroom/2025/11/live-translation-on-airpods-expands-to-the-eu/){rel=""nofollow""}. Apple expands device and language support over time; check Apple's support pages for the current state. All facts checked August 2026.* # Speak & Translate alternatives (2026): what to pick when one phrase isn't enough Search for the "Speak and Translate" voice translator and you find something surprising: it isn't one app but a name carried by **dozens of different apps** across the App Store and Google Play, from different developers. And for all the different icons, they belong to one category: the **phone phrase translator**. Tap the mic, say a sentence, hear the translation read back. That category is genuinely good at its job. The trouble starts when the job outgrows it: a business call, a meeting with a foreign supplier, an online lesson, a family spread across two countries. This guide separates things honestly: what these apps do well, where the whole category stops by design rather than by flaw, and the right alternative **per job** — not one "best alternative" for everyone. --- ## What "Speak and Translate" apps do well - **A quick spoken phrase**: a menu, a taxi, a question in the street. Tap, speak, play it to the other person. - **The camera**: photograph a sign or a menu and translate it. - **Offline** (in many of them): a travel mode with no data needed. If those are your jobs, don't look for an alternative — keep a well-rated one on your phone for travel. The fair comparison starts when the job itself changes. --- ## Where the whole category stops — by design The same three limits sit in every app of this category, whatever the developer: 1. **The relay.** The mechanic is always: you speak → it waits → it translates → it reads aloud → the other person answers into the same phone → it waits → it translates. For two sentences, fine. For a real discussion, you're taking turns operating a translator instead of talking — and the conversation runs twice as long. 2. **One device for two people.** The phone passes back and forth or sits between you. Add a third person and the arrangement collapses — there is no concept of a "room" at all. 3. **No calls, no meetings.** These apps can't join your call with a client abroad, or your team meeting on screen. They were built for the street, not for the working table. None of that is a flaw — it's the category's boundary. It becomes a problem only when we ask a phrase tool to run a conversation. --- ## The right tool for each job | Job | Right tool | | ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------- | | A spoken phrase while traveling | A "Speak and Translate"-style app — any well-rated one | | A face-to-face exchange, both people equipped | Translator earbuds ([the three kinds of voice translator](https://intermind.com/blog/mutarjim-sawti)) | | A paragraph, an email | Text translators: Google Translate and its peers | | A whole document (PDF, DOCX) | A file translator — InterMIND, for one, [translates files in the meeting across 30 languages](https://intermind.com/features/realtime-translation) | | **A voice conversation — a call, a meeting, several people** | **Live meeting translation** — everyone speaks and hears their own language | The first row and the last are the two that get confused in practice. The ladder is always the same: phrase → text → document → **conversation**. Each rung has its own tool, and the last one is the least known — so it gets its own section. --- ## The rung most people don't know: the full conversation The real need behind many voice-translator searches isn't a phrase in the street — it's a **conversation**: a business meeting with a foreign party, a job interview in a second language, a family call across two countries. That job has its own category — live meeting translation — where the meeting itself is multilingual: - You speak Arabic; the other side hears you in English, French, or Turkish **in under a second** — and answers in their language while you hear Arabic. **24 languages** on voice, chat, and shared notes, in any mix, with no forced English anchor. - **Per-listener audio**: five participants, five languages, one meeting — no phone changing hands. - **Free meetings up to 60 minutes** — enough for a real meeting, not a token trial; guests join from one browser link, no signup. - **Web, desktop, iOS and Android** — on the phone too, but as a full meeting, not as a device two people take turns with. - **Quality published, not promised**: per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark) — see how your pair stands before you rely on us. And if what you like about phrase apps is that they *show rather than promise*: [`/demo`](https://intermind.com/demo) does the same with voice — speak, and hear the live translation yourself. --- ## The verdict Don't look for one "better alternative" — name your job. A phrase while traveling? Keep a "Speak and Translate" app on your phone; the category does its work. Text or a document? Each has its tool. But if the job that outgrew your app is the **conversation** — the meeting, the call, the lesson — the alternative isn't a better phrase app but a different category entirely: [live meeting translation](https://intermind.com/blog/real-time-meeting-translation). Phrase, text, document, conversation. Four jobs, four tools — and now you know which is which. --- ## FAQ **What is the best alternative to Speak and Translate?** Depends on which job outgrew it. Another phrase in the street: any well-rated app of the same category — there is no meaningful "best" among dozens of near-identical ones. A face-to-face exchange where both people are equipped: translator earbuds. Text or a document: a text/file translator. A call, an online lesson, or a meeting with several people: that's a different category — live meeting translation, where every participant speaks and hears their own language. **Can a Speak and Translate app translate a phone call or an online meeting?** The category is built for in-person, turn-taking use: even Google's own documentation for its Conversation mode describes taking turns speaking into one device. None of these apps joins a call or a meeting as a participant. If the conversation happens on a screen — a client call, a team meeting, an interview — you need a tool that lives in the meeting itself: per-listener translated audio, translated chat, translated notes. **What's the difference between a phrase translator and live meeting translation?** A phrase translator is a relay: speak, wait, play the translation, pass the phone. Live meeting translation is simultaneous and per-participant: five people, five languages, each speaking and hearing their own, at once. InterMIND does this across 24 languages on voice, chat, and shared notes — with per-language-pair quality published at [/benchmark](https://intermind.com/benchmark), and [free meetings up to 60 minutes](https://intermind.com/pricing) to test it on a real conversation. --- *Sources: [Google Translate — hear speech-to-speech translations (Conversation mode)](https://support.google.com/translate/answer/6142474){rel=""nofollow""} — the turn-taking mechanic of phrase translation, documented by the category's largest example — checked August 2026. InterMIND facts: [per-surface language breakdown](https://intermind.com/blog/how-many-languages-do-you-support), [live benchmark](https://intermind.com/benchmark).* — The Mind.com Team # Best AI translation tools for conferences, events, and meetings (2026): an honest comparison If you typed *"best AI translation tools for conferences,"* *"real-time interpretation software,"* or *"which tools support multilingual simultaneous interpretation,"* you've probably noticed the listicles all blur together. Every tool claims "real-time," "AI-powered," and "multilingual," and most of them mean genuinely different things by it. One subtitles a webinar. One streams a human interpreter's audio to attendees' phones. One is a $300 earbud. These are not the same product, and picking the wrong category is the most expensive mistake here. But there's a deeper split the listicles miss entirely — and it's the one that actually matters once the call is over. **Almost every tool on every list translates one thing: the spoken moment.** Someone talks, you hear it in your language, and that's the whole product. The instant the words stop, the translation stops. The chat is still in the speaker's language. So are the shared notes. So is the contract someone dropped in. So is the follow-up. So is the support thread when something breaks. A meeting is not just the audio. It's the messages, the notes, the documents, the notifications, the help you read mid-call, the conversation with support afterward, and the record you keep. **The honest question isn't "how good is the voice" — it's "how much of the meeting does it actually translate?"** That's the axis this guide is built on, and it's where the field separates hard. So this guide does the part the listicles skip: it names the three *jobs* people mean, gives you the questions that tell them apart — including the surface-coverage one nobody asks — and *then* compares named tools. We make one of them ([InterMIND](https://intermind.com){rel=""nofollow""}), and we'll say where it fits and where it doesn't — but the questions below are vendor-neutral and work on any tool, including ours. > This is the comparison companion to our foundational guide, [*Real-time meeting translation: how it works, and how to evaluate one*](https://intermind.com/blog/real-time-meeting-translation). If you want the deeper "how does this work under the hood" version, start there. New to interpretation as a category? Start with [*interpretation vs translation and the types of interpretation*](https://intermind.com/blog/interpretation-vs-translation) for the definitions, then [*Simultaneous interpretation: booth, RSI, or AI*](https://intermind.com/blog/simultaneous-interpretation-guide) for the modes and the costs. --- ## First: the three jobs hiding under one search Almost every tool in this space does one of three jobs well. Naming them is half the decision. 1. **Simultaneous interpretation delivery** — get *audio* (a human interpreter's, or a machine's) to a room or to attendees' devices, in real time, often one-directional (a stage to an audience). Think large events, parliaments, webinars. Tools: Interprefy, KUDO, Boostlingo, Akouo, Verspeak. 2. **Conversational meeting translation** — a working meeting where *several* people each speak, type, read, and listen in their *own* language, both directions, at once. Think a sales call, a standup, a partner negotiation. This is the hardest job and the smallest category. 3. **Caption / transcript translation** — translate the *text* of what's said: live subtitles, post-call transcripts, AI notes. Think Zoom/Teams/Meet captions, Otter, AI notetakers. A tool can handle job 1 and be useless for job 2. A captioning add-on (job 3) is not interpretation at all — it's reading, not hearing. Decide your job first. --- ## The questions that actually separate tools Run any candidate through these. They cut through the marketing faster than any feature matrix. The last one is the one no listicle asks — and it's usually the deciding one. ### 1. One speaker, or everyone at once? Event tools optimize for **one source → many listeners** (a speaker on stage, an audience listening). Meeting tools have to handle **N people each speaking and listening in different languages, simultaneously, both directions.** If your use case is a four-person call where everyone talks, a one-directional event platform will feel wrong no matter how good its audio is. ### 2. Do listeners *hear* it, or *read* it? Captions (job 3) are a reading experience — subtitles, not audio. They're great for accessibility and webinars where one person presents. They're poor for a discussion, because you can't read four people's subtitles and still react to each other. If you need spoken translation, rule out anything whose "translation" is text-only. ### 3. Machine, or human-in-the-loop? KUDO, Interprefy, and Boostlingo are built around **routing human interpreters** (with AI as an option). That's the right answer for a UN-grade session where a mistranslation is a liability. It's the wrong cost structure for a Tuesday standup. AI-only tools (Wordly, DeepL Voice, InterMIND) trade certified-human accuracy for instant, per-meeting, no-booking availability. Know which trade you're making. ### 4. Whose voice comes out? Most machine tools replace every speaker with **one generic synthetic narrator** — eight people, one robot voice. A few keep the *speaker's own voice* via zero-shot voice synthesis, so a listener hears the translation in a voice recognizably the speaker's. In a real conversation that's the difference between a discussion and a transcript read aloud. (We wrote up why this is hard and how it works in [*Speak in your own voice — in a language you don't speak*](https://intermind.com/blog/own-voice-translation).) ### 5. How much of the meeting does it actually translate? *(the one nobody asks)* This is the question that should be first, not last. Voice is the demo; it's not the meeting. A real working session generates **a whole communication surface** around the audio: - **The chat** — links, decisions, side-questions typed while someone else talks. - **The shared notes** — the agenda, the action items, the doc everyone edits live. - **The documents** — the contract, the deck, the spreadsheet dropped in for review. - **The in-product help** — what you read when you can't find a setting mid-call. - **The support conversation** — what happens, days later, when something breaks. - **The after-record** — the summary, the digest, the transcript you actually keep and forward. Most tools translate *the audio and nothing else*. Everyone hears the call, then opens a chat log, a notes pane, and a follow-up email all still in a language half the room can't read. The translation evaporated the moment the talking stopped. Ask any candidate plainly: **after the audio, what else comes back in my language?** If the answer is "captions," you have a voice tool with a transcript bolted on — not a translated meeting. This single question reorders most shortlists. ### 6. What happens to the audio — and where does it run? For anything regulated — legal, medical, HR, finance — ask plainly: **is the call recorded or the voice stored, and does any of it leave your jurisdiction?** Some tools retain audio for model training; some store a voiceprint to do voice cloning; some send your meeting content to a US-hosted model the moment they generate a summary. This is a procurement gate, not a nice-to-have. (Our own answer: the live session retains nothing, and nothing derived from a meeting touches a US-domiciled model — see [*the GDPR audit*](https://intermind.com/blog/gdpr-audit-what-we-closed) and [*where one meeting actually runs*](https://intermind.com/blog/where-one-intermind-meeting-actually-runs).) --- ## The contenders, sorted by job The tools below are the names that come up most for conference and meeting translation in 2026. We've grouped them by the three jobs above so you compare like with like. ### For large events & simultaneous interpretation delivery (job 1) - **Interprefy** — remote-simultaneous-interpretation (RSI) platform. Its documentation states [190+ languages and 6,000+ language combinations](https://www.interprefy.com/){rel=""nofollow""} with human interpreters, and [80 languages for its AI speech translation](https://knowledge.interprefy.com/what-languages-can-interprefy-ai-translate-from-and-to){rel=""nofollow""} (speech and captions). ([InterMIND vs Interprefy, feature by feature.](https://intermind.com/compare/interprefy)) - **KUDO** — RSI platform with an AI option: the [KUDO AI Speech Translator](https://kudo.ai/solutions/kudo-ai-speech-translator/){rel=""nofollow""} page states real-time audio and captions in 60+ languages, alongside its human-interpreter delivery. - **Boostlingo** — interpreter network and delivery platform: [its site](https://boostlingo.com/){rel=""nofollow""} states 10K qualified interpreters, 275+ languages 24/7, on-demand phone/video interpreting (OPI/VRI), plus AI captions ("Boostlingo AI Pro"). - **Akouo / Verspeak** — deliver interpreter audio to attendees' own phones over the web, for in-room and hybrid events without receiver hardware. **Pick one of these if:** you're running a conference, webinar, or formal multilingual session with an audience — especially if you need or already use human interpreters. If your "event" is an internal all-hands or an investor call — an audience, but everyone still needs to follow, and often ask, in their own language — that's a job-2 problem wearing a job-1 badge. See how InterMIND handles [global town halls](https://intermind.com/use-case/town-halls) and [investor-relations calls](https://intermind.com/use-case/ir-calls). And if the event itself is the job — a webinar, a conference session, a training — the InterMIND version of job 1 is at [events & webinars](https://intermind.com/use-case/events); worship services, a format of their own, are at [church translation](https://intermind.com/use-case/church). ### For everyday multilingual meetings (job 2) This is the category where question 5 — *how much of the meeting?* — does the most work, because these tools look alike in a voice demo and diverge sharply once the call has chat, notes, and documents in it. - **Wordly** — AI-only, real-time translation for meetings and events. Its documentation states [60+ languages and 3,000+ language pairs](https://www.wordly.ai/translator-languages){rel=""nofollow""}, with attendees choosing text, audio, or both; [its FAQ](https://www.wordly.ai/faq){rel=""nofollow""} also states translated transcripts and session summaries from the Wordly Portal. Chat, shared-notes, and document translation are not stated in its public documentation (checked August 2026). ([InterMIND vs Wordly, feature by feature.](https://intermind.com/compare/wordly)) - **DeepL Voice** — DeepL's real-time speech translation. [The product page](https://www.deepl.com/en/products/voice){rel=""nofollow""} states live captions in 40+ languages inside Microsoft Teams, Zoom Meetings, and Google Meet — and lists voice-to-voice support as "coming soon" (checked August 2026). Document translation is a separate DeepL product, not part of the meeting flow. - **InterMIND** — what we build. AI-only, **conversational**meeting translation where the whole meeting — not just the audio — comes back in each participant's language, both directions, at once. The point of difference is surface coverage: - **Voice** — 24 languages, per-viewer translated audio with sub-second latency, in the **speaker's own voice** via a zero-shot ASR → MT → TTS cascade, not a single robot narrator. ([The real-time translation feature](https://intermind.com/features/realtime-translation); [how the pipeline works](https://intermind.com/blog/inside-the-translation-pipelines).) - **Chat & shared notes** — every message and every keystroke in the notes pane translated live, per viewer, in the same 24 languages, with per-language edit diffs. - **Documents** — drop a PDF, DOCX, PPTX, or XLSX into the chat and each participant gets it back in their language with formatting intact — **30 languages** via the DeepL Document API. (The honest per-surface language breakdown is [here](https://intermind.com/blog/how-many-languages-do-you-support).) - **In-product help & support, in your language** — the help assistant answers in the language you write in, and customer support replies are drafted in the client's language. The conversation around the product is multilingual too, not just the call. - **The after-record** — the post-meeting AI summary/digest is generated for you, and (like everything above) the meeting content stays on EU-hosted models with zero data retention — **no meeting data reaches a US-domiciled model.** - **Quality is published, not claimed** — the production voice pipeline is scored monthly against FLORES-200 with the full per-language-pair distribution at [/benchmark](https://intermind.com/benchmark), and you can [run the live demo](https://intermind.com/demo) on your own audio. **Pick one of these if:** your "conference" is really a working meeting — a call where multiple people need to talk, type, read, and decide *with* each other across languages, and where the chat, notes, documents, and follow-up need to be readable too, not just the audio. ### For captions, transcripts & notes (job 3) - **Zoom** — translated captions in 36 languages, each participant picking their own ([Business Plus/Enterprise, or a paid add-on](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0060844){rel=""nofollow""}); the newer [Voice translator](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0084896){rel=""nofollow""} renders translated *audio* per participant (AI Companion, US-cluster accounts, desktop app 7.0+); and [audio channels for up to 20 human interpreters](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0064768){rel=""nofollow""} you source yourself. Full breakdown: [Zoom live translation](https://intermind.com/blog/zoom-live-translation). - **Microsoft Teams** — [translated captions in 31 listed languages](https://support.microsoft.com/en-us/office/use-live-captions-in-microsoft-teams-meetings-4be2d304-f675-4b57-8347-cbd000a21260){rel=""nofollow""} (organizer needs Teams Premium or Copilot); the [Interpreter agent](https://support.microsoft.com/en-us/office/interpreter-in-microsoft-teams-meetings-and-calls-c7efe2bb-535d-42ab-a5c4-d2d91619b46d){rel=""nofollow""} translates spoken audio into spoken audio, optionally simulating the speaker's voice (listener needs a Copilot license); and [pre-configured channels for up to 16 language pairs of human interpreters](https://support.microsoft.com/en-us/office/use-language-interpretation-in-microsoft-teams-meetings-b9fdde0f-1896-48ba-8540-efc99f5f4b2e){rel=""nofollow""}. Full breakdown: [Teams live translation](https://intermind.com/blog/teams-live-translation). - **Google Meet** — [translated captions](https://support.google.com/meet/answer/10964115){rel=""nofollow""} on Business Standard+ Workspace editions, and [Gemini speech translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""} that translates the spoken audio "in a voice like yours" (Google AI Pro/Ultra or qualifying Workspace plans). Full breakdown: [Google Meet live translation](https://intermind.com/blog/google-meet-live-translation). - **Otter, and AI notetakers generally** — transcribe and summarize, sometimes translate the transcript afterwards. This is recording and notes, not live interpretation: nothing changes what participants hear during the call. We compared the category in [Otter alternatives](https://intermind.com/blog/otter-ai-alternatives) and [Fireflies alternatives](https://intermind.com/blog/fireflies-ai-alternatives), with per-tool language counts and sources. **Pick one of these if:** you mainly need a translated transcript or subtitles, and live two-way spoken translation isn't the requirement. ### A note on hardware (Timekettle et al.) Earbud translators show up in these searches, so for completeness: [Timekettle's W4 Pro product page](https://www.timekettle.co/products/timekettle-w4-pro-ai-interpreter-earbuds-new){rel=""nofollow""} states 52 languages and 106 accents online, for in-person conversation plus call/video translation through its companion app. It's a personal device: the translation happens in your ears, not in the meeting — other participants' experience is unchanged. A different category from the meeting platforms above; decide which of the three jobs you're hiring for first. --- ## The surface-coverage table Question 5 — *how much of the meeting comes back in your language?* — as one table. Every cell is what the vendor's own public documentation states, checked August 2026 (links in the sections above and in Sources below). "Not stated" means we found no claim for that surface in the vendor's public docs — check with the vendor before buying if it matters to you. | Tool | Translated audio (live) | Translated captions | Chat messages | Live shared notes | Documents in-meeting | Post-meeting record | | ------------------- | --------------------------------------------------------------------------- | ------------------------------------------------------ | ----------------------------------------------- | --------------------------- | -------------------------------------- | ------------------------------------------ | | **Interprefy** | Yes — human interpreters (190+ languages) or AI (80) | Yes (AI, 80 languages) | Not stated | Not stated | Not stated | Not stated | | **KUDO** | Yes — human interpreters or AI (60+ languages) | Yes (60+ languages) | Not stated | Not stated | Not stated | Not stated | | **Wordly** | Yes (AI; 60+ languages, 3,000+ pairs) | Yes | Not stated | Not stated | Not stated | Translated transcripts + session summaries | | **DeepL Voice** | "Coming soon" (per product page) | Yes (40+ languages, in Teams/Zoom/Meet) | Not stated | Not stated | Separate DeepL product | Not stated | | **Zoom** | Yes — Voice translator (AI Companion; US cluster, desktop 7.0+) | Yes (36 languages; Business Plus/Enterprise or add-on) | Manual per-message (38 languages) | Not stated | Not stated | Not stated | | **Microsoft Teams** | Yes — Interpreter agent, can simulate speaker's voice (Copilot license) | Yes (31 languages; Premium/Copilot) | Per-message, optional auto (inline translation) | Not stated | Not stated | Not stated | | **Google Meet** | Yes — Gemini speech translation, "in a voice like yours" (qualifying plans) | Yes (Business Standard+) | Not stated | Not stated | Not stated | Not stated | | **InterMIND** | Yes — 24 languages, speaker's own voice | Yes (24, per viewer) | Automatic per viewer (24 languages) | Yes — live, per viewer (24) | Yes — PDF/DOCX/PPTX/XLSX, 30 languages | Summary + digest, in your language | Two honest notes on our own row: the per-surface language counts differ (voice/chat/notes 24, documents 30, website 17) — [the full breakdown and why](https://intermind.com/blog/how-many-languages-do-you-support); and voice quality is published monthly against FLORES-200 at [/benchmark](https://intermind.com/benchmark) rather than claimed here. --- ## A quick decision shortcut - **Conference with an audience + you want human interpreters** → Interprefy / KUDO / Boostlingo. - **Working meeting, several people, everyone talks, both directions, AI-only** → Wordly / DeepL Voice / [InterMIND](https://intermind.com){rel=""nofollow""} — and here the differentiators are *own-voice* output, *whole-surface* coverage (chat, notes, documents, support, the after-record — not just audio), and *published* quality numbers. Test those specifically. - **You just need translated captions or a translated transcript** → your existing Zoom/Teams/Meet, or an AI notetaker. The honest meta-point: "best AI translation tool for conferences" has no single winner because "conference" hides three different jobs — and within the meeting job, most tools translate the spoken moment and stop. Name your job, then ask how much of the meeting actually comes back in your language. The shortlist writes itself. --- ## FAQ **What is the best real-time interpretation software?** Depends on which of the three jobs you're hiring for. Delivering interpretation to an event audience: Interprefy (80 AI languages, 190+ with human interpreters), KUDO (60+), or Wordly (60+) — all three document audio plus captions. A working meeting where everyone speaks, types, and reads in their own language, both directions: that's the conversational job — InterMIND covers voice, chat, notes (24 languages), and documents (30) per participant. Translated subtitles on a call you already run: Zoom (36 caption languages), Teams (31), or Meet handle it natively, with plan gates listed above. **Which tools support multilingual simultaneous interpretation?** With human interpreters: Interprefy (6,000+ language combinations), KUDO, Boostlingo (275+ languages on-demand), plus Zoom (audio channels for up to 20 interpreters you hire) and Teams (16 pre-configured language pairs). AI-only simultaneous: Interprefy AI (80 languages), KUDO AI (60+), Wordly (60+), InterMIND (24, in the speaker's own voice, both directions). **What are the best AI translation tools for conferences and meetings in 2026?** For conferences with an audience: Wordly, Interprefy AI, and KUDO AI all document 60–80 language live translation as audio and captions. For multilingual working meetings: InterMIND translates the whole meeting per participant — audio in the speaker's voice, chat, live notes, and dropped documents. For captions inside your existing platform: Zoom, Teams, and Meet each ship translated captions natively (gates in the table above), and DeepL Voice adds 40+ caption languages inside all three. **Is there an AI meeting assistant that transcribes and translates?** Notetakers (Otter, Fireflies, and the category we compared [here](https://intermind.com/blog/fireflies-ai-alternatives)) transcribe live and can translate the transcript afterwards — nothing changes what participants hear during the call. Wordly documents translated transcripts and summaries after its sessions. InterMIND does both jobs in one: live translated audio/chat/notes during the meeting, then a summary and digest generated in each participant's language. **Do Zoom, Microsoft Teams, or Google Meet translate meetings by themselves?** Yes, with plan gates. Zoom: translated captions in 36 languages (Business Plus/Enterprise or a paid add-on) and a Voice translator that renders translated audio (AI Companion, US-cluster accounts). Teams: translated captions in 31 languages (Premium/Copilot) and the Interpreter agent for spoken-audio translation (Copilot license). Meet: translated captions (Business Standard+) and Gemini speech translation on qualifying plans. Details and setup for each: [Zoom](https://intermind.com/blog/zoom-live-translation), [Teams](https://intermind.com/blog/teams-live-translation), [Meet](https://intermind.com/blog/google-meet-live-translation). --- ## See it for yourself We'd rather you test than take our word. For the meeting-translation job (job 2), the fastest way to judge any tool — ours included — is to put your own meeting through it: talk, then check whether the chat, the notes, and the doc came back in your language too. - **[Try the live demo](https://intermind.com/demo)** — runs InterMIND's production voice pipeline on your audio, in any of 24 languages. - **[Read the benchmark](https://intermind.com/benchmark)** — monthly FLORES-200 scores, full per-pair distribution, no cherry-picking. - **[How to evaluate any real-time translator](https://intermind.com/blog/real-time-meeting-translation)** — the vendor-neutral foundation behind this guide. --- *Sources: [Interprefy — AI translation languages](https://knowledge.interprefy.com/what-languages-can-interprefy-ai-translate-from-and-to){rel=""nofollow""}, [Interprefy — platform](https://www.interprefy.com/){rel=""nofollow""}, [KUDO — AI Speech Translator](https://kudo.ai/solutions/kudo-ai-speech-translator/){rel=""nofollow""}, [Boostlingo](https://boostlingo.com/){rel=""nofollow""}, [Wordly — translator languages](https://www.wordly.ai/translator-languages){rel=""nofollow""}, [Wordly — FAQ](https://www.wordly.ai/faq){rel=""nofollow""}, [DeepL — Voice](https://www.deepl.com/en/products/voice){rel=""nofollow""}, [Zoom — viewing captions in another language](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0060844){rel=""nofollow""}, [Zoom — Voice translator](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0084896){rel=""nofollow""}, [Zoom — Language Interpretation](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0064768){rel=""nofollow""}, [Zoom — translating messages in Chat](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0058113){rel=""nofollow""}, [Microsoft — live captions in Teams meetings](https://support.microsoft.com/en-us/office/use-live-captions-in-microsoft-teams-meetings-4be2d304-f675-4b57-8347-cbd000a21260){rel=""nofollow""}, [Microsoft — Interpreter in Teams](https://support.microsoft.com/en-us/office/interpreter-in-microsoft-teams-meetings-and-calls-c7efe2bb-535d-42ab-a5c4-d2d91619b46d){rel=""nofollow""}, [Microsoft — language interpretation in Teams](https://support.microsoft.com/en-us/office/use-language-interpretation-in-microsoft-teams-meetings-b9fdde0f-1896-48ba-8540-efc99f5f4b2e){rel=""nofollow""}, [Microsoft — inline message translation](https://learn.microsoft.com/en-us/microsoftteams/inline-message-translation-teams){rel=""nofollow""}, [Google Meet — translated captions](https://support.google.com/meet/answer/10964115){rel=""nofollow""}, [Google Meet — Speech Translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""}, [Timekettle — W4 Pro](https://www.timekettle.co/products/timekettle-w4-pro-ai-interpreter-earbuds-new){rel=""nofollow""}, checked August 2026. Vendors expand plans and language lists over time; check their pages for the current state. InterMIND facts: [per-surface language breakdown](https://intermind.com/blog/how-many-languages-do-you-support), [live benchmark](https://intermind.com/benchmark).* # The other bundle: languages around your communication, not apps around English Walk into any growing company from Jakarta to Bogotá and the software decision has usually already been made — not tool by tool, but as a bundle. Google Workspace or Microsoft 365: mail, calendar, documents, spreadsheets, storage, and a meeting app, one subscription, one admin console. The bundle *is* the pitch. Nobody buys Meet or Teams on its merits; they arrive in the box. It's a genuinely good deal on its own terms. But look at what sits at the center of the box. Every app in the bundle exists to serve communication — mail is communication, documents get written to be discussed, meetings are where the rest converges. And that communication layer, in both bundles, quietly assumes something about your team: **that it operates in one language.** In practice, in English. ## What the bundle assumes This isn't a rhetorical jab — it's documented product design, and we track it closely because it's our market: - **Google Meet's speech translation runs between English and five European languages, one pair per meeting** — every documented pair includes English. A meeting between a Portuguese speaker and an Indonesian speaker isn't in the design space (checked August 2026). - **Microsoft Teams gates translated captions behind the organizer's Premium or Copilot license, and its AI Interpreter behind a per-listener Copilot license, metered at 20 hours per user per month** — ten languages, scheduled meetings only. Translation is an enterprise add-on, priced per seat, on top of the bundle you already bought (checked August 2026). - **Nothing translated survives either platform's meetings.** Teams documents that captions aren't saved and recordings keep original audio only; Google documents that speech translation is unavailable in recordings and no audio is saved. The record of your meeting — the thing your team rereads, cites and onboards from — speaks whatever language the meeting was held in. Read the three facts together and the shape of the bundle is visible: the apps are many, and the communication layer underneath them defaults to one language — with translation documented as a paid, per-seat layer on top of the subscription you already bought. The bundle is apps around English-default communication. If your team happens to think in English, you'll never notice. If it thinks in Spanish, Portuguese, Turkish, Indonesian, Vietnamese — the bundle works, but *you* do the translating, in your head, every day, at whatever it costs you in precision and standing. We've written about what that costs: [the false fluency trap](https://intermind.com/blog/false-fluency-trap). ## Invert it Now run the bundling logic in the other direction. Keep communication at the center — meetings, chat, notes, documents, the history they leave — and instead of bundling *more apps* around it, bundle *languages* around it: - The meeting, live, in [every participant's own language](https://intermind.com/features/realtime-translation) — 24 languages for voice, chat and shared notes, each participant choosing theirs independently. - The [chat and its history](https://intermind.com/features/multilingual-chat), readable in each member's language — not just the messages arriving now, but the record. - [Documents](https://intermind.com/features/document-translation), translated in the same space — 30 languages. - The [recap after every meeting](https://intermind.com/features/recap), delivered to each signed-in participant in their own language. - And the whole thing persisting: the meeting becomes a channel, and [a week later it still exists in your language](https://intermind.com/blog/the-meeting-a-week-later) — which is the part no caption feature can give you, because captions die with the call. That's the entire product idea. Not an office suite with translation sprinkled on top; a communication space with languages as the load-bearing layer. One subscription on the host's side — [participants pay nothing](https://intermind.com/blog/free-for-participants). ## The sovereignty clause There's a second assumption inside the incumbent bundles, quieter than the language one: **the bundle decides where your communication lives.** Your meetings, transcripts and files reside in the vendor's cloud, under the vendor's jurisdiction, processed by the vendor's AI under terms you inherit rather than choose. For a café chain that may not matter. For a bank, a clinic network, a government supplier — or simply a company that has watched data-transfer rules shift every few years — it's a standing question mark over the whole box. We treat jurisdiction as part of the product: an [EU data path at every hop, zero data retention on the AI features, and user-controlled recording](https://intermind.com/features/security) — verifiable, not asserted. The same logic as publishing [monthly per-pair translation quality](https://intermind.com/benchmark) instead of adjectives: claims you can check are the only claims worth shipping. ## What we're not Honesty clause, because bundle comparisons invite overreach: **InterMIND does not replace your office suite.** We don't do mail, spreadsheets or slide decks, and we don't plan to. Keep Workspace or 365 for those — most of our customers do. The claim is narrower and sharper: the *communication* inside that stack — the meetings, the chat, the record they leave — deserves a layer built for the languages your team actually thinks in, rather than an English-default layer with translation sold back per seat. Two bundles, two philosophies. Theirs: more apps around one language. Ours: [every language around the communication](https://intermind.com) — in a jurisdiction you choose. ## See the difference in one meeting - **[Try the live demo](https://intermind.com/demo)** — no signup; hear a meeting where every listener gets their own language, simultaneously. - **[See the benchmark](https://intermind.com/benchmark)** — translation quality per language pair, published monthly on real traffic. - **[Compare with your current bundle](https://intermind.com/compare/google-workspace)** — feature-by-feature against [Google Workspace](https://intermind.com/compare/google-workspace) and [Microsoft Teams](https://intermind.com/compare/microsoft-teams), every competitor cell sourced. --- ## FAQ **Can InterMIND replace Google Workspace or Microsoft 365?** No, and it doesn't try to — there's no mail, spreadsheets or slides. InterMIND replaces the *communication layer*: multilingual meetings, chat, shared notes, document translation, and the persistent translated history they produce. It runs alongside your office suite. **What does "languages around communication" mean concretely?** That translation is the room's default behavior, not an add-on: voice in 24 languages per participant live, chat and notes translated per viewer including history, documents in 30 languages, and a recap after every meeting in each member's own language — with no per-listener licenses. The per-pipeline numbers differ on purpose; [here's why](https://intermind.com/blog/how-many-languages-do-you-support). **Isn't translation in Meet and Teams good enough for an international team?** Check the documented fences against your actual meetings: Meet translates speech between English and five languages, one pair per meeting; Teams interprets in ten languages for listeners holding Copilot licenses, 20 hours a month, scheduled meetings only — and neither keeps anything translated after the call (checked August 2026). If your meetings fit inside those fences, the bundled features may serve you well. If they don't, that's the gap this post is about. --- *Sources: [Google Meet — Learn about Speech Translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""}, [Google Workspace — plans and pricing](https://workspace.google.com/pricing){rel=""nofollow""}, [Microsoft — Interpreter in Teams meetings and calls](https://support.microsoft.com/en-us/office/interpreter-in-microsoft-teams-meetings-and-calls-c7efe2bb-535d-42ab-a5c4-d2d91619b46d){rel=""nofollow""}, [Microsoft Learn — Teams Premium licensing](https://learn.microsoft.com/en-us/microsoftteams/teams-add-on-licensing/licensing-enhance-teams){rel=""nofollow""}, [Microsoft 365 — business plans](https://www.microsoft.com/en-us/microsoft-365/business){rel=""nofollow""}, [Microsoft — Use live captions in Teams meetings](https://support.microsoft.com/en-us/office/use-live-captions-in-microsoft-teams-meetings-4be2d304-f675-4b57-8347-cbd000a21260){rel=""nofollow""}. Vendors change bundles, licensing and language lists over time; check their pages for the current state. All facts checked August 2026.* # The false-fluency trap: when "good-enough" English is worse than no English > **TL;DR — Knowing a language poorly is worse than not knowing it at all.** > > *Not* knowing it makes the gap visible: the room agrees on an interpreter and reorganizes around it. > *Poorly* knowing it hides the gap: fluent-sounding errors get minuted as agreed; the cost surfaces weeks later in the contract review. > > Aviation regulators acted on this mechanism in 2008. Medical research has been counting it since 2003. The boardroom hasn't. Every cross-border meeting today defaults to a shared language nobody fully owns — almost always English. The arrangement *feels* like it works. The rest of this post is the evidence that it works less than people think, drawn from three independent literatures (aviation safety, medical interpretation, international-business research), and an honest engagement with the strongest published counterargument (the *Foreign-Language Effect*). --- ## How many people in the room are operating in a second language? Of the roughly 1.5 billion English speakers in the world, about **400 million are native**; the remaining 1.1 billion learned English as a second or additional language (David Crystal, *English as a Global Language*, Cambridge University Press, 2nd ed., 2003; Ethnologue speaker estimates). In any cross-border business meeting the ratio is usually worse than the global average, because English is over-represented in professional contexts. EF Education First publishes an annual [English Proficiency Index](https://www.ef.com/wwen/epi/){rel=""nofollow""} covering more than 100 countries. "Very high proficiency" is defined there as the ability to *"use nuanced language in social, professional, and academic situations."* That is the bar most measured populations *do not* clear. The median country sits in the "Moderate" band; "High" and "Very High" are concentrated in a small set of Northern European and a few East Asian economies. In practice, the average global business meeting is staffed mostly by people operating somewhere between **B1 and C1** on the Common European Framework — fluent enough to function in routine conversation, *not* fluent enough to argue contract language, regulatory nuance, or technical edge cases without loss. That is the population the rest of this post is about. --- ## Failure mode 1 — False precision Fluent-sounding non-native speech is the most expensive failure mode in this space, because it removes the signal that anything needs checking. Lev-Ari and Keysar showed that listeners judge information delivered in a non-native accent as **less credible**, even when the content is identical and the speaker is reading from a script ([*Why don't we believe non-native speakers?*, Journal of Experimental Social Psychology, 2010](https://doi.org/10.1016/j.jesp.2010.05.025){rel=""nofollow""}). The bias is automatic, robust, and one-directional. The *opposite* half of the same bias is the more expensive one in business. When a non-native speaker produces a sentence that *sounds* fluent — grammar correct, intonation natural — listeners assume comprehension is symmetric: that the speaker understood the surrounding discussion as well as they appear to have produced their own sentence, and that the listener understood the speaker as well as they appear to have spoken. Both assumptions are routinely wrong. Tenzer, Pudelko, and Harzing surveyed employees across 15 multinational teams and found that perceived language proficiency systematically inflated trust attributions, with the imputed competence regularly turning out to be misplaced ([*The impact of language barriers on trust formation in multinational teams*, Journal of International Business Studies, 2014](https://doi.org/10.1057/jibs.2013.64){rel=""nofollow""}). The mechanism in the room: - The speaker emits a sentence whose grammar is correct and whose meaning is *off by 15%.* - Listeners hear the fluency, not the 15%. - Nobody asks the clarifying question, because the sentence "sounded fine." - The 15% gap is written into the minutes as agreed. A confidently wrong sentence is more dangerous than an obvious gap, because **the gap earns a follow-up question; the false clarity gets minuted.** This is *false precision* — the load-bearing cost of operating in shared imperfect English. --- ## Failure mode 2 — Self-censorship of nuance The person most likely to know the answer is often the person least equipped to express it in the working language. Volk, Köhler, and Pudelko reviewed the cognitive-neuroscience literature on L2 processing in multinational corporations and reported measurable additional load on working memory, processing speed, and emotional regulation when professionals operate in a non-native language ([*Brain drain: The cognitive neuroscience of foreign language processing in multinational corporations*, Journal of International Business Studies, 2014](https://doi.org/10.1057/jibs.2014.26){rel=""nofollow""}). The first thing to disappear under that load is *nuance* — the qualifiers, hedges, conditional clauses, and counter-arguments that a native speaker deploys without thinking. The observable consequence is documented in the international-business literature: senior subject-matter experts who would dominate a meeting in their native tongue become the **quietest people in the room** when forced to operate in English ([Tsedal Neeley, *Global Business Speaks English*, Harvard Business Review, 2012](https://hbr.org/2012/05/global-business-speaks-english){rel=""nofollow""}; Neeley, *[The Language of Global Success](https://press.princeton.edu/books/hardcover/9780691175379/the-language-of-global-success){rel=""nofollow""}*, Princeton University Press, 2017). They are not silent because they have nothing to add. They are silent because the cost of expressing their actual view in English — finding the right verb tense, hedging without sounding evasive, qualifying without sounding weak — exceeds the cost of staying quiet. The decision is then made by whoever *can* express themselves fluently. Who are not, in general, the people who know most. Hinds, Neeley, and Cramton called this *language as a lightning rod*: language proficiency becomes a proxy for status, and status determines who speaks ([*Language as a lightning rod*, Journal of International Business Studies, 2014](https://doi.org/10.1057/jibs.2013.62){rel=""nofollow""}). **The most knowledgeable person in the room becomes the least articulate.** --- ## Failure mode 3 — Fluency outranks authority In any negotiation conducted in shared English, the native-English side holds an advantage *before any substance is exchanged.* The non-native side spends part of its cognitive budget on language production; the native side spends all of it on substance. This is a measurable processing asymmetry, not a personality effect. The neuroscience review cited above reports working-memory penalties of roughly 20–30% on equivalent tasks when the same person operates in L2 versus L1. Translated into a live negotiation: the non-native side is running on perhaps 70% of its cognitive budget while the native side runs on all of it. The visible consequence is not that the non-native side speaks *less*. The consequence is that the non-native side **agrees more.** Hedges drop. "I think we could probably consider" collapses into "OK." The native side gets the wording it wanted; the non-native side feels it got *most* of what it wanted; the gap only surfaces in the contract review weeks later. This is the part of the language tax that is hardest to see in the moment and most expensive to fix afterwards. --- ## The strongest evidence: where "good-enough" English literally kills The three failure modes above are well-attested in research literatures, but skeptical readers can dismiss them as soft-science findings about feelings. The strongest evidence for the thesis comes from the two domains where second-language English has been measured against a hard outcome — and where regulators have already acted on what the data shows. ### Aviation — when ICAO regulated the problem **Tenerife, 27 March 1977.** A KLM 747 and a Pan Am 747 collided on the runway at Los Rodeos. 583 people died — still the worst accident in aviation history. The investigation identified several contributing causes, including a non-standard radio exchange. KLM Captain Van Zanten, operating in second-language English under heavy time pressure, told the tower *"we are now, uh, at takeoff."* The tower acknowledged with *"OK."* Van Zanten's phrasing was ambiguous between *"we are in the takeoff position"* and *"we are in the process of taking off"*; the Pan Am crew, also second-language English speakers, were still on the runway. Subsequent reforms by ICAO and national authorities introduced standardized phraseology specifically to remove this class of ambiguity ([Spanish Aviation Authority final report; SKYbrary case study](https://skybrary.aero/accidents-and-incidents/b742-b741-tenerife-airport-spain-1977){rel=""nofollow""}). **Avianca 052, 25 January 1990.** A Boeing 707 ran out of fuel and crashed near Cove Neck, New York, killing 73. The flight crew, operating in second-language English with New York ATC, reported being *"running out of fuel"* — a phrase that does not exist in ICAO standard phraseology and that ATC did not interpret as a declared emergency. The standard term *"fuel emergency"* or *"minimum fuel"* was never used. The NTSB's final report identified the failure of the flight crew to use standard phraseology among the probable causes ([NTSB Aircraft Accident Report AAR-91/04](https://www.ntsb.gov/investigations/AccidentReports/Reports/AAR9104.pdf){rel=""nofollow""}). **The regulatory response.** In 2003, ICAO adopted **Language Proficiency Requirements (LPRs)** — Annex 1 (Personnel Licensing) and Annex 10 (Aeronautical Telecommunications) — making demonstrated English proficiency a licensing requirement for international flight crew and air traffic controllers worldwide. A six-level rating scale was defined; **Level 4 ("Operational") is the minimum** for international operations. Below Level 4, the licence is not valid for international flight. The implementation deadline was March 2008 ([ICAO Doc 9835, *Manual on the Implementation of ICAO Language Proficiency Requirements*](https://store.icao.int/en/manual-on-the-implementation-of-icao-language-proficiency-requirements-doc-9835){rel=""nofollow""}). The regulator's reasoning here is the thesis of this post, written into binding international law: *non-native English at "good-enough" level is unsafe in high-stakes communication; either everyone speaks the language at a defined operational standard, or they don't get to fly the airplane.* ### Medicine — when the comparison is measurable The other domain with a hard outcome is medical interpretation, and here the comparison is even cleaner: same clinical encounter, same patient, *with* versus *without* a trained interpreter. Glenn Flores and colleagues audio-recorded 57 emergency-department encounters with Spanish-speaking limited-English-proficient patients in two pediatric EDs and counted every interpretation error and its potential clinical consequence ([Flores et al., *Errors of medical interpretation and their potential clinical consequences: a comparison of professional versus ad hoc versus no interpreters*, Annals of Emergency Medicine, 2012](https://doi.org/10.1016/j.annemergmed.2012.01.025){rel=""nofollow""}). Headline findings: - **1,884 interpretation errors** identified across the 57 encounters. - **18% of all errors had potential clinical consequences.** - The clinically consequential error rate dropped to **12%** when a **professional interpreter with ≥100 hours of training** was used. - The rate was **22%** with **professional interpreters with <100 hours of training**, **20%** with **ad hoc interpreters** (family members, untrained bilingual staff), and **20%** with **no interpreter at all**. Two things to read out of this. First: ad-hoc interpretation by a fluent-sounding bilingual person who is not a trained interpreter performs **no better than no interpreter at all** on clinical-consequence rates. The fluency does not translate into accuracy on the parts that matter. Second: the protective effect only kicks in once interpretation crosses a *defined training threshold* — exactly the same regulatory pattern as ICAO Level 4. "Good enough" is not a category that exists in this data. Either you cross the threshold or you don't. Earlier work by the same author had already shown the same pattern in primary care ([Flores et al., *Errors in medical interpretation and their potential clinical consequences in pediatric encounters*, Pediatrics, 2003](https://doi.org/10.1542/peds.111.1.6){rel=""nofollow""}). The 2012 study quantified it against a defined competency standard. ### Why this matters in the boardroom The boardroom is not the cockpit and it is not the ED. The stakes per minute are lower; the consequences arrive in months, not seconds; the cost is in dollars and reputation rather than lives. But the *mechanism* is identical. A non-native speaker operating in a shared imperfect language emits fluent-sounding utterances whose semantic content has shifted; listeners receive the fluency as a signal of accuracy; the gap is not detected at the moment of emission; the gap surfaces only when the artifact built on top of it (the procedure, the dispatch instruction, the contract clause) is executed against reality and turns out to mean something different than the room believed. Aviation and medicine matter for the boardroom because they are the two domains where the cost of running this mechanism was *quantified*, and a regulator agreed it was unacceptable. They are the natural experiment for the thesis. --- ## The cost in business — where there is data Business doesn't have a Tenerife or a Flores study with the same scientific cleanliness. What it has is survey data — self-reported, mixed-mechanism, but at scale. **The Economist Intelligence Unit, 2012.** A survey of more than 500 senior executives across 51 countries — *Competing across borders: How cultural and communication barriers affect business.* Among the findings: - **49% of respondents reported that misunderstandings have stood in the way of major international transactions, incurring significant losses for their company.** - **64% reported that differences in language and culture make it difficult to gain a foothold in unfamiliar markets.** - **67% reported that miscommunication is interfering with their international business efforts.** The report does not split "language miscommunication" from "cultural miscommunication" cleanly, but for the half of respondents reporting *failed major transactions*, the cost line is real even if the mechanism inside is mixed. The business literature lacks the controlled "with vs without interpreter" comparison that medicine has, and the regulatory backstop that aviation has — but the underlying mechanism is the one the cleaner literatures have already pinned down. --- ## "But doesn't operating in L2 make people better decision-makers?" The strongest published counterargument to this post's thesis is the **Foreign-Language Effect**, demonstrated by Keysar, Hayakawa, and An at the University of Chicago ([*The Foreign-Language Effect: Thinking in a Foreign Tongue Reduces Decision Biases*, Psychological Science, 2012](https://doi.org/10.1177/0956797611432178){rel=""nofollow""}). Their experiments showed that when people considered classic decision-theory problems — framing effects, loss aversion, the Asian disease problem — *in their second language*, they exhibited **less** of the standard cognitive bias than when considering the same problems in their L1. Subsequent work has replicated the effect across multiple language pairs and decision-bias paradigms (see also Costa et al., *Cognition*, 2014). This is a real finding. So why doesn't it refute the thesis? Three reasons: 1. **The Foreign-Language Effect is about individual reasoning, not multi-party communication.** Keysar et al. measured what happens when one person decides a problem alone in their head, in their L2, with no listener and no communication partner. The three failure modes in this post are all properties of *interaction* — false precision and self-censorship and the fluency-authority gradient only exist because there is more than one person in the room. The Foreign-Language Effect doesn't enter that domain at all. 2. **The effect operates on emotional/heuristic biases, not on semantic accuracy.** The mechanism is that L2 creates emotional distance from the framing of the problem, so the listener falls back on more analytical (System 2) processing. Useful for reducing loss aversion. Not useful for *transmitting* a contract clause without semantic drift, which is what the boardroom is doing. 3. **Aviation and medical regulators have already weighed both effects against each other.** ICAO and the medical-interpretation literature both know the L2 reasoning literature exists. Neither field decided that the bias-reduction benefit was worth the false-precision cost. Both went the other way: define a proficiency standard, enforce it, and require interpreters or standardized phraseology for everyone below the standard. The honest summary: operating in L2 makes *you* — alone, in your head — slightly more rational on certain framing tasks. It makes *the room* you are in less able to transmit information accurately. The Foreign-Language Effect is a property of individual cognition; the three failure modes above are properties of multi-party communication. They don't cancel. --- ## Why "good-enough" machine translation does not fix any of it A generic machine-translation layer slotted on top of an English meeting addresses none of the three failure modes, and in some configurations makes them worse: - **False precision compounds.** A good-enough MT layer also emits fluent-sounding output, now layered on top of fluent-sounding non-native English. Two stacks of unverified fluency separate source intent from listener comprehension. - **Self-censorship persists.** If the working language of the meeting is still English and translation only serves the *listeners*, the speakers still pay the L2 cost. They still drop the nuance. The translation pipeline preserves the loss faithfully. - **The fluency–authority gradient flips, but does not flatten.** A poorly tuned translation layer just moves the advantage to whoever has the best engine in their corner, not to whoever knows most. The fix is not "add translation on top of an English meeting." The fix is to remove the requirement that *any* participant operate in a language they do not fully own. That is a different architectural choice — and the one we're building toward. --- ## What changes when every participant speaks their native language The structural change is simple to state and hard to engineer: 1. **Every participant speaks their own native language** — no L2 cognitive load, no nuance drop, no self-censorship. 2. **Every participant hears every other participant in their own native language**, with sub-second latency and tone preserved. 3. **The translation layer is auditable**: source utterance, target utterance, per-language transcripts exported as a single bundle, with per-pair quality measured and [published monthly on real traffic](https://intermind.com/benchmark) rather than asserted as a marketing number. All three failure modes are linked to the *same* requirement — that someone in the room operate in a shared imperfect language. Removing that requirement removes them together. Adding translation on top of that requirement does not. That is the difference the next class of cross-border meeting tooling is going to have to be measured against. Aviation regulators didn't accept "everyone speaks pretty good English" as the answer. Medical regulators didn't either. The boardroom should not be the last domain that does. --- ## A reading list The load-bearing sources for this post, in roughly the order they were used: **On the cognitive and social mechanism** - Lev-Ari, S., & Keysar, B. (2010). [*Why don't we believe non-native speakers?*](https://doi.org/10.1016/j.jesp.2010.05.025){rel=""nofollow""} Journal of Experimental Social Psychology, 46(6), 1093–1096. - Tenzer, H., Pudelko, M., & Harzing, A.-W. (2014). [*The impact of language barriers on trust formation in multinational teams.*](https://doi.org/10.1057/jibs.2013.64){rel=""nofollow""} Journal of International Business Studies, 45(5), 508–535. - Hinds, P. J., Neeley, T. B., & Cramton, C. D. (2014). [*Language as a lightning rod: Power contests, emotion regulation, and subgroup dynamics in global teams.*](https://doi.org/10.1057/jibs.2013.62){rel=""nofollow""} Journal of International Business Studies, 45(5), 536–561. - Volk, S., Köhler, T., & Pudelko, M. (2014). [*Brain drain: The cognitive neuroscience of foreign language processing in multinational corporations.*](https://doi.org/10.1057/jibs.2014.26){rel=""nofollow""} Journal of International Business Studies, 45(7), 862–885. - Neeley, T. (2012). [*Global Business Speaks English.*](https://hbr.org/2012/05/global-business-speaks-english){rel=""nofollow""} Harvard Business Review. - Neeley, T. (2017). [*The Language of Global Success.*](https://press.princeton.edu/books/hardcover/9780691175379/the-language-of-global-success){rel=""nofollow""} Princeton University Press. - Keysar, B., Hayakawa, S. L., & An, S. G. (2012). [*The Foreign-Language Effect: Thinking in a Foreign Tongue Reduces Decision Biases.*](https://doi.org/10.1177/0956797611432178){rel=""nofollow""} Psychological Science, 23(6), 661–668. **Aviation** - ICAO. [*Manual on the Implementation of ICAO Language Proficiency Requirements (Doc 9835).*](https://store.icao.int/en/manual-on-the-implementation-of-icao-language-proficiency-requirements-doc-9835){rel=""nofollow""} International Civil Aviation Organization. - NTSB. [*Avianca, The Airline of Colombia, Boeing 707-321B, HK 2016, Fuel Exhaustion, Cove Neck, New York, January 25, 1990 (AAR-91/04).*](https://www.ntsb.gov/investigations/AccidentReports/Reports/AAR9104.pdf){rel=""nofollow""} - SKYbrary. [*Tenerife airport disaster, 1977 — accident case study.*](https://skybrary.aero/accidents-and-incidents/b742-b741-tenerife-airport-spain-1977){rel=""nofollow""} **Medicine** - Flores, G., Abreu, M., Barone, C. P., Bachur, R., & Lin, H. (2012). [*Errors of medical interpretation and their potential clinical consequences: a comparison of professional versus ad hoc versus no interpreters.*](https://doi.org/10.1016/j.annemergmed.2012.01.025){rel=""nofollow""} Annals of Emergency Medicine, 60(5), 545–553. - Flores, G., et al. (2003). [*Errors in medical interpretation and their potential clinical consequences in pediatric encounters.*](https://doi.org/10.1542/peds.111.1.6){rel=""nofollow""} Pediatrics, 111(1), 6–14. **Business** - Economist Intelligence Unit. (2012). *Competing across borders: How cultural and communication barriers affect business.* Sponsored by EF Education First. **Population** - Crystal, D. (2003). *English as a Global Language* (2nd ed.). Cambridge University Press. - EF Education First. [*English Proficiency Index*](https://www.ef.com/wwen/epi/){rel=""nofollow""} (annual). — The Mind.com Team **Aviation said it in 2008. Medicine has been counting it for twenty years. The boardroom is the last room still pretending that "good enough" is good enough.** # The 9 best Fireflies.ai alternatives in 2026 (tested for multilingual meetings) Fireflies.ai is a good product at the thing it was built for: it joins your call, transcribes it in 100+ languages, and hands you a searchable summary afterward. If you're reading this, something about it isn't fitting — the per-seat price at team scale, the storage caps on the lower plans, the bot in every meeting, or a need it simply wasn't built for. This guide compares the 9 alternatives people actually shortlist in 2026, with verified pricing and limits. Full disclosure up front: **we build one of them** ([InterMIND](https://intermind.com){rel=""nofollow""}), and it's the right choice only for a specific case — meetings that happen in more than one language. For everything else, we'll tell you honestly which competitor to pick. Every card links to the vendor's pricing; numbers below were checked in July 2026 and can drift. > Want the deeper 1-on-1 comparisons? See [InterMIND vs Fireflies](https://intermind.com/compare/fireflies) and [InterMIND vs Otter](https://intermind.com/compare/otter), the head-to-head [Otter vs Fireflies](https://intermind.com/blog/otter-vs-fireflies), or the category-level [buyer's guide to AI meeting translation tools](https://intermind.com/blog/best-ai-conference-translation-tools). Leaving Otter instead? [The 9 best Otter.ai alternatives](https://intermind.com/blog/otter-ai-alternatives). --- ## The verdict table | Tool | Best for | Free plan | Paid from\* | Languages | | ---------------------------------------------------- | --------------------------------------------- | ------------------- | ---------------------- | ------------------ | | [InterMIND](https://intermind.com/#tool-1) | Meetings in 2+ languages, translated live | Yes (Basic) | per-user, 14-day trial | 24 voice / 30 docs | | [Otter.ai](https://intermind.com/#tool-2) | English-first teams who live in transcripts | 300 min/mo | \~$8.33/user/mo | 6 | | [Fathom](https://intermind.com/#tool-3) | Individuals who want free unlimited recording | Unlimited recording | \~$16/mo | 38 | | [tl;dv](https://intermind.com/#tool-4) | Sales teams that clip and coach | Unlimited recording | \~$18/seat/mo | 30+ | | [Notta](https://intermind.com/#tool-5) | Asian-language and translated transcripts | 120 min/mo | \~$9/mo | 58 | | [MeetGeek](https://intermind.com/#tool-6) | Auto-capture across a whole team | 3 h/mo | $9.99/user/mo | 100+ | | [Read AI](https://intermind.com/#tool-7) | Engagement metrics + notes | 5 meetings/mo | $15/user/mo | 16+ | | [Krisp](https://intermind.com/#tool-8) | Noisy environments, call centers | Trial | \~$8/user/mo | 16+ | | [Built-in notetakers](https://intermind.com/#tool-9) | Teams already paying for a suite | Included | with suite plan | varies | \*Lowest advertised per-user price on annual billing, July 2026. Month-to-month is typically 1.5–2× higher. --- ## How to choose in 30 seconds - **Your meetings are English-only and you want a record** → [Otter](https://intermind.com/#tool-2) if you live in transcripts, [Fathom](https://intermind.com/#tool-3) if you want free and unlimited. Don't overthink it. - **You transcribe a lot of non-English audio** → [Notta](https://intermind.com/#tool-5) (58 languages, translated transcripts) or [MeetGeek](https://intermind.com/#tool-6) (100+). - **Your meetings happen in more than one language at once** — people going quiet, or limping along in shared English → a notetaker records the problem; it doesn't fix it. You need live translation: that's [InterMIND](https://intermind.com/#tool-1). - **You're mostly annoyed by the price** → your platform's [built-in notetaker](https://intermind.com/#tool-9) may already be good enough. --- ## The 9 best Fireflies.ai alternatives []{#tool-1} ### 1. InterMIND — when the meeting itself isn't in one language Every other tool on this list does the same job as Fireflies: record now, read later. InterMIND does a different job — it translates the meeting **while it's happening**. Each participant picks a language when they join, and from then on they *hear* every speaker in it (in the speaker's own voice, sub-second latency), and *read* the chat, the shared notes, and dropped-in documents in it too. The AI summary at the end is just the last surface, not the product. - **Price:** free Basic plan; Pro and Business are per-user with a 14-day trial — see [pricing](https://intermind.com/pricing). - **Languages:** 24 for live voice, chat, and notes; 30 for document translation. [The honest per-surface breakdown.](https://intermind.com/blog/how-many-languages-do-you-support) - **Limits:** plans meter voice minutes and chat words on a rolling 30-day window, shared across the team. - **Best for:** recurring meetings where the room doesn't share a fluent common language — international teams, cross-border sales, distributed standups. - **Honest caveat:** if every meeting you run is in one language, InterMIND is the wrong tool — a notetaker below is cheaper and built for exactly that. Two things you can check rather than take on faith: translation quality is [published monthly per language pair](https://intermind.com/benchmark), not marketed as a language count, and [the live demo](https://intermind.com/demo) runs the production pipeline on your own voice. []{#tool-2} ### 2. Otter.ai — the default English notetaker The most-searched Fireflies alternative: transcription in six languages, live collaboration on the transcript, OtterPilot for auto-joining. The trade Fireflies users notice first: Otter caps *transcription minutes* (Fireflies caps storage instead), and its language list is short. - **Price:** free 300 min/mo (30 min per conversation); Pro $8.33/user/mo billed annually ($16.99 monthly), 1,200 min/mo; Business $19.99/user/mo annually ($30 monthly), unlimited meetings up to 4 h each. - **Languages:** 6 — English, Spanish, French, German, Japanese, Chinese. - **Best for:** English-heavy teams who want the transcript to be the workspace. - **Vs Fireflies:** better live-transcript UX; far fewer languages (6 vs 100+) and metered minutes. []{#tool-3} ### 3. Fathom — the best free recorder Fathom's free plan is genuinely unlimited — recordings, transcription, and instant summaries at no cost, which makes it the easiest "just leave Fireflies" move for individuals. - **Price:** free unlimited recording + transcription; Premium $16/mo billed annually ($20 monthly); Team $15 and Business $25/user/mo annually (2-user minimum). - **Languages:** 38 for transcription. - **Best for:** individuals and small teams who want a free, low-friction recorder. - **Vs Fireflies:** more generous free tier; lighter on team knowledge-base features. []{#tool-4} ### 4. tl;dv — sales teams on a budget tl;dv records and transcribes free without limits, then charges for the AI layer (summaries, multi-meeting reports, CRM sync, coaching). Its 2026 positioning is squarely sales-team enablement. - **Price:** free plan with unlimited recording/transcription but tightly limited AI summaries; Pro ~~$18/seat/mo billed annually (~~$29 monthly); Business \~$59/seat/mo annually. - **Languages:** 30+ for transcription. - **Best for:** sales teams that clip calls and coach reps without paying Gong prices. - **Vs Fireflies:** cheaper way to record everything; the useful AI features sit behind Pro. []{#tool-5} ### 5. Notta — the multilingual transcriber Notta's pitch is the transcript in another language: 58 transcription languages, translated transcripts on paid plans, and bilingual/real-time translation views as add-ons. If your problem is "my recordings aren't in English," Notta is the notetaker to shortlist. - **Price:** free 120 min/mo (short per-recording cap); Pro from \~$9/mo billed annually ($13.99 monthly), 1,800 min/mo; Business \~$16.67/seat/mo annually. Real-time translation is a paid add-on (from \~$6/mo). - **Languages:** 58 for transcription; transcript translation on paid plans. - **Best for:** individuals working across Asian and European languages who need the *record* translated. - **Vs Fireflies:** better translation of the transcript; still after-the-fact — the meeting itself stays untranslated. []{#tool-6} ### 6. MeetGeek — set-and-forget auto-capture MeetGeek auto-joins everything on the calendar and quietly builds a searchable meeting library with summaries and workflow automations — the "capture the whole team by default" pick. - **Price:** free 3 h of transcription/mo; Pro $9.99/user/mo, 20 h/mo; Business $17/user/mo, unlimited. - **Languages:** 100+ with automatic language detection. - **Best for:** managers who want every team call captured and summarized without anyone pressing record. - **Vs Fireflies:** similar breadth at a slightly lower entry price; smaller ecosystem of integrations. []{#tool-7} ### 7. Read AI — meeting analytics on top of notes Read AI layers engagement and sentiment metrics on top of transcripts and summaries — who talked, who tuned out, how the meeting *went*, plus reports that chase you into email and Slack. - **Price:** free 5 meeting transcripts/mo (1 h cap); Pro $15/user/mo billed annually ($19.75 monthly); Enterprise $22.50, Enterprise+ $29.75/user/mo annually. - **Languages:** 16+ for meeting reports. - **Best for:** leaders who want meeting-culture metrics, not just minutes. - **Vs Fireflies:** adds analytics Fireflies doesn't have; costs more per seat and covers fewer languages. []{#tool-8} ### 8. Krisp — notes plus noise cancellation Krisp started as the noise-cancellation app and grew an AI notetaker on top. Because it processes audio on-device, it works across *any* calling app without a bot joining — and the transcript quality benefits from the cleaned-up audio. - **Price:** Core \~$8/user/mo billed annually ($16 monthly) with unlimited AI notes; Advanced \~$15/user/mo annually; free 7-day trial. - **Languages:** 16+ for transcripts and summaries. - **Best for:** noisy environments, call centers, and the bot-averse. - **Vs Fireflies:** no meeting bot — noise removal runs on-device before the notetaker; lighter on search and knowledge-base features. []{#tool-9} ### 9. The built-in notetakers (Zoom, Teams, Meet) Before paying any third party, check what your suite already includes: Zoom AI Companion, Copilot in Teams, and Gemini in Meet all summarize meetings on their paid tiers now, and each platform has some level of translated captions. We've written up exactly how far each one goes and where it stops: [Zoom](https://intermind.com/blog/zoom-live-translation), [Teams](https://intermind.com/blog/teams-live-translation), [Google Meet](https://intermind.com/blog/google-meet-live-translation). - **Price:** bundled with the suite plan you may already pay for. - **Best for:** teams that live inside one platform and need "good enough" notes. - **Vs Fireflies:** free-ish and zero-setup; weaker cross-meeting search, and captions ≠ translation of the actual conversation. --- ## The question none of these lists ask Every comparison above — ours included — can be reduced to price × languages × limits. But there's a split that matters more, and it's the reason this guide exists: **A notetaker solves recall. It does not solve comprehension.** If your meetings happen in one language, recall is the whole job, and you should pick from cards 2–9 with confidence. But if half the room is limping along in someone else's language, a better transcript is a record of a meeting that already failed. The people who couldn't follow the discussion live won't be rescued by a summary of it — [translated after the fact, quality unverified](https://intermind.com/blog/why-translation-quality-marketing-is-broken). That's the case InterMIND was built for, and it's why it looks odd on this list: it's not a better notetaker, it's [live, per-participant translation](https://intermind.com/blog/real-time-meeting-translation) of the voice, chat, notes, and documents — with the notes as a by-product. If that's your actual problem, no notetaker on this page solves it. --- ## FAQ **Is there a completely free Fireflies alternative?** Fathom is the most generous: unlimited free recording, transcription, and basic summaries. tl;dv also records and transcribes free but tightly limits AI summaries. Fireflies' own free plan transcribes without limits but caps *storage* at 400 min per team. **What's the best Fireflies alternative for non-English meetings?** For transcribing recordings in other languages: Notta (58 languages) or MeetGeek (100+). For meetings that *happen* in several languages at once and need live translation rather than an after-the-fact transcript: [InterMIND](https://intermind.com/compare/fireflies). **Does Fireflies.ai translate meetings?** No. Fireflies transcribes in 100+ languages and can detect the language automatically, but it doesn't translate the live meeting — participants still need a common language during the call. **Otter or Fireflies — which is better?** Otter has the better live-transcript workspace; Fireflies covers far more languages (100+ vs 6) and prices storage instead of minutes. English-only and transcript-centric: Otter. Multi-language recordings or unlimited transcription: Fireflies. We wrote the full head-to-head: [Otter vs Fireflies](https://intermind.com/blog/otter-vs-fireflies). **Do I even need a third-party notetaker?** Maybe not — Zoom, Teams, and Meet all bundle AI summaries on paid tiers now. Third-party tools still win on cross-meeting search, CRM sync, and working across platforms. --- ## Try the part that's hard to fake If language — not note-taking — is your real bottleneck, the only test that matters is your own language pair, on your own audio: - **[Run the live demo](https://intermind.com/demo)** — InterMIND's production pipeline on your voice, any of 24 languages. - **[Read the benchmark](https://intermind.com/benchmark)** — monthly per-pair quality scores, full distribution, no cherry-picking. - **[InterMIND vs Fireflies, feature by feature](https://intermind.com/compare/fireflies)** — the detailed 1-on-1. — The Mind.com Team --- *Sources: vendor pricing pages ([Otter.ai](https://otter.ai/pricing){rel=""nofollow""}, [Fireflies.ai](https://fireflies.ai/pricing){rel=""nofollow""}, [Fathom](https://fathom.ai/pricing){rel=""nofollow""}, [tl;dv](https://tldv.io/app/pricing/){rel=""nofollow""}, [Notta](https://www.notta.ai/en/pricing){rel=""nofollow""}, [MeetGeek](https://meetgeek.ai/pricing){rel=""nofollow""}, [Read AI](https://www.read.ai/pricing){rel=""nofollow""}, [Krisp](https://krisp.ai/pricing/){rel=""nofollow""}), checked July 2026. Prices are annual-billing rates unless noted and may change.* # Is InterMIND free for meeting participants? Short answer: **yes — for every participant**. If someone invites you to an InterMIND meeting, everything on your side is free: joining, speaking and reading in your own language, receiving translated documents, getting the recap afterwards. Whether you have an InterMIND account or not changes *what you can keep after the call* — it never changes the price, which for a participant is zero either way. This page exists because "free" on a landing page is usually followed by an asterisk. Here is the whole asterisk, spelled out. ## The model: one plan per room, not one plan per person InterMIND is priced like a room, not like a seat at the table. **The host's team has a plan. Everyone the host invites just joins.** That split is structural, not promotional: - **The host is the only person with a billing relationship** — and only on paid plans. A host on the free Basic plan pays nothing either; paid plans exist for teams that need more than Basic's limits. - **Participants draw on the host's plan.** When a participant reads chat in their own language or receives a translated document, that usage counts against the host's plan limits — never against the participant. - **Participants have no billing surface at all.** There is no paywall, no upgrade prompt, and no card form on the participant's path — with an account or without one. If you have ever set up a multilingual call on a per-seat product, this is the difference to notice: you don't need every participant licensed or provisioned. You need one plan — the host's. ## What every participant gets, free Any participant in an InterMIND meeting — invited by link, at no cost: - **hears and reads the meeting in their own language** — live, whatever language each other participant is speaking or typing; - **speaks and writes in their own language**, and every other participant gets it in theirs; - **receives shared documents translated** into their language, when the host's plan includes document translation. You pick a language once, at the door. Nothing else to configure, nothing to pay. ## With an account or without one — what actually differs You can join a meeting without signing up at all: click the link, enter a name, pick a language. InterMIND calls that joining **as a guest**, and it exists so that a first-time invitee is never blocked at the door. To be precise about what it is: a guest is simply a participant who isn't signed in — nothing more. The difference shows up **after** the call, because everything InterMIND produces from a meeting needs somewhere to be delivered: - **Signed-in participants** get the recap in their own language, the links to the recording and meeting artifacts, and the meeting's chat history in the channel — readable in their language, for as long as the channel lives. - **Guests get none of that.** A guest identity is temporary by design — it is deleted within about 24 hours (which is also a privacy feature: nothing about a guest is retained). There is no account to deliver a recap to, so the meeting ends and nothing follows. So the honest advice is the opposite of "stay a guest": **create an account — it's free.** Signing up costs nothing, there is no card involved, and it is the only thing standing between you and the recap, the recording, and the history in your language. Guest mode is for the moment you're already late to a call — not a way to use InterMIND. ## Where the limits actually live Every limit in InterMIND belongs to the host's plan — the full row-by-row breakdown is on the [pricing page](https://intermind.com/pricing), but the shape is: - **Basic (free plan):** free forever for the host too — multilingual meetings with chat translation capped at 10,000 words per rolling 30-day window; document translation not included. - **Paid plans:** lift the chat-translation cap and add document translation as a metered feature — counted in distinct files over a rolling 30-day window (10 on Pro, 30 on Business, 100 on Enterprise). None of these numbers are a participant's problem. They bound what the *room* can do; participants simply use whatever the room provides. ## Why we price it this way A multilingual meeting only works if **everyone** is in it. The moment the other side of the call has to pay or get provisioned, half of your invitees show up untranslated — and the meeting quietly falls back to whatever language dominates. Making participation free is not generosity; it is the only configuration in which the product does what it promises. ## Try it first - **[Try the live demo](https://intermind.com/demo)** — join a real translated meeting from your browser, no signup needed. - **[See every limit side by side](https://intermind.com/pricing)** — the full plan comparison, row by row. ## FAQ **Do meeting participants need an InterMIND account?** Not to join — a link is enough. But post-meeting delivery is tied to an account: the recap in your language, recording links, and the channel history go to signed-in participants. The account is free, so the practical answer is: join however you like, sign up if you want to keep anything. **Who pays for the translation in a meeting?** The host's team plan covers the room. Participant usage — translated voice, chat, documents — draws on the host's plan limits. Participants themselves have no billing relationship with InterMIND at all, and a host on the free Basic plan pays nothing either. **Is there a free plan for hosts too?** Yes. The Basic plan is free and includes multilingual meetings with chat translation capped at 10,000 words per rolling 30-day window. Document translation starts on paid plans. The full breakdown is on the [pricing page](https://intermind.com/pricing). **What does a guest miss compared to a signed-in participant?** Inside the meeting — nothing: translation of voice, chat, and documents works the same. After the meeting — everything: recap, recording links, and history are delivered to signed-in participants only, and guest identities are deleted within about 24 hours. **Do participants need to install an app?** No — joining from a desktop or mobile browser works with just the link. Native mobile apps for iOS and Android exist for people who prefer them, but they are optional. **Can a participant be charged by accident?** No. There is no payment surface on the participant path — no card form, no upgrade prompt, no paywall. Billing exists only inside the host team's settings. # We finished our GDPR audit. Here's what we actually closed. A few weeks ago we wrote that ["GDPR-compliant" on a video tool's homepage means less than you think](https://intermind.com/blog/gdpr-compliant-video-conferencing) — that GDPR is a set of obligations on *you*, the data controller, that a vendor either helps you meet or quietly leaves on your desk. The honest way to back that claim is to do the work on our own side and show it, line by line. So we did. We ran a full audit of the InterMIND codebase against the obligations that fall on us as a data processor, fixed every gap that had code behind it, and verified each one against the running product. This post is the close-out report — not a badge, a checklist with our answers. We're deliberately not claiming "100% GDPR-certified." GDPR isn't a certificate you pass — and we won't wave an ISO badge we don't yet hold. What we *can* say: the architectural and process obligations a DPO works through now have concrete, verifiable answers, each checked against the running code. --- ## What we closed ### Right to erasure (Art. 17) — the cascade actually runs Deleting your account doesn't just deactivate it. `POST /api/user/delete-account` runs a real cascade: it nukes your meetings → participants, messages, conferences, transcriptions; it sweeps your storage blobs out of Tigris **before** the database cascade so nothing is left orphaned — chat attachments *and* video-recording files, both columns; and it cancels your Stripe subscriptions and deletes the Stripe customer. On-demand deletion is there too — drop a channel or a message from the UI and it's gone. Anonymous (guest) accounts get their own deletion endpoint plus a background sweep every 6 hours, under a monitored cron. The audit surfaced one gap here — recording blobs that the database cascade dropped but storage kept — and we closed it: erasure now leaves nothing behind in object storage. ### Retention (Art. 5(1)(e)) — a documented criterion Art. 5(1)(e) doesn't require an automatic time-to-live. It requires a **defined retention criterion**. Ours is now written into the Privacy Policy: data is kept until you or your team owner delete it, and deleting your account erases everything. That's the same model collaboration tools like Slack and Notion run on — persistence is the expected behavior, and you stay in control of it. The criterion is stated, not implied. ### Analytics consent (Art. 6/7) — opt-out by default A Usercentrics consent banner (shown to EU visitors) gates analytics, and PostHog ships with `opt_out_capturing_by_default: true` — nothing is captured until consent is given, not the other way around. ### Data portability (Art. 20) — a real export `GET /api/user/export` builds a ZIP of your meetings, messages, recordings, and translations, with a 7-day download window and automatic cleanup. Access, deletion, and portability are tooling that works, not promises in a policy. ### No meeting content reaches a US-domiciled model The single biggest flow of meeting content — live voice and chat translation — runs on **our own engine in France**, never a third-party LLM. The post-meeting AI steps that *do* use a general-purpose model (the digest, the note-editor's generative actions) run on **EU-hosted Mistral with zero-data-retention**, pinned so hard that the request fails rather than fall back to a non-ZDR or US host. We also scrubbed participant names and utterance text out of the conference browser logs that PostHog session-recording could otherwise capture. The full vendor-by-vendor map is in [*Where one InterMIND meeting actually runs*](https://intermind.com/blog/where-one-intermind-meeting-actually-runs). ### Transparency — sub-processors and processing records, published The [sub-processor list](https://intermind.com/legal/subprocessors) is live, with what each vendor does and where it's domiciled — not "available on request." Behind it sits a Record of Processing Activities (ROPA) built from the live schema: 11 processing operations, the security measures on each, and the erasure / portability paths. Our [Privacy Policy](https://intermind.com/privacy) and [Terms](https://intermind.com/terms) now run under our own legal entity, with the real processing chain described. ### EU runtime — pinned, not promised Every runtime hop a meeting takes is in the EU: app and APIs on Vercel Frankfurt, the meeting server on Fly Paris, application data in Neon Postgres (AWS Frankfurt), errors on Sentry EU, analytics on PostHog EU, email via Resend Ireland. Object storage on Tigris is now **pinned to EU regions** (Frankfurt + Amsterdam) — every new write lands in the EU regardless of where the user is. The full architecture is on our [security page](https://intermind.com/docs/security). --- ## Why this matters for your procurement For most EU buyers — German Mittelstand, regulated teams running standard GDPR DPAs — the data-residency question now has a direct answer: the data doesn't leave the EU at runtime, erasure works, the retention criterion is stated, and the sub-processor list is on the table. That's a much shorter conversation than "let us get back to you on where the data goes." For French *souveraineté numérique* and SecNumCloud-grade procurement, vendor corporate domicile is itself a criterion — a deeper conversation about deployment topology that we'll have honestly rather than oversell. And the one thing we won't do is wave an ISO certificate we don't yet hold: certifications are on the roadmap, but our answer to the checklist is **architectural and verifiable today**. --- ## See it for yourself - [GDPR-compliant video conferencing: the full DPO checklist](https://intermind.com/blog/gdpr-compliant-video-conferencing) — the seven things to make any vendor answer, and where Zoom sits. - [Where one InterMIND meeting actually runs](https://intermind.com/blog/where-one-intermind-meeting-actually-runs) — the vendor map, with the gaps named. - [Security & Privacy](https://intermind.com/docs/security) · [Sub-processors](https://intermind.com/legal/subprocessors) · [Privacy Policy](https://intermind.com/privacy) · [Terms](https://intermind.com/terms) - **[Try the live demo](https://intermind.com/demo)** — run the live, EU-runtime, multilingual pipeline on your own audio. GDPR compliance isn't a badge you buy — it's work you do and can show. This is ours, line by line. If your DPO needs an answer this post doesn't give, write us. — The Mind.com Team --- *Sources: [Regulation (EU) 2016/679 (GDPR)](https://eur-lex.europa.eu/eli/reg/2016/679/oj){rel=""nofollow""} — the articles audited against; internal claims verified against the InterMIND codebase and running product; checked August 2026.* # GDPR-compliant video conferencing (2026): what it actually takes (and a Zoom alternative that translates) "GDPR-compliant" on a video platform's homepage is close to content-free. GDPR is not a certification you pass — it's a set of obligations on *you*, the data controller, that a vendor either helps you meet or quietly leaves on your desk. The useful question isn't "is this tool GDPR-compliant?" It's "**what does this tool make me responsible for, and where does my meeting data actually go?**" > **🔒 Update — we did this to ourselves.** Since this post went up, we ran InterMIND's own codebase through the full checklist below and closed every item — erasure, retention, EU runtime, sub-processors — verified against the code, not asserted on a page. **→ [Read the audit close-out, line by line](https://intermind.com/blog/gdpr-audit-what-we-closed).** This post is the plain-terms version of that question: the checklist a DPO actually works through, where Zoom sits on it (fairly — it's more compliant-capable than the internet implies), and where an EU-runtime alternative changes the answer. If you want the broader "how should real-time meetings work" frame, that's the [pillar guide](https://intermind.com/blog/real-time-meeting-translation); this one is specifically about the data-protection layer. --- ## The plain-terms checklist Strip the marketing and "GDPR-compliant video conferencing" comes down to seven things you can actually verify: 1. **A signed DPA.** A Data Processing Addendum that names the vendor as your processor, with documented purposes and instructions. No DPA, no lawful processing — full stop. 2. **A real sub-processor list.** Every third party that touches meeting data — transcription, storage, email, analytics, AI features — named, with what they do and where they're domiciled. 3. **Where data is processed at runtime.** The physical region your audio, transcripts, recordings, and metadata are handled in. This is what most data-residency clauses are actually about. 4. **The international-transfer mechanism.** If any data leaves the EEA, on what legal basis? Standard Contractual Clauses (SCCs), adequacy, or a residency setup that avoids the transfer entirely. This is the Schrems II question, and it doesn't go away because a homepage says "compliant." 5. **Security posture you can audit.** ISO 27001, ISO 27701, SOC 2 — independent attestations, not self-assertions. 6. **Data-subject rights tooling.** Can you actually fulfil access, deletion, and portability requests for meeting data, or only in theory? 7. **Retention and deletion you control.** Recordings, transcripts, and AI summaries deleted on your schedule, not the vendor's default. A tool is "GDPR-compliant" for your purposes only when all seven have concrete answers. Most homepages answer zero of them. --- ## Where Zoom actually stands (fairly) Zoom is more GDPR-capable than its reputation suggests, and it's worth being accurate about that rather than scoring cheap points: - It publishes a **global DPA with the 2021 EU Standard Contractual Clauses** built in. - It offers **EU Data Residency** — "Zoom EU Infrastructure" — for Enterprise and Education customers, covering Meetings, Webinar, Chat, Phone, and Contact Center. - It holds **ISO 27001 and, as of February 2026, ISO 27701** (the privacy-management extension), plus published EU public-sector sovereignty controls. So a properly configured Zoom deployment, on the right tier, can satisfy most of the checklist. The honest caveats are about defaults and domicile, not capability: - **EU residency is a tier-and-config add-on, not the default.** It's on Enterprise/Education and has to be turned on. The free and lower paid tiers don't get it. "We use Zoom" and "we use Zoom configured for EU residency" are different procurement facts. - **Zoom is a US-domiciled company.** Even with SCCs and EU residency, the *vendor's corporate domicile* is in scope for CLOUD-Act and sovereignty-grade evaluations. For standard GDPR this is manageable with SCCs; for French *souveraineté numérique* or SecNumCloud-grade procurement, corporate domicile is itself a criterion, and that's a harder conversation. - **The meeting still happens in one language.** Zoom's translated captions and AI Companion help, but the room doesn't become genuinely multilingual — every participant hearing the meeting live in their own language. If your compliance problem is also a *language* problem (cross-border audits, multi-site CAPA reviews), that gap is unsolved regardless of where the data lives. The short version: Zoom **can** be GDPR-compliant when configured for it. Whether that's *sufficient* depends on how sovereignty-sensitive your buyer is — and whether the meetings are multilingual. --- ## The EU-runtime alternative We built InterMIND so the data-residency answer is the default, not a configuration project — and so the meeting can actually be multilingual. On the checklist above, the part that's structural rather than promised is **where the meeting runs**. As verified against the live deployment: - **Every runtime hop is in the EU.** App and server APIs on Vercel in Frankfurt (`fra1`); the meeting WebSocket server on Fly in Paris (`cdg`); application data in Neon Postgres on AWS Frankfurt (`eu-central-1`); recordings on Tigris, EU-pinnable; error/analytics on Sentry EU and PostHog EU; transactional email via Resend in Ireland. - **The translation engine is our own code, in France** (OVH) — not a US general-purpose model we resell. The single biggest flow of meeting content stays EU-resident and never touches a third-party LLM. Document translation goes to **DeepL in Cologne** — also an EU company. - **The sub-processor list with corporate-domicile detail ships with the DPA** as standard practice, not on request. **No meeting content reaches a US-domiciled model.** Everything derived from a meeting that needs a language model stays on EU processors: the *post-meeting* AI digest (topics, decisions, action items) runs on EU-hosted Mistral with zero-data-retention; the post-meeting summary and the AI note-editor translate through our own EU engine, and the editor's generative actions (fix, extend, simplify) run on the same EU Mistral. The one place a US model is still in the loop touches **no meeting data**: our public translation-quality benchmark judges machine translations of [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""} reference sentences — a fixed public dataset, not anyone's call. The full vendor-by-vendor map is in [*Where one InterMIND meeting actually runs*](https://intermind.com/blog/where-one-intermind-meeting-actually-runs), and the build-vs-buy view is in [*What one InterMIND meeting is built from*](https://intermind.com/blog/what-one-intermind-meeting-is-built-from). One thing we **don't** claim: we are not going to wave an ISO certificate we don't yet hold. Our answer to the checklist is architectural — EU runtime, own engine, transparent sub-processors, a real DPA — and we'd rather you verify that than trust a badge. Certifications are on the roadmap; the runtime is real today. We ran our own codebase through this exact checklist and closed each item against the code — [the close-out report, line by line](https://intermind.com/blog/gdpr-audit-what-we-closed). And the part Zoom's residency settings can't add: the meeting is **multilingual by design** — every participant live in their own language across 24 languages, voice and chat and notes, with [translation quality published openly](https://intermind.com/benchmark) instead of asserted. How that works is in the [pillar guide](https://intermind.com/blog/real-time-meeting-translation); the regulated-meeting use case is in [*multilingual compliance meetings*](https://intermind.com/blog/multilingual-compliance-meetings). --- ## The checklist as a table The seven items against the two tools this post discusses in detail. Cells state only what each vendor documents (Zoom links in Sources below; InterMIND claims verified line-by-line in [the audit close-out](https://intermind.com/blog/gdpr-audit-what-we-closed) and [the runtime map](https://intermind.com/blog/where-one-intermind-meeting-actually-runs)) — for any other vendor, make them answer the left column in writing. | Checklist item | Zoom (documented) | InterMIND (documented) | | -------------------------------- | --------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | | Signed DPA | Global DPA with 2021 EU SCCs | DPA; sub-processor list ships with it | | Sub-processor list | Published | Ships with the DPA, incl. corporate domicile per vendor | | Runtime region | EU Data Residency on Enterprise/Education, has to be enabled | EU at every hop by default: Frankfurt (app/API/DB), Paris (meeting server), own translation engine in France | | International-transfer mechanism | SCCs | Meeting data stays in the EEA — no transfer to paper over | | Auditable security posture | ISO 27001; ISO 27701 as of Feb 2026 | No ISO certificate yet — the answer is architectural and open to verification, certifications on the roadmap | | AI features and meeting content | Translated captions / AI Companion (see [our Zoom breakdown](https://intermind.com/blog/zoom-live-translation)) | No meeting content reaches a US-domiciled model; digest/summary on EU processors, zero retention | | Multilingual meeting | Captions translate text; the room stays one-language live | 24 languages live per participant: voice, chat, notes | --- ## FAQ **Is Zoom GDPR compliant?** It can be, on the right tier and configuration: Zoom publishes a global DPA with the 2021 EU SCCs, offers EU Data Residency for Enterprise and Education customers, and holds ISO 27001 and (since February 2026) ISO 27701. The honest caveats: EU residency is a tier-and-config add-on rather than the default, and Zoom remains a US-domiciled company — manageable under standard GDPR with SCCs, but a real criterion in sovereignty-grade procurement. **What makes a video conferencing tool GDPR compliant?** Seven verifiable things, not a homepage badge: a signed DPA, a real sub-processor list, a documented runtime region, a lawful international-transfer mechanism, independently audited security, working data-subject-rights tooling, and retention you control. A vendor is compliant *for your purposes* when all seven have concrete answers in writing. **Does GDPR require video meetings to stay in the EU?** No — GDPR requires a lawful transfer mechanism (SCCs, adequacy) when data leaves the EEA, not residency as such. Residency matters because it removes the transfer question entirely — the post-Schrems II analysis gets short when the data never leaves. That's the difference between "compliant with paperwork" and "nothing to paper over." **Which video conferencing tools run entirely in the EU?** Ask any vendor for their runtime map: region per hop — app, APIs, meeting media, database, recordings, analytics, email, AI features. InterMIND publishes exactly that ([every hop, vendor by vendor](https://intermind.com/blog/where-one-intermind-meeting-actually-runs)): Vercel Frankfurt, Fly Paris, Neon Frankfurt, translation engine on OVH France, digest on EU-hosted Mistral with zero retention. Whatever tool you evaluate, the map — not the marketing page — is the answer. --- ## Which buyer is this for - **German Mittelstand and regulated EU teams** running standard GDPR DPAs: an EU-at-every-hop runtime answers the residency question directly, without a tier upgrade or a residency project. The transfer-mechanism conversation gets a lot shorter when the data doesn't leave. - **French public-sector and *souveraineté*-grade procurement**, where vendor corporate domicile is itself part of the spec: that's a deeper conversation about deployment topology, and one we'll have honestly rather than oversell. - **US-domestic and APAC buyers**: residency usually isn't the constraint — latency from your region is. Different problem; tell us and we'll plan for it. --- ## Try it, then read the data map - **[Try the live demo](https://intermind.com/demo)** — run the live, multilingual pipeline on your own audio and hear what "EU-runtime *and* multilingual" actually feels like. - [InterMIND vs. Zoom](https://intermind.com/compare/zoom) — feature-by-feature, honestly. - [*Where one InterMIND meeting actually runs*](https://intermind.com/blog/where-one-intermind-meeting-actually-runs) — the full vendor map, gap included. "GDPR-compliant video conferencing" is a checklist, not a badge. Whichever tool you pick — ours or another — make the vendor answer all seven questions in writing. The ones that can are worth your time; the ones that can't are selling you a homepage. — The Mind.com Team --- *Sources on Zoom: [Zoom — GDPR](https://www.zoom.com/en/trust/gdpr/){rel=""nofollow""}, [Zoom — EU Data Residency / privacy in Europe](https://www.zoom.com/en/blog/zoom-privacy-europe/){rel=""nofollow""}, [Zoom — ISO 27701](https://www.zoom.com/en/trust/legal-compliance/iso-27701/){rel=""nofollow""}, checked August 2026. Vendor offerings and tiers change; verify the current configuration against Zoom's trust pages. InterMIND runtime facts are verified against the live deployment as described in the linked data-map post.* # Give every guest their language Think about what a good host does. They don't hand you a chair kit and an Allen key. They don't serve one dish and post the recipe for anyone who can't eat it. Whatever the guest needs to be fully present — a seat, a plate, a place at the table — the host provides it, for everyone, before anyone has to ask. Now look at how meetings handle language. The default meeting is furnished in exactly one language — usually the organizer's, or the corporate house language, which is to say English. Everyone else is welcome, of course. They're just expected to bring their own furniture: follow at speed, formulate replies under pressure, chase nuance in real time, and do it all silently, because the one thing worse than working in your second language is announcing that you're working in it. A meeting is precisely the room where professionals hide that effort hardest. The fix the industry keeps reaching for is personal equipment: an app *you* switch on, captions *you* enable, a feature marketed to the person who "needs help." That framing fails socially even when it works technically, because it asks one participant to raise a hand and mark themselves. Assistive tools with a public stigma get quietly declined — that's why the same function succeeds or fails depending on its symbolism. Nobody wants to be the person visibly compensating. There is an older, better model, and everyone already respects it: **the interpreter booth**. When the UN seats delegates, or any serious international body convenes, interpretation isn't an accessory some delegate switches on in shame. It's *infrastructure the host provides* — a mark of the occasion's seriousness and the host's respect. No delegate is "using assistance." The room simply speaks your language, because you were invited properly. ## Hosting is the fix That model — translation as something the **host provides for everyone**, not something a participant requests for themselves — is the entire idea behind InterMIND, and it changes three things at once. **It changes the social meaning.** When the room itself runs in every participant's language, using yours carries no signal. Nobody switched anything on; nobody is compensating. Language becomes like the lighting — provided, universal, unremarkable. And the norm it establishes is honest, because it applies to natives too: **everyone is smarter in their native language.** The English native who never has to spend cycles decoding accented, compressed second-language English gets sharper counterparts and loses nothing. Subtitles, after all, are watched mostly by people who hear fine. **It changes who benefits.** Here's the inversion that surprises people: speaking your own language is a *courtesy to the listener*. Formulate in the language you think in, and your counterpart — hearing you in theirs — receives your actual reasoning instead of the simplified version that survives translation-in-your-head. You speak generously; they lose nothing of what you meant. We've written about how much gets lost the other way: [the false fluency trap](https://intermind.com/blog/false-fluency-trap). **It changes the economics.** Hospitality means the guest never sees a bill. In InterMIND, the host's plan covers the room: [participants pay nothing and install nothing](https://intermind.com/blog/free-for-participants) — they click a link, pick their language once, and the meeting simply happens in it. Voice, [chat](https://intermind.com/features/multilingual-chat), notes, [documents](https://intermind.com/features/document-translation). One decision by one person furnishes the room for everyone. (That's also, practically, why this model spreads: you don't have to convince every participant — you have to be a good host once, and everyone at the table experiences it.) ## The hospitality extends past the meeting A dinner ends when the guests leave. A working relationship doesn't — and this is where hosting-in-your-language stops being a nicety and becomes structural. In InterMIND the meeting persists as a channel: the chat and its history readable in each member's language, the [recap](https://intermind.com/features/recap) delivered to every member in their own language after the call, the documents in the same space. Your guest wasn't accommodated for an hour; they became **a member who reads the record in their language forever**. Next week's dispute, next month's onboarding, the decision someone cites in October — all of it exists in their language too, not just in the host's. We measured the industry against exactly this bar: [how much of a meeting still exists in your language a week later](https://intermind.com/blog/the-meeting-a-week-later). That's the difference between translating a *moment* and furnishing a *space*. Moments expire. A space you provided keeps being provided. ## What this asks of you Almost nothing — which is the point. Hosting multilingually isn't a program you roll out; it's one habit: 1. **Create the room and send the link.** Guests join from a browser, no account, free. 2. **Let people pick their language** — including you. Say what you actually think, in the language you think it in. 3. **Let the space do the follow-through.** Everyone gets the recap in their language; the history stays readable in theirs. The counterpart who used to nod along now argues back. That's what you invited them for. ## Host one meeting this way - **[Try the live demo](https://intermind.com/demo)** — no signup; hear four perspectives of one room, each in its own language. - **[Start a meeting](https://intermind.com/meetings)** — send one link, give every guest their language. - **[See the benchmark](https://intermind.com/benchmark)** — the translation quality you'd be offering your guests, measured monthly, per language pair, in public. --- ## FAQ **What does it mean to host a multilingual meeting?** It means the organizer provides translation for the whole room — the way venues provide chairs and conference hosts provide interpreters — instead of leaving each participant to arrange their own comprehension. In InterMIND the host creates the meeting; every guest joins by link, picks a language, and hears, reads and later rereads everything in it. **Do my guests need an account or a paid plan?** No. Guests join by link in a browser, need no account, and pay nothing — the host's plan covers the room. Each guest picks their own language independently; the recap arrives in it afterwards. **Why provide translation for participants who speak good English?** Because the room gets smarter, not just kinder. People reason, argue and catch nuance best in the language they think in — natives included, since they stop decoding everyone else's second-language English. Provided-for-everyone is also what removes the stigma: when translation is the room's default, using your language signals exactly nothing. --- # Google Meet live translation: how it works, and where it stops If you searched "Google Meet live translation," the honest short answer is: **yes, Meet can translate a live meeting — two different ways — and which one you get depends on your plan.** Both are real and both are good at a specific job. Neither is built for a room where four nationalities each need to hear the meeting in their own language at the same time. This post explains how each works, how to turn it on, and exactly where the ceiling is — so you can tell which side of it your meeting lands on. > This is the platform how-to companion to our foundational guide, [*Real-time meeting translation: how it works, and how to evaluate one*](https://intermind.com/blog/real-time-meeting-translation). If you want the category-wide version of "how should this work," start there. For the Zoom and Microsoft Teams versions of this post, see [Zoom](https://intermind.com/blog/zoom-live-translation) and [Teams](https://intermind.com/blog/teams-live-translation). New to interpretation as a category? Start with [*Simultaneous interpretation: booth, RSI, or AI*](https://intermind.com/blog/simultaneous-interpretation-guide). --- ## The two features, and why they're not the same thing Google Meet has shipped live translation in two distinct forms. People conflate them constantly, so separate them first: ### 1. Translated captions (text) The older feature. The speaker talks; Meet shows **on-screen caption text translated into another language**. It's a reading experience — subtitles, not audio. Available on **Business Standard, Business Plus, Enterprise Standard, and Enterprise Plus** Workspace editions. Useful when one person presents and others follow along in text, or for accessibility. It does not change what anyone *hears*. ### 2. Speech translation (audio) The newer, Gemini-powered feature, rolled out through 2025–2026. This one translates the **spoken audio itself**, in near real-time, "in a voice like yours" — so a listener hears the speaker's words in their own language, in a voice resembling the speaker's. This is the headline feature, and it's genuinely impressive. It's also where the limits live, so it's worth being precise about them. --- ## How to turn on speech translation in Google Meet 1. You (and the feature) must be on a qualifying plan: **Google AI Pro or Ultra**, or Workspace **Business Standard/Plus**, **Enterprise Standard/Plus**, **Frontline Plus**, or **AI Pro for Education**. On Workspace, an admin may need to enable it. 2. In a meeting, open the captions/translation controls and turn on **speech translation**, then pick the language pair. 3. Wait for the translation indicator on the speaker's tile to clear before you respond — there's a built-in delay of a few seconds for the translation to complete. That's the whole setup. The friction isn't the toggle; it's what the toggle can and can't do. --- ## The limits that actually decide if it fits This is the part the feature announcements don't lead with. As of August 2026, Meet's speech translation is, by Google's own documentation: - **English-anchored.** Translation runs **between English and French, German, Italian, Portuguese, or Spanish** — and that's the full list. Every pair includes English. There is no German↔Japanese, no Spanish↔Arabic, no path that doesn't route through English. If your meeting's languages aren't "English plus one of those five," speech translation doesn't cover it. - **One language pair per meeting.** The whole meeting runs on a single pair at a time. A room with an English, a Spanish, and a German speaker can't have each person hear their own language simultaneously — you pick one pair, and that's the meeting. - **Beta, with hard caps.** It's a beta feature: **90 minutes per meeting**, not available in live streams or recordings, participants under 18 can't use it, and meeting-room hardware can only listen, not translate. - **A few seconds of delay.** Google tells users to wait for the translation symbol to disappear before replying. That's fine for taking turns; it's not the sub-second, talk-over-able latency a fast back-and-forth needs. On privacy, Google is clear and good here: no audio is saved, and no models are trained on your voice. None of this makes the feature bad. It makes it **a bilingual, English-anchored tool** — built for an English speaker and a Spanish speaker taking turns, not for a genuinely multilingual room. --- ## The structural ceiling, in one sentence **Google Meet translates *a pair of languages, anchored on English, one at a time.* A genuinely multilingual meeting needs *every participant in their own language, at once.*** Those are different architectures, not different settings — which is why "more languages" can't be a future toggle on the pair-at-a-time model; it's a different way of building the room. If your meetings are "English plus one other language, two parties, under 90 minutes," Meet's speech translation may be all you need, and it's built into a tool you already have. If your meetings have three or more languages in the room — or two that aren't English — you've hit the ceiling, and the question becomes which tool was built for the other side of it. --- ## When you've outgrown the pair-at-a-time model This is the job we built InterMIND for: not bilingual-with-English, but **every participant in their own language, simultaneously, live.** Concretely, where Meet's speech translation does one English-anchored pair, InterMIND does: - **24 languages live**, on voice, chat, and shared notes — with no requirement that English be one of them. A French speaker and a Japanese speaker can meet without anyone touching English. ([The full end-to-end translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Per-listener audio.** Each participant hears the meeting in their own picked language at the same time — five people, five languages, one room. (How that actually works under the hood is in [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **A sub-second latency budget**, because the design target is a conversation that flows, not turn-taking around a translation indicator. - **Document, chat, and notes translation too** — 30 languages for files dropped in chat — not just spoken audio. - **Quality you can audit.** We publish per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark), rather than asking you to take a launch post's word for it. We're not claiming Meet is bad — for its bilingual case, voice-preserving speech translation is a strong feature. We're claiming it's a different shape of meeting. The honest comparison is in [*Real-time meeting translation*](https://intermind.com/blog/real-time-meeting-translation) and across our [platform comparisons](https://intermind.com/compare). --- ## Try the other side of the ceiling - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and hear per-listener translation instead of one English-anchored pair. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, with the methodology written down. - [InterMIND vs. the platforms you know](https://intermind.com/compare) — honest, feature-by-feature. - Running a webinar or training across languages? The event workflow: [events & webinars](https://intermind.com/use-case/events). - Just need a record, not live translation? The notetaker head-to-head: [Otter vs Fireflies](https://intermind.com/blog/otter-vs-fireflies). Google Meet live translation is a real, useful feature with a clearly drawn edge. Knowing where that edge is — English-anchored, one pair, 90 minutes — is the whole decision. --- ## FAQ **Does Google Meet translate audio or only captions?** Both, on different plans. Translated captions show the speech as text in your language (Business Standard/Plus, Enterprise Standard/Plus). Speech translation — the Gemini-powered feature — translates the spoken audio itself, in a voice resembling the speaker's, on Google AI Pro/Ultra and qualifying Workspace editions. **Which languages does Google Meet speech translation support?** As of August 2026: between English and French, German, Italian, Portuguese, or Spanish. Every pair includes English, and a meeting runs on one language pair at a time. **Is Google Meet speech translation free?** No. It requires Google AI Pro or Ultra, or a qualifying Workspace edition (Business Standard/Plus, Enterprise Standard/Plus, Frontline Plus, or AI Pro for Education), and on Workspace an admin may need to enable it. **How long can a translated Google Meet meeting run?** Speech translation is a beta capped at 90 minutes per meeting. It also doesn't work in live streams or recordings, and participants under 18 can't use it. — The Mind.com Team --- *Sources: [Google Meet — Learn about Speech Translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""}, [Google Workspace — Gemini updates](https://blog.google/products/workspace/google-workspace-gemini-may-2025-updates/){rel=""nofollow""}, [Google Meet — Use translated captions](https://support.google.com/meet/answer/10964115){rel=""nofollow""}. Google expands supported languages over time; check the help pages for the current list. All facts checked August 2026.* # Hindi to English voice translator — and English to Hindi — live, in a real conversation Searching for a **Hindi to English voice translator** usually means one of two very different situations. Either you need one phrase spoken aloud — a tourist moment, a quick exchange — or you need a **conversation** to work across the two languages: a call with a client who speaks only English, an interview, an online lesson, a parent on the other side of the arrangement. The first problem was solved years ago. The second is where most tools quietly give up — and it's the one this guide is really about. One more thing the search phrase hides: a real conversation is never one-directional. "Hindi to English" and "**English to Hindi**" are the same need — you speak Hindi and they hear English, they answer in English and you hear Hindi. Any tool that only relays one direction at a time turns the conversation into a queue. --- ## The three kinds of Hindi–English voice translator ### Phone phrase apps Google Translate's conversation mode, Apple Translate, and dozens of "speak and translate" apps. Tap the mic, speak Hindi, it speaks English — then pass the phone over. Genuinely good for one sentence at a time, and most are free. The mechanic is a relay: speak → transcribe → translate → speak back, one direction at a time, one shared device. For a menu or an auto stand, perfect. For a discussion, the relay is the bottleneck — you operate a translator in turns instead of talking. ### Earbud translators AirPods Live Translation, Pixel Buds, and dedicated translator earbuds. You hear the other person translated in your ear — nicer than a screen, but **one-to-one, in person, and usually needing matching gear on both sides**. Built for a traveler and a host standing face to face, not for a call. ### Live meeting translation A different architecture: the *call itself* is bilingual. You speak Hindi and hear Hindi; they speak English and hear English — **both directions at once, continuously, under a second behind**. No phone passing, no turn-taking, and it works for three or ten participants as well as for two. --- ## Which one do you need? - **One phrase, in person** — a phone app. Don't over-buy; the free ones are good. - **Face to face, both people equipped** — earbuds are the nicest experience. - **A call, a meeting, an interview, a lesson — any real back-and-forth** — live meeting translation. The other two kinds technically "work" and quietly double the length of every exchange. The Hindi–English pair lands in the third case more often than most: the common scenario isn't a tourist phrase but **work and family across the two languages** — a client call where your English-speaking counterpart has no Hindi, a remote interview, an online class, a household video call spanning two countries. --- ## Where InterMIND fits InterMIND is the third kind: live meeting translation, with Hindi ↔ English as one live pair among many. - **Hindi and English, both directions, live** — you speak Hindi, each listener hears their own picked language; they answer in English, you hear Hindi. Continuous, per-listener, sub-second. - **24 languages on voice, chat and shared notes** — so a three-way call can run Hindi ↔ English ↔ Arabic or any other mix, nothing forcibly routed through English. ([The full translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Guests join from one link** — browser, no signup. Send the link to your counterpart; their side needs no setup at all. - **Web, desktop, iOS and Android** — the phone is a full meeting participant, not a shared device passed across the table. - **Free meetings up to 60 minutes** — enough for a real call, not a token trial. - **Quality published, not promised** — per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark), methodology included. Check how Hindi ↔ English stands before you rely on it. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, Hindi to English or English to Hindi, and *hear* the translation instead of reading promises about it. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic. - Start with the category guide: [*Real-time meeting translation*](https://intermind.com/blog/real-time-meeting-translation). - Comparing specific language pairs? See the voice-pair guides for [English ↔ Italian](https://intermind.com/blog/traduttore-inglese-italiano-vocale), [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale) and [Romanian ↔ Italian](https://intermind.com/blog/traduttore-rumeno-italiano-vocale) — or the [Arabic voice translator](https://intermind.com/blog/mutarjim-sawti) guide. — The Mind.com Team --- *Sources: [Google Translate — Live translate](https://support.google.com/translate/answer/6142474){rel=""nofollow""}, [Apple — Live Translation with AirPods](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google — Translate with Pixel Buds](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. Vendors change features and language lists over time — check their pages for the current state.* # "How many languages do you support?" — and why our honest answer is six numbers, not one It is the first question on every procurement call and the first question in every demo. We get it three times a week: > "How many languages does InterMIND support?" The honest answer is: **it depends on which surface you mean.** We have six of them, and they have different language lists for good reasons. This page is the definitive answer so we stop giving inconsistent ones. --- ## The short version | Surface | Languages | | -------------------------------------------------------------------------------------------- | --------: | | Platform API (`createParticipant` accepts) | **24** | | Real-time voice translation in meetings | **24** | | Real-time chat-message translation | **24** | | Real-time shared-notes translation | **24** | | On-demand file translation in chat (DeepL Document API) | **30** | | Website UI (localized at [intermind.com](https://intermind.com){rel=""nofollow""}) | **17** | > **The product is 24 languages.** Files: **30**. Site: **17**. API engine: **24**. Voice, chat, and notes share the same 24 — the full engine list. Files are wider because the file pipeline can be. The website is narrower because translating a marketing site is a different cost than translating in-product copy. Everything below is why those numbers are not the same. --- ## Why there isn't a single number A translation platform is not one product. It is at least three: 1. **A real-time speech pipeline** — audio in, translated audio out under one second, in a meeting where the latency budget is brutal. 2. **A real-time text pipeline** — chat messages and shared notes, character-by-character, propagated to every viewer in their language. 3. **An asynchronous document pipeline** — a 40-page PDF dropped into chat, translated as a whole file with formatting and structure preserved. Each pipeline has its own engine, its own latency budget, and its own per-pair quality envelope. A language that is production-grade for documents (slow, careful, many passes) can be unusable for live voice (fast, one shot, no retries). Treating those as "one language count" is what produces marketing pages that say "200+ languages" and tell you nothing. So we publish per-surface. --- ## 1. Platform API — 24 languages This is the widest list, because it is the engine surface. The [`createParticipant` endpoint](https://api.mind.com/#operation/createParticipant){rel=""nofollow""} accepts these 24 codes: `ar, cs, da, de, en, es, fi, fr, hi, hu, is, it, ja, ko, nl, no, pl, pt, ro, ru, sv, tr, uk, zh` If you are an integrator building on Mind directly, you can hand any of those 24 to a participant. That **does not** mean every one of those 24 scores the same on our quality bar. It means our engine can emit text in that language. The product shows all 24 and publishes each pair's quality openly (next section). --- ## 2. Voice, chat, and notes inside the product — 24 languages When a user joins an InterMIND meeting, the language picker shows all **24** engine languages — including **Arabic (`ar`)** and **Hindi (`hi`)**, which we previously held back. We used to remove those two, because some of their pairs don't clear the median-quality bar we publish on [`/benchmark`](https://intermind.com/benchmark). We've changed the trade-off: instead of hiding a language wholesale, we surface it and let the public per-pair numbers tell the truth. Every pair — strong or weak — is measured on real traffic and deep-linkable by URL, so a user (or an auditor) can check `en→ar` before they rely on it rather than discover its quality mid-meeting. This is still the honest version, just tuned. Most of the category lists everything the model can emit and stays silent on quality; the opposite extreme (hide anything below a bar) makes a real, in-demand language invisible. We do neither: show the full list, publish the number next to it. If a pair isn't good enough for your use case, the benchmark says so out loud. The same 24-language list applies across the three real-time surfaces: - **Voice translation** — per-viewer translated audio with sub-second latency. Each participant picks their own listening language at the start of the call. - **Chat messages** — every message translated as it is typed; edits produce per-language diffs ([see v1.2](https://intermind.com/blog/intermind-v1-2-release)). - **Shared notes** — character-by-character live translation of the host's notes pane, per viewer, with diff history. One picker, one set of 24 languages, three places it shows up. --- ## 3. On-demand file translation — 30 languages Drop a PDF, DOCX, PPTX, or XLSX into the chat. Each participant can request the file in their language. The translated copy is returned as the same file format with structure preserved (tables, headings, lists). This surface uses the **DeepL Document API**; our file pipeline maps **30 target languages** — wider than our real-time voice pipeline. If you can translate the PDF to Estonian on DeepL today, you can translate it to Estonian in InterMIND chat today. The file list includes some languages our real-time pipeline does not — for example **Bulgarian, Greek, Estonian, Indonesian, Lithuanian, Latvian, Slovak, Slovenian**. Both **Arabic** and **Hindi** are now on the real-time voice/chat/notes list as well as, for Arabic, the file list; a French participant can both listen to the meeting in Arabic and request the contract PDF in Arabic. The one real-time language **not** on the file list: **Hindi.** Our document pipeline doesn't map it yet, so Hindi is available for live voice, chat, and notes but not for on-demand file translation — the reverse of the old Arabic asymmetry. We flag this in the file picker rather than hide it. Why the file list is bigger in general: - Documents are **asynchronous**. There is no one-second budget, so the pipeline can afford a slower, more careful engine that handles more pairs well. - DeepL is a dedicated document-translation engine, and we use it directly for files. We do not try to route documents through the same engine that runs voice. Different problem, different tool. --- ## 4. Website UI — 17 languages [intermind.com](https://intermind.com){rel=""nofollow""} is currently shipped in **17 locales**: English, German, Spanish (Latin America), French, Italian, Portuguese (Brazilian), Dutch, Polish, Ukrainian, Chinese (Simplified), Russian, Japanese, Korean, Turkish, Arabic, Indonesian, and Vietnamese. Two of those — **Indonesian and Vietnamese** — are localized on the public site but are **not** yet available for in-meeting translation: our speech engine doesn't cover them. So you can read the landing pages, blog, and docs in Indonesian, but if you join a meeting, voice and chat translation won't run in Indonesian yet. We flag this in the language switcher rather than let you discover it in a live call. The other fourteen locales are full products — site and in-meeting both. Why narrower than the file list: - A marketing site has its own translation cost — and its own quality bar. Landing pages, FAQs, pricing, legal copy. It is not free. - We ship locales where traffic and pipeline justify the maintenance, not for every language the product supports. - The site list leads the product in two places (id, vi) on purpose: localizing the public surface is a standalone capture layer even before the speech engine reaches a language. This list will grow as pipeline grows. It is not capped on principle. It is capped on cost. --- ## The pattern, and how to read it on any vendor page Every multilingual vendor has these same surfaces. Most of them quote you the widest one and let you discover the others on your own. The honest version is: - The **engine** list is widest, because it is what the model emits. - The **product** list is narrower, because not every engine output meets a product bar. - The **document** list often differs from the **voice** list, because the engines are different. - The **website** list is narrowest, because translating your own site costs you cash. When you evaluate any platform in this category, ask for the four numbers separately. If a vendor refuses to split them — or doesn't know — you are not getting a useful answer. --- ## What we will not pretend - **Some pairs are weaker than others, and we say so.** Arabic and Hindi are back in the voice/chat/notes picker, but not every pair scores the same. Rather than hide a whole language, we publish per-pair quality on [`/benchmark`](https://intermind.com/benchmark) — a weak pair like `en→ar` shows its real number instead of being quietly dropped. Check the pair you need before you rely on it. - **The file list (30) being wider than the voice list (24) is a real gap, not a feature.** A French user can request an Estonian translation of the PDF but cannot listen to the meeting in Estonian. We will not paper over this by quoting the bigger number and hoping you don't notice. - **The site leading the product in Indonesian and Vietnamese is a gap, not a feature.** You can read the site in those two but not yet run a meeting in them. We flag it in the switcher instead of letting you find out mid-call. We close it when the speech engine adds them. --- ## Try it yourself - **[Try the live demo](https://intermind.com/demo)** — runs the production voice + chat pipeline against your audio, in any of the 24 product languages. The same pipeline that scores [`/benchmark`](https://intermind.com/benchmark). - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month quality on real traffic. Every pair in the picker, strong or weak, deep-linkable by URL. - **[Read the methodology](https://intermind.com/benchmark/methodology)** — what the numbers are, what they aren't, who the judge is. Six numbers, six surfaces, one engine. That is the honest answer to "how many languages do you support." --- ## FAQ **How many languages does InterMIND translate live?** 24 — the same list across voice, chat, and shared notes: `ar, cs, da, de, en, es, fi, fr, hi, hu, is, it, ja, ko, nl, no, pl, pt, ro, ru, sv, tr, uk, zh`. Every pair's quality is published per month at [`/benchmark`](https://intermind.com/benchmark). **How many languages does document translation support?** 30, via the DeepL Document API — drop a PDF, DOCX, PPTX, or XLSX into chat and request it in your language with formatting preserved. The file list is wider than the live list (it adds e.g. Bulgarian, Greek, Estonian) but misses Hindi. **Why is the website in fewer languages than the product?** The site ships in 17 locales because localizing a marketing site is a separate cost from in-product translation. Two site locales (Indonesian, Vietnamese) lead the product on purpose — flagged in the language switcher. — The Mind.com Team --- *Sources: [DeepL — supported languages](https://developers.deepl.com/docs/getting-started/supported-languages){rel=""nofollow""}, [Mind API — createParticipant](https://api.mind.com/#operation/createParticipant){rel=""nofollow""}; product surface lists verified against the shipped code, checked August 2026.* # Inside the four translation pipelines that run InterMIND The old `/product/overview/how-it-works` page on mind.com is several major releases out of date. It describes a single "translation engine" the way most vendor pages do — one big arrow from "you speak" to "they hear." That picture was already a simplification two years ago. Today it is wrong. The truth is that InterMIND runs **four separate translation pipelines**, each solving a different problem with a different engine, a different latency budget, and a different quality envelope. They share a language picker. They do not share an engine. This is the updated answer to "how does it work." > **A companion piece:** [*"How many languages do you support?"*](https://intermind.com/blog/how-many-languages-do-you-support) covers what each pipeline *covers* (24 / 24 / 30 / 17). This post covers what each pipeline *does* — and why it is its own thing. --- ## Why "one engine for everything" is a lie A live meeting platform has at least four jobs to do at once, and they pull in incompatible directions: 1. **Real-time voice** — audio in, translated audio out, under one second, every viewer in their own language. The hard constraint is latency. 2. **Real-time chat text** — short messages, fast, with edits and quotes and HTML structure preserved. 3. **Real-time shared notes** — character-by-character collaborative typing, with structural hierarchy (lists, headings, checkboxes) that has to survive translation. 4. **Asynchronous document files** — a 40-page PDF dropped into chat. No latency budget. The hard constraint is *fidelity* — formatting, tables, page numbers, font. You can build one giant LLM call that tries to do all four. We tried. It is bad at all four. The latency budget for voice means the model can't think; the fidelity budget for documents means the model has to. A chat edit needs a diff in the viewer's language; a 40-page PDF needs format preservation that no token-streaming model gives you. So we run four. Here is each one. --- ## Pipeline 1: Real-time voice translation **The problem:** A participant speaks French. Another participant joined in German, a third in Brazilian Portuguese, a fourth in Japanese. Each one needs to hear the speaker in their own language, in their own ear, with a delay short enough to keep eye contact possible. **The budget:** Sub-second end-to-end. Anything past \~1.2 seconds and the conversation breaks — people start talking over the translation, and the meeting drifts toward "let's just switch to English." ### How the audio actually moves ![Voice translation pipeline: the speaker's browser does ASR locally via Mind SDK, ws-server fans the transcript out to the translation engine over one WebSocket per target language present in the room, and each viewer receives their own translated audio track.](https://intermind.com/blog/inside-the-translation-pipelines-voice.svg) A few things worth naming explicitly: - **ASR runs in the speaker's browser**, not on a central server. We use the Mind SDK locally; this saves a round-trip and gives us the source-language transcript with the lowest possible delay before translation can even start. - **Translation is not one fan-out.** We hold a pool of WebSocket connections to our translation engine, **one per target language present in the room**. If three participants picked German, German shares one connection. If nobody picked Arabic, no Arabic connection is opened. The pool drops idle connections after five minutes. This is why a four-language meeting costs the same as a forty-language meeting up to the point of who actually showed up — we never translate to languages no participant is listening in. - **Synthesized speech is per-viewer.** Each participant receives their own translated audio track, mixed against the original speaker's video. They are not watching a master "translated meeting" — they are watching the *same meeting*, with their personal audio channel translated to their picked language. This is why two people in the same physical room can each plug in headphones and hear different languages. ### Why this matters when a meeting goes sideways In a 60-minute call with eight languages, things break in interesting ways: WebSockets drop, ASR temporarily mis-transcribes a proper noun, one participant's network gets jittery. The architecture above is what lets us isolate failures: one viewer's audio glitching does not affect the other seven, because the translation engine never produced "the translation" in the first place — it produced eight, in parallel, and only the affected one has to recover. The engine itself is ours, hosted on our own infrastructure. We do not route real-time voice through third-party general-purpose LLMs. The latency budget rules them out; the data-residency story rules them out for the regulated customers who actually care. > **What we publish about voice quality:** [/benchmark](https://intermind.com/benchmark) runs the production voice pipeline against [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""} sentences for every published language pair, monthly. The judge is named (Gemini 2.5 Flash primary, Claude Sonnet 4 fallback). The full distribution — median, p10, p90, min, max, sample size — is on the page. See [the methodology](https://intermind.com/benchmark/methodology) for what those numbers do and don't measure. --- ## Pipeline 2: Real-time chat translation **The problem:** Every chat message in the meeting, translated for every participant in their own language, as it is sent. Plus edits — and edits need to look like edits, not like re-translations. **The budget:** Fast, but not sub-second. A chat message can take half a second to appear in another language without anyone caring. What people care about is whether the translation is *right* and whether edits make sense. ### What the chat pipeline actually does Each message goes through the same translation engine the voice pipeline uses — but with different pre- and post-processing: - **HTML structure is preserved.** Chat supports rich text (paragraphs, lists, quotes, bold, italic). We convert to plain text for the model, translate, then re-wrap the result in the original tags. The model never sees the HTML — it sees clean prose. - **Quotes are translated independently.** If you reply to a message and quote it, the `[QUOTE]…[/QUOTE]` block and the new content are translated as separate units, so the model can't confuse the two. - **Long messages get chunked.** We split on paragraph boundaries at 1,000 characters per chunk. Each chunk is its own translation call. We do *not* feed 4,000-character novels to the model in one shot — the failure modes (truncation, lost paragraphs, mid-sentence cut-offs) are too ugly. - **Translation is lazy.** We use an IntersectionObserver: a message is only translated when it scrolls into the viewer's viewport. Switching languages in a long-running channel used to replay every translation API call from the history. Now it doesn't. ### The interesting part: edits as diffs In v1.2 we changed how chat edits behave for viewers in another language. The old behavior was: someone edits a message, we re-translate the whole thing, you see a fresh paragraph and have to spot what moved. The new behavior: 1. The original message was already translated to your language. 2. When the sender edits, we re-translate the *new* version. 3. We compute the diff between **your previous translation** and **your new translation**, in your language. 4. We show that diff inline — same way Git shows you what changed. So when "review by Tuesday" becomes "review by Thursday" in English, your Spanish-reading colleague sees **martes → jueves** highlighted, not a re-translated paragraph they have to re-read. This required treating the chat pipeline as a *stateful* per-viewer cache, not a stateless translate-on-request endpoint. Documents and voice don't need this. Chat does. --- ## Pipeline 3: Real-time shared-notes translation **The problem:** The host opens a shared-notes pane and starts typing. Every participant sees the notes in their language, character-by-character, with the structure of the document — headings, nested lists, checklists, code blocks — intact. **The budget:** Same as chat (\~half a second), but with two extra constraints: - **The thing being translated changes mid-translation.** The host is still typing. A naive system that translates "the whole document" each keystroke produces flicker and burns the API budget. We translate at the granularity of the *changed unit*, not the whole document. - **Structure must survive.** If you ask a translation model to translate a markdown blob with three nested lists, you get back something that *looks* like the original but with subtly flattened hierarchy, renumbered items, or moved indentation. We do not let the model see the whole blob. ### How the notes pipeline differs from chat The structural preservation is the main thing. We translate **each list item independently** rather than as one document. The model sees: > "Compliance review — Q2 deliverables" — not: > "# Project plan\n## Quarter\n- Compliance review — Q2 deliverables\n- Vendor scoring\n - Tier 1 vendors..." The wrapping document — the `
    `, the headings, the indentation — is rebuilt on the client side using the same structure the original document had, with each leaf node swapped for its translation. The model never gets to "improve" the hierarchy. Notes also use the same per-viewer diff model as chat edits: if the host changes a line, viewers in other languages see the changed words highlighted, not a fresh paragraph. --- ## Pipeline 4: Asynchronous document translation **The problem:** Someone drops a 40-page PDF, a Word doc, a PowerPoint deck, or an Excel sheet into chat. Each participant can request a copy in their own language. The translated file must look like the original — same fonts, same tables, same page numbers, same headers, same charts in place. **The budget:** No real-time constraint. A minute is fine. Two minutes is fine. The constraint is **fidelity** — if the translated PDF doesn't look like the original, the recipient won't trust it. ### Why this pipeline does not share an engine with voice A general LLM, even a very good one, will hand you back a translated *text* of a document. It will not hand you back a translated *PDF* with the same layout. The model has no concept of "page break that has to line up with the source" or "table cell that has to keep its column width." For this surface we use the **DeepL Document API** directly. It is purpose-built for translating *files as files*, not *prose extracted from files*. DeepL handles: - PDF (with layout preservation) - DOCX, DOC - PPTX - XLSX The document is uploaded to DeepL's pipeline, translated server-side with formatting intact, and returned as the same format. We then upload the result to our object storage and surface it back in chat as a downloadable attachment. ### What this costs and why we don't hide it DeepL bills a minimum of 50,000 characters per document — roughly one US dollar per file on the Pro tier, regardless of whether the document is one page or thirty. We absorb that cost rather than charging per file; it shows up in the meeting's translation usage as **billed characters**, converted to word-units that match the way the rest of the product reports translation activity. We picked DeepL for this surface because translating *files as files* is exactly the job it was built for — we did not try to build a better one. The same is not true the other way around — DeepL does not run a live-voice pipeline of the kind we built for meetings. Different problems; different tools. The honest version of "what powers InterMIND translation" is "the right engine per pipeline" — not "our engine, everywhere." ### Languages this pipeline covers that voice does not The document pipeline reaches **30 languages**, vs. 24 for voice. The extras include: Bulgarian, Greek, Estonian, Indonesian, Lithuanian, Latvian, Slovak, Slovenian. (Arabic used to sit in this list while its voice quality was below our bar; it's now in the real-time picker, with its per-pair scores public on [`/benchmark`](https://intermind.com/benchmark) like every other language. The asymmetry now runs the other way for Hindi — live on voice, not yet on files.) That asymmetry is real. It means a French participant in a meeting can request the contract PDF in Estonian even though they cannot listen to the meeting in Estonian. We flag it in the picker rather than smooth it over with a single number. The reasoning is in the [language-count post](https://intermind.com/blog/how-many-languages-do-you-support). --- ## Where the pipelines meet The four pipelines do not run in isolation. A meeting room is where they touch each other, and the seams matter: - **A chat message with a document attachment** triggers the chat pipeline for the text and the document pipeline for the file. The participant in another language sees the message translated immediately and the attachment translation arriving asynchronously as a downloadable. - **A shared note that quotes a transcript line** crosses notes ↔ voice. The transcript is what the voice pipeline produced for the sender's language; the note translation produces a per-viewer copy of that quote in everyone else's language, with its source attribution preserved. - **A transcript exported after the meeting** runs the chat-style text pipeline over the full conversation, producing a per-language file that participants can download. This is the same code path as chat translation, just batched. The language picker is one piece of UI. The infrastructure underneath is four pipelines, talking to each other. --- ## What we deliberately do not try - **No "unified translation model."** We are not building one model that does voice, chat, notes, and documents. The latency vs. fidelity trade-off doesn't have a winner. We use the right engine per surface. - **No silent re-routing.** If the file pipeline can't translate to Hindi today, we don't quietly fall back to the voice engine and pretend it worked — the file picker flags the gap instead of hiding it. - **No "we translate to 200 languages."** Our engine emits 24. The live surfaces ship all 24, documents 30 — and instead of one marketing-friendly number, the per-pair quality that has to stand in front of an auditor is published at [`/benchmark`](https://intermind.com/benchmark), weaker pairs included. --- ## Try it yourself - **[Try the live demo](https://intermind.com/demo)** — runs the live voice pipeline against your audio, in any of the 24 product languages. The same pipeline that scores [`/benchmark`](https://intermind.com/benchmark). - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month quality on real traffic. Every pair in the picker, strong or weak, deep-linkable. - **[Read the methodology](https://intermind.com/benchmark/methodology)** — what the numbers are, what they aren't, who the judge is. Four pipelines, four engines, one meeting room. That is the honest replacement for the old `how-it-works` page. — The Mind.com Team --- *Sources: [DeepL — supported languages](https://developers.deepl.com/docs/getting-started/supported-languages){rel=""nofollow""}, [DeepL — usage count and billing](https://support.deepl.com/hc/en-us/articles/360020685720-Usage-count-and-billing-in-DeepL-API){rel=""nofollow""} (the 50,000-character per-file minimum), [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""}; internal pipeline facts verified against the shipped code, checked August 2026.* # The Meeting Room That Doesn't Switch to English You schedule a cross-border meeting. People join from three time zones. Within ninety seconds someone says **"let's keep this in English so we're all on the same page"** — and you watch the most senior people in the room go quiet. That switch is a tax. On time, on confidence, on what actually gets decided. We've spent the last six weeks rebuilding InterMIND around the principle that **no one in the room should have to switch.** Not for the conversation. Not for the chat. Not for the notes you're typing live. Not for the edit history you'll review tomorrow. Here's what shifted. --- ## 1. Per-Viewer Everything, Not Just Voice Live voice translation has always been the headline — 21 languages, sub-second latency, every participant picks their own language at the start of the call. v1.2 makes everything **around** the voice multilingual too, and keeps it in sync, per viewer, in real time. (For why "21" and not the bigger numbers other surfaces show, see [the breakdown](https://intermind.com/blog/how-many-languages-do-you-support).) ### Chat: every edit shows you what changed, in your language Every message in chat is translated as it's typed. The change in v1.2: when someone edits a message, we don't throw out their previous translation and serve you a fresh one. We compute the diff **between the previous translated text and the new translated text**, in your language — and show it inline. So when your colleague edits "we should review by Tuesday" to "we should review by Thursday," your Spanish reviewer doesn't see a re-translated paragraph and guess what moved. They see *martes → jueves*, highlighted. Same precision as a Git diff, but rendered in whichever of the 21 languages they read. If the edit changes only formatting — bold, italic, list markers — we re-translate anyway, because format-only edits are often where meaning shifts (a comma becomes a clause break). And the previous translation stays on screen while the new one streams in, so you never look at blank space waiting for the API to come back. ### Notes: live, character-by-character, per language The host opens the shared-note pane during the meeting and starts typing. Every participant sees the note in their language, character-by-character, with sub-second propagation. Edit a line and viewers get a per-language diff — same model as chat edits. We also rewrote note translation to **translate list items individually instead of the whole document**. A nested checklist in English doesn't preserve its structure if you translate the markdown blob — the model flattens the hierarchy. Translating each `
  • ` independently keeps the indentation, the numbering, and the inline formatting intact in every language. ### Attachments: documents inside chat speak the room Drop a Google Doc, OneDrive link, or PDF into chat. v1.2 renders them inline: - **Linked docs** — title, preview, domain badge, no raw URL bleeding through. Editable docs open in a new tab with the right locale hint (`hl=`), so a Brazilian Portuguese reader doesn't land in an English Docs UI. - **Image uploads** — thumbnail on upload, intrinsic-grid responsive layout in the message bubble so images don't blow out at narrow widths. - **Three-tier document access** — *edit*, *view*, or *no-access* — with auto-share when permissions allow. Participants without Drive access still see a clean read-only preview, not a "permission denied" page. ### Lazy translation: faster, cheaper, less noisy We used to translate every message in your history when you switched languages. For long-running team channels that meant replaying hours of API calls for messages you'd never scroll back to. v1.2 uses an IntersectionObserver: translations happen on demand, when a message enters the viewport. Switching languages mid-meeting is now instant. We also stop chat-list polling immediately on 401, which silenced a runaway error stream you would never have noticed but we did. ### Voice transcription: sanitized before render Transcription text is sanitized before rendering — no HTML injection through a creative pronunciation. It sounds obvious. It also wasn't true two months ago. --- ## 2. The Note Editor Caught Up With the Conversation Notes and chat messages share a rich editor with: - **Slash menu** for headings, lists, code blocks - **AI Improve** transforms — pick a paragraph, rewrite for tone, expand, summarize, fix grammar — directly in the editor, no copy-paste round trip - **@-mentions** wired into participant data: type `@` and you see the live participant list - **Emoji picker** that doesn't drop out of position on narrow widths - **Image upload** straight into the message body AI Improve transforms run as single-chain inserts with no output-token cap (we hit that ceiling more than once during testing). Output is converted from markdown to HTML via `marked`, so the editor never shows you a raw `### Heading` artifact. A shared note can be opened in a new tab — useful when the meeting is on one monitor and the document needs a full second screen. Sidebar resize no longer gets stuck behind the iframe. --- ## 3. The Meeting Room Finally Has the Boring Parts Right A translation feature is meaningless if the meeting room around it doesn't hold up. v1.2 closes a long list of gaps you only notice when they fail. | Capability | What it does | Why it matters | | -------------------------------- | --------------------------------------------------------------------------- | ------------------------------------------------------------ | | **Host-only permissions** | Mic, camera, screenshare, chat — gated per participant from More → Settings | Moderation surface stops being a free-for-all | | **Conference guide modal** | Keyboard shortcuts and button explanations in one panel | New users find features without a tour | | **Note sharing on stage** | Pin a note next to your camera overlay; camera stays visible | Walk through a doc without losing the room | | **Two-mode read receipts** | *Present* (currently watching) vs *returned* (saw it later) | Beats the binary "delivered" lie | | **Pre-join deferred approval** | Guests see an approval-pending screen before they appear | No surprise joins, no awkward expulsions | | **Ad-hoc copy-link nudge** | Host joins alone? Toast asks if they want to copy the link | Closes the "am I the only one here, was the link sent?" loop | | **Raise-hand fix** | The button works again | It was broken; now it isn't | | **Forced dark theme in meeting** | Dark only while the meeting room is mounted | Less brightness whiplash mid-call | | **1-hour Basic cap** | Free meetings end at 60 min — more headroom than Zoom Free's 40 | Honest pricing surface | Translation runs full quality on every plan, including Free. --- ## 4. Teams That Don't Fight the Tool We rewrote the role and seat model from scratch on a strict **1-user-1-org** invariant. No lazy team creation. No ghost orgs. Owners cancel subscriptions; they don't get asked to transfer to nowhere. | Role | Scope | | ---------- | ---------------------------------------------------- | | **Owner** | Billing, domain management, every action below | | **Admin** | Invite, role changes, license grants — no billing | | **Member** | Meetings, chat, their own profile | | **Solo** | Personal contacts, personal meetings — no team scope | Licenses are assignable per seat, Zoom-style. Bulk operations from the Users table — make admin, make member, grant license, revoke license — apply across selection with a clear warning when only part of the bulk action will fit available seats. License switching is optimistic, no spinner waiting for the server. If you grant a seat and the next click would exceed the plan, a contextual **Buy more licenses** CTA appears exactly where you'd next try to grant one. Domain Management is gated to Business+ and clearly labeled, not silently hidden. Pending invitations clean themselves up. Self-service is real: members can't accidentally delete themselves, can't grant themselves admin, can't bulk-action their own row. Owners see Owner-only UI; members don't see it at all. The guest-access flow now lives outside the team boundary entirely. Hosts can approve a guest into the meeting without adding them to the team; the host's personal contacts get a separate *Add to contacts* option from the participants sidebar. --- ## 5. Honest Cuts, Honest Numbers A release post is also a chance to be straight about what didn't make it. **We disabled Arabic, Turkish, and Hindi for live voice/chat/notes during the alpha cycle.** Translation quality for those three didn't meet our bar at the time. We'd rather ship languages that work than pad the count with footnotes. (Update: all three have since returned to the live picker — voice/chat/notes is 24 languages today, with per-pair quality published on [`/benchmark`](https://intermind.com/benchmark) instead of hidden behind the count. See [the full per-surface breakdown](https://intermind.com/blog/how-many-languages-do-you-support).) The picker names for Spanish, Portuguese, and Chinese now disambiguate region and script, because nobody should have to guess whether "Chinese" means simplified or traditional. **We retired `staging.intermind.com`.** `main` is the only long-lived branch. Production deploys are manual. CI on push to `main` runs lint + unit + integration on a Vercel preview and stops there — we don't auto-promote, which means we don't ship to prod by accident either. **We migrated the translation server from Amsterdam to Paris.** Measurable latency improvement for European users, with retry logic added in the same pass to absorb upstream provider blips. **We collapsed two Sentry projects into one** and split by environment tag. Two projects sounded organized; in practice it doubled the dashboards we had to check before believing anything. **Replay session sampling went down** because the previous rate was burning the entire org quota. The 4xx HTTP errors fall out of capture for the same reason — most were not actionable. We publish a **[public benchmark](https://intermind.com/benchmark)** with translation latency, quality scores, and the methodology. The methodology page includes a disclosure: before 2026-04-23, our chat harness was mis-counting some translation events, which inflated reported latency. We found it, fixed it, kept the bad numbers in the historical chart, and labeled them. That's the kind of disclosure that should be expected from anyone publishing a benchmark. In production we run **hourly synthetic E2E checks** through Checkly — full chat round-trip and full voice round-trip, including an LLM-scored translation-quality assertion — plus **Web Vitals monitoring** via Sentry Metric Alerts on every public page. When something breaks, we know within minutes, not the next morning. --- ## 6. Eight Use Cases — Try Before You Sign Up Telling people "we translate" doesn't land. We built eight use-case pages with animated demos for the scenarios that actually drive cross-border meeting need: 1. **Regulated multilingual meetings** — CAPA reviews, inspection prep, multi-site quality calls. The audit trail is multilingual end-to-end. 2. **Cross-border sales calls** — close in the buyer's language, keep your team's notes in yours. 3. **Multi-site engineering reviews** — design walkthroughs across plants where the technical vocabulary is the hard part, not the small talk. 4. **Remote technical interviews** — assess the engineer, not their English. 5. **Multilingual customer support** — every ticket conversation, in every language, with one rep. 6. **Language tutoring** — instructor speaks naturally; learner gets dual-language subtitles and chat. 7. **Online events and webinars** — one host, many language tracks, no relay interpreters. 8. **Remote consulting** — billable hours that don't get spent translating jargon. The [/use-case](https://intermind.com/use-case) listing page has the full set with CSS-animated demos. The [/demo](https://intermind.com/demo) route lets you experience translation against a scripted conversation — no account, no sign-up. --- ## What Else Shipped Compact list, for completeness: - **Site-wide localization** for ES / PT / RU / ZH with locale-prefix routing and `hreflang` on every public route - **Auto-translate i18n on commit** — we edit `en.json`, the other four locales are translated and re-staged in the same pre-commit hook (Gemini Flash → Sonnet fallback) - **Single translation-language override** — interface language and translation language used to be two settings that diverged; now they're one - **Bug reports with attachments** — on-demand screenshot, conference-button context, email delivery with a Sentry thread ID - **In-app reply to bug reports** — when we reply to your Sentry issue, the message lands back in your feed via an `@reply` convention - **G2 social proof carousel** with curated reviews on the landing page - **Sitemap + robots.txt** generation for SEO; redirected stray legacy URLs - **In-app docs with AI chat** — full knowledge base injected, search, walkable from [/docs](https://intermind.com/docs) - **Adaptive default page** — sidebar reorders based on which area you actually use --- ## Free Still Works The free-for-everyone alpha ran through June 2026; billing is now active. The Basic plan stays free forever — live voice and chat translation included, with a 1-hour meeting cap and no credit card. If you're going to test InterMIND for your team, every paid plan starts with a 14-day trial — you get the full Business experience without paying for it. → [Try the demo](https://intermind.com/demo) — no sign-up → [Create a free account](https://intermind.com/login) — 30 seconds → [See the use cases](https://intermind.com/use-case) — find your scenario — The Mind.com Team **Your team doesn't switch to English in the hallway. The meeting shouldn't make them switch either.** --- *Sources: [Zoom — plans and pricing](https://zoom.us/pricing){rel=""nofollow""} (Basic plan 40-minute meeting cap), checked August 2026. This is a release post — product numbers reflect v1.2 as shipped in May 2026; the current per-surface language counts are in [How many languages do you support?](https://intermind.com/blog/how-many-languages-do-you-support).* # The Language Barrier Ends Here You spend years learning a language. You still can't close a deal in Mandarin. You still need an interpreter for the Tokyo call. You still lose nuance in every cross-border conversation. **What if you didn't have to?** Today we're launching the public alpha of InterMIND — and giving it away for free until June 2026. --- ## Why InterMIND Exists We stopped memorizing phone numbers when smartphones made it unnecessary. We abandoned manual calculations when calculators became ubiquitous. We no longer memorize directions since GPS navigation emerged. Language learning is one of the last inefficient allocations of human cognitive potential. The average person spends 600–1,000 hours to reach basic proficiency in a new language. Fluency requires 2,000+ hours. These are hours that could be invested in your actual expertise. **We believe language shouldn't determine who you can work with.** Geography shouldn't limit your ambitions. Culture shouldn't be a competitive disadvantage. InterMIND eliminates the need to spend years mastering foreign languages for practical communication — just as calculators eliminated mental arithmetic. Learn more about [our mission](https://mind.com/resources/company/about){rel=""nofollow""}. --- ## Not Translation. Not Interpretation. Something New. InterMIND is **conversational telepathy** — you think in English, they hear perfect Mandarin. They respond in Japanese, you understand every nuance. It preserves your voice, your tone, your personality — in any language. It captures context, cultural subtext, business intent. Unlike traditional translation tools that convert words, InterMIND **interprets meaning**, adapts tone, and facilitates seamless multilingual dialogue as if the language barrier didn't exist. > Speak naturally. Be understood perfectly. Close more deals. **How it works:** your speech is recognized, cleaned up (removing filler words, correcting errors), translated with context and cultural awareness, and synthesized back into natural speech — all in under 3 seconds, matching the speed of professional simultaneous interpreters. Learn more about [the technology behind InterMIND](https://mind.com/product/overview/how-it-works){rel=""nofollow""}. --- ## What You Get Today ### Speak and Listen in Your Language Join a video call, pick your language, start talking. Everyone hears the conversation in their native language — **21 languages**, sub-second latency, live subtitles. No plugins, no setup, no interpreters. ([Why "21" and not "24" or "30"](https://intermind.com/blog/how-many-languages-do-you-support).) Chat messages are translated too — automatically, as you type, before you even hit send. ### HD Video Meetings — No Compromises InterMIND is a full video conferencing platform, not a translation add-on. HD video, screen sharing, cloud recording, laser pointer, reactions, hand raise — everything you'd expect, with translation built into the core. Create ad-hoc meetings with shareable links. Invite guests without requiring them to sign up — they get in with just a name, with host approval. ### Team Chat That Breaks Language Walls Persistent channels with every message auto-translated. See the original and translation side by side. Share files up to 500 MB with link previews. Real-time delivery via WebSocket — no refreshing, no waiting. ### Enterprise-Ready from Day One Google Workspace SSO with domain-based auto-join. Microsoft directory integration. Passwordless email sign-in. Role-based access control and bulk invitations. Export your data anytime. Built-in docs with search and AI chat assistant at [/docs](https://intermind.com/docs). --- ## Free During Alpha — No Credit Card Required Every feature. Every plan tier. Zero cost until June 2026. We're building InterMIND in the open and we need your feedback. During the alpha period, all users get unrestricted access to the full platform. See [pricing and plans](https://intermind.com/pricing) for what comes after the trial. --- ## Roadmap | Milestone | Date | What to expect | | ------------------------------- | ----------------- | ------------------------------------------------------- | | **v1.0.0-alpha** (this release) | March 1, 2026 | Full platform — free for all users | | **Extended trial** | Until June 2026 | No charges, no credit card needed | | **v1.0.0-beta** (paid) | June 2026 | Billing activated, new features based on alpha feedback | | **v1.0.0** (GA) | September 1, 2026 | General availability — stable, production-ready release | --- ## Get Started in 30 Seconds 1. **Try the demo** — no sign-up required: [/demo](https://intermind.com/demo) 2. **Create a free account**: [/login](https://intermind.com/login) 3. **Start your first translated meeting** Have feedback? Use the in-app bug report (with screenshot capture) or reach out via support chat. We're not building a translator. **We're building a bridge between worlds.** — The Mind.com Team **Stop pretending you understood that meeting.** Everyone speaks their language. Everyone understands. Instantly. # Interpretation vs translation: what's the difference — and every type of interpretation explained (2026) **Interpretation is the conversion of spoken (or signed) language into another language, live, while the communication is happening. Translation is the conversion of written text from one language to another, done after the text exists.** The American Translators Association compresses it into one line: [translators do the writing, interpreters do the talking](https://www.atanet.org/client-assistance/translator-vs-interpreter/){rel=""nofollow""}. The two words get used interchangeably in everyday speech — "we need a translator for the meeting" almost always means *interpreter* — but they name different professions, different skills, different tools, and different buying decisions. This guide gives you the working definitions, a side-by-side comparison, every type of interpretation with the situations each one fits, and where AI now sits in the picture. --- ## What is interpretation? **Interpretation is real-time language conversion of speech.** An interpreter listens to a speaker in one language and renders the meaning — not the word-for-word text — in another language, either while the speaker is still talking (simultaneous) or in the pauses between their sentences (consecutive). The output is ephemeral: it exists in the moment, for the people in the room or on the call, and is gone when the meeting ends. Because it happens live, interpretation allows no second draft. The interpreter carries everything — vocabulary, tone, cultural context, the speaker's intent — in working memory, at the speaker's pace. That is why the United Nations, where sessions run in [six official languages rendered simultaneously into the other five](https://www.un.org/dgacm/en/content/interpretation){rel=""nofollow""}, treats conference interpreting as one of the most demanding language professions, and why professional simultaneous interpreters [work in pairs and swap roughly every 30 minutes](https://akjournals.com/abstract/journals/084/9/2/article-p261.xml){rel=""nofollow""}. ## What is translation? **Translation is the conversion of written text.** A translator works on a document that already exists — a contract, a website, a subtitle file, a medical record — and produces a new written document in the target language. The work happens after the fact, with time to research terminology, consult reference material, use translation-memory tools, and revise. The output is durable: a text that can be reviewed, certified, and reused. That review step is the practical dividing line. A translated contract can be checked word by word before anyone signs it. An interpreted negotiation cannot — which is why the two professions certify differently and price differently, and why "can I get it in writing?" is a translation question even when the meeting itself was interpreted. ## Interpretation vs translation: the difference at a glance | | Interpretation | Translation | | ------------------- | ----------------------------------------------- | ----------------------------------------- | | **Medium** | Spoken or signed language | Written text | | **When it happens** | Live, in real time | After the source text exists | | **Direction** | Often both ways in one session | Usually one way per assignment | | **Deliverable** | Ephemeral — heard, then gone | Durable — a document you can review | | **Time to work** | Seconds (simultaneous) to a pause (consecutive) | Hours to weeks, with revision | | **Tools** | Booth/console, receivers, or a platform | CAT tools, translation memory, glossaries | | **Fidelity target** | Meaning and intent, at speaking pace | Full accuracy, reviewable word by word | | **Typical pricing** | Per interpreter, per day or hour | Per word or per page | Two consequences follow directly from the table. First, the skills barely overlap: an outstanding contract translator may be unable to interpret a meeting, and vice versa — hiring one to do the other's job is a category error, not a compromise. Second, machine assistance entered the two fields at different speeds: machine translation of text has decades of history, while machine *interpretation* — live speech, both directions, at conversation pace — only became practical recently. We cover that shift [below](https://intermind.com/#ai-interpretation). ## Types of interpretation "Types of interpretation" mixes two independent axes that are worth separating: the **mode** (how the interpreting itself is timed and delivered to the listener) and the **delivery format** (where the interpreter is and how the audio reaches you). ### The modes | Mode | How it works | Typical setting | | -------------------------- | --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | | **Simultaneous** | Interpreter renders the speech while the speaker is still talking; listeners hear it seconds behind | Conferences, multilingual meetings, broadcasts, the UN | | **Consecutive** | Speaker pauses every few sentences; interpreter renders the segment | Medical consultations, depositions, site visits, interviews | | **Whispered (chuchotage)** | Simultaneous interpretation whispered directly to one or two listeners, no equipment | A delegate or executive in a room running in another language | | **Liaison (bilateral)** | Interpreter alternates between both languages in a back-and-forth exchange | Negotiations, escorted visits, small two-party meetings | | **Sight translation** | A written document read aloud in another language on the spot | Forms and letters inside a medical or legal appointment | **Simultaneous interpretation** is the mode that scales: the meeting runs at full speed, any number of languages can run in parallel (one channel per language), and listeners choose their channel. Its historical cost — soundproof booths, two interpreters per language, day rates — is what made it an event-day service rather than an everyday one. We wrote a full plain-language guide to it: [Simultaneous interpretation: what it is, how it works, and when AI can do it](https://intermind.com/blog/simultaneous-interpretation-guide). **Consecutive interpretation** trades time for simplicity: no equipment, but the meeting takes roughly twice as long, because everything is said twice. It shines where precision per sentence matters more than pace — a diagnosis, a sworn statement, a contract clause read aloud. **Whispered and liaison** interpretation are small-scale variants of the two modes above — worth knowing by name because agencies quote them separately, but nearly every real decision reduces to *simultaneous vs consecutive*. ### The delivery formats | Format | What it is | What it changes | | ----------------------------- | ----------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | **On-site** | Interpreter physically present (booth or beside the listener) | The traditional baseline: highest setup cost, hardware in the room | | **OPI — over-the-phone** | Interpreter joins by phone, on demand | Minutes-based pricing, audio only, no visual context | | **VRI — video remote** | Interpreter joins by video call | Adds visual context (critical for sign language); common in hospitals | | **RSI — remote simultaneous** | Interpreters work from a hub or home via a platform; audio streams to each listener's device | Removes booths, receivers, and travel from simultaneous interpreting | | **AI interpretation** | A speech pipeline (recognition → translation → synthesis) interprets continuously, per listener | Removes the booking and the per-day rate; interpretation becomes a software feature | The formats are a cost story. On-site simultaneous is the most expensive configuration in language services: per interpreter, per language, per day, plus equipment. OPI and VRI made *consecutive* interpretation available on demand — a hospital doesn't schedule a Tagalog interpreter for a possible emergency; it dials one. RSI did the same restructuring for *simultaneous*: the interpreters are still professionals at day rates, but the booth, the receivers, and the travel are gone. And AI interpretation removes the remaining constraint — the booking itself — which changes *which meetings get interpreted at all*. The weekly sync between Berlin, São Paulo, and Tokyo was never going to book two interpreters per language pair; with AI it doesn't have to. ## Which type do you need? Match the situation, not the terminology: - **A conference, summit, or multilingual event** → simultaneous. Human (on-site or RSI) for headline stakes and budget to match; AI for the sessions, breakouts, and events that a booth budget would exclude. What that looks like in practice: [events and conferences](https://intermind.com/use-case/events). - **A medical appointment or patient conversation** → consecutive, usually via VRI or OPI on demand; in several jurisdictions qualified interpreters in healthcare are a legal requirement, not a courtesy. For multilingual telehealth and staff meetings, AI interpretation covers the everyday layer: [healthcare](https://intermind.com/use-case/healthcare). - **A working meeting, webinar, or recurring cross-border call** → this is the space human interpretation never reached (nobody books a booth for a stand-up) and where [AI simultaneous interpretation](https://intermind.com/features/simultaneous-interpretation) is the native fit: every participant picks a language, the meeting runs at full speed. - **A court hearing, sworn deposition, or certified proceeding** → an accredited human interpreter, full stop. Where the interpretation *is* the legal record, machine output doesn't qualify — and vendors claiming otherwise should worry you. - **A document — contract, report, website, patent** → that's translation, not interpretation. Different professional, different certification, priced per word. (If the document comes up *inside* a live meeting, that's the one place the two jobs meet — see below.) []{#ai-interpretation} ## Where AI fits: machine interpretation AI entered translation and interpretation asymmetrically. Machine translation of text is decades old and everywhere. Machine **interpretation** — live speech in, live speech out, both directions, at conversation pace — became practical only when speech recognition, machine translation, and voice synthesis each got fast enough to chain with sub-second latency. We build one of these systems, so here is the honest shape of the category rather than a pitch: - **What AI interpretation does today:** continuous simultaneous interpretation of a meeting, per listener, with no booking and no per-day rate. In InterMIND's case that is 24 languages of live voice, with each speaker's **own voice** preserved via zero-shot synthesis rather than one synthetic narrator ([how that works](https://intermind.com/blog/own-voice-translation)) — plus the parts of a meeting a human interpreter was never asked to cover: chat messages, shared notes, and documents dropped into the call, translated inline. - **What to demand from any vendor, ours included:** published per-language-pair quality, not a language count. "Supports 60 languages" says nothing about *your* DE↔EN. Ours is measured monthly on FLORES-200 and published in full at [/benchmark](https://intermind.com/benchmark) — or [run the live demo](https://intermind.com/demo) and judge with your own ears. - **What AI interpretation does not do:** replace accredited interpreters in certified settings. Courts, treaty negotiations, and sworn proceedings belong to humans. AI's territory is the enormous space below that bar — the meetings that never got interpretation because booths and day rates priced them out. The major meeting platforms sit at intermediate points: Zoom and Teams offer [interpretation](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0064768){rel=""nofollow""} [channels](https://support.microsoft.com/en-us/office/use-language-interpretation-in-microsoft-teams-meetings-b9fdde0f-1896-48ba-8540-efc99f5f4b2e){rel=""nofollow""} for human interpreters you source and pay yourself, all three offer translated captions on qualifying plans, and Google Meet ships [Gemini speech translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""} on qualifying plans. If you're comparing tools rather than learning the category, the buyer's guide does that comparison with sources: [best AI translation tools for conferences and meetings](https://intermind.com/blog/best-ai-conference-translation-tools). --- ## FAQ **What is the difference between interpretation and translation?** Interpretation converts spoken (or signed) language live, while the communication happens; translation converts written text, after the text exists. An interpreter works in real time with no revision pass; a translator works on a document with time to research and revise. They are separate professions with separate certifications. **What are the main types of interpretation?** By mode: simultaneous (rendered while the speaker talks), consecutive (rendered in the speaker's pauses), whispered/chuchotage (simultaneous, whispered to one or two listeners), and liaison (back-and-forth in small exchanges). By delivery: on-site, over-the-phone (OPI), video remote (VRI), remote simultaneous (RSI), and AI interpretation. **What is simultaneous interpretation?** Interpretation delivered in parallel with the speaker — listeners hear their language a few seconds behind, and the meeting never pauses for translation. It's how the UN runs sessions in six languages and how multilingual conferences keep one agenda. Full guide: [simultaneous interpretation explained](https://intermind.com/blog/simultaneous-interpretation-guide). **Is an interpreter the same as a translator?** No. An interpreter works with live speech; a translator works with written text. The skills barely overlap — one profession trains working memory and real-time delivery, the other trains research, writing, and revision. Someone who does both exists, but each role is hired, certified, and priced separately. **Which is harder, interpretation or translation?** They are hard differently. Interpretation compresses everything into real time — no dictionary, no second draft, meaning rendered at the speaker's pace, which is why simultaneous interpreters work in pairs and rotate every \~30 minutes. Translation demands a different rigor: full accuracy on a text that will be reviewed word by word, sometimes certified and legally binding. **Can AI do interpretation?** For everyday meetings, webinars, and events — yes: AI systems now interpret live speech continuously, per listener, in both directions. Quality varies by vendor and language pair, so verify rather than assume: check published per-pair benchmarks ([ours is here](https://intermind.com/benchmark)) or [test it live on your own voice](https://intermind.com/demo). For certified settings — courts, sworn proceedings — accredited human interpreters remain the only qualifying option. **Do Zoom, Teams, or Google Meet include interpretation?** Partially. Zoom and Teams provide interpretation channels for human interpreters you hire yourself, and all three platforms offer translated captions on qualifying plans; Meet additionally ships Gemini speech translation on qualifying plans. None of it is turnkey full-meeting interpretation — the per-platform breakdowns are here: [Zoom](https://intermind.com/blog/zoom-live-translation), [Teams](https://intermind.com/blog/teams-live-translation), [Meet](https://intermind.com/blog/google-meet-live-translation). --- ## See the difference, live Definitions explain the category; hearing it settles it. - **[Run the live demo](https://intermind.com/demo)** — speak, and hear yourself interpreted into another language, in your own voice, on the production pipeline. - **[Read the benchmark](https://intermind.com/benchmark)** — per-language-pair quality, measured monthly, published in full. - **[AI simultaneous interpretation, in detail](https://intermind.com/features/simultaneous-interpretation)** — how live interpretation works inside an InterMIND meeting. — The Mind.com Team --- *Sources: [ATA — Translator vs. Interpreter](https://www.atanet.org/client-assistance/translator-vs-interpreter/){rel=""nofollow""}, [United Nations DGACM — Interpretation](https://www.un.org/dgacm/en/content/interpretation){rel=""nofollow""}, [Chmiel — Boothmates forever? On teamwork in a simultaneous interpreting booth](https://akjournals.com/abstract/journals/084/9/2/article-p261.xml){rel=""nofollow""} (two interpreters per booth, \~30-minute turns), [Zoom — Using language interpretation in meetings and webinars](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0064768){rel=""nofollow""}, [Microsoft Teams — Use language interpretation in meetings](https://support.microsoft.com/en-us/office/use-language-interpretation-in-microsoft-teams-meetings-b9fdde0f-1896-48ba-8540-efc99f5f4b2e){rel=""nofollow""}, [Google Meet — Speech Translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""}, checked August 2026. Vendors change plans and language lists over time — check their pages for the current state.* # Is your meeting confidential? People say things in meetings they would never put in an email. Salary bands. A candid read on a partner. The negotiation floor. Legal strategy. That's what meetings are *for* — the medium where you can still think out loud. Which makes it odd how rarely anyone asks where all that thinking-out-loud actually goes. The words leave your mouth, become audio, hit a server, possibly feed a model, possibly become a transcript, possibly get retained under a policy you agreed to by clicking "join." "Confidential" on a vendor homepage tells you nothing about any of that. We put the word on our homepage too. So here's what we think it should be required to mean — three questions concrete enough to check, and our answer to each. Not because you should trust the answers, but because every one of them is verifiable, which is the only kind of confidentiality claim worth making. ## Question 1: Who processes your words, besides the people in the room? A translated meeting is, unavoidably, a processed meeting. Speech becomes text, text becomes another language, an AI writes the recap. The confidentiality question isn't *whether* machines touch your words — it's **which machines, run by whom, under what terms**. Our answer: voice and chat translation run on **our own engine, on dedicated servers in France** — not routed through generic third-party translation APIs, so your negotiation doesn't transit a consumer service. The AI features that do use an external model (the recap) are pinned to an **EU provider under zero-data-retention terms** — your meeting content is processed and immediately discarded, not stored, not used for training. No meeting content reaches a US-domiciled model. Document translation goes through exactly one named sub-processor (DeepL's Document API, for file-format support), disclosed on our sub-processor list rather than "available on request." The checkable part: every vendor that touches meeting data is [named on the page, with what it does and where it's domiciled](https://intermind.com/features/security). ## Question 2: What remains afterwards — and who decides? A meeting leaves traces: chat history, files, recaps, sometimes a recording. Confidentiality isn't the absence of a record (a record in your language is half the product's value — [we've argued it's the half the industry forgets](https://intermind.com/blog/the-meeting-a-week-later)). It's **the record being yours**: - **Recording is a host decision.** It's available when the host turns it on — not a default, not a vendor's background process. - **Retention is yours to end.** Data is kept until you or your team owner delete it — the criterion is written in the privacy policy, not implied. - **Deletion actually deletes.** Erasing an account cascades across the database and sweeps every storage blob, recordings included. We audited our own codebase to make sure the cascade runs — and [published what we found and fixed](https://intermind.com/blog/gdpr-audit-what-we-closed). ## Question 3: Whose laws apply to the room? Where data lives decides who can compel it and which rules protect it. For a bakery's stand-up, academic. For a bank, a clinic, a bidder in a public tender — the first question procurement asks. Our answer: **EU at every runtime hop.** The translation engine runs in France, meeting orchestration in Paris, the database in Frankfurt; storage and analytics are EU-resident too. That's a jurisdiction you can name in a compliance review, not a region-of-convenience that shifts with the vendor's load balancer. ## The quieter half: data never collected One more property, easy to miss: **your guests join by link, with no account.** No sign-up means no profile, no address book upload, no identity graph of everyone who ever attended your meetings. The least exposed data is the data that was never collected — and it's also just [how hospitality should work](https://intermind.com/blog/give-every-guest-their-language). ## Make any vendor answer the same three questions Take this list into your next procurement call — it works on us too: 1. Which companies process meeting content, and is the list published? 2. Does any AI feature retain meeting content after processing? Under what terms — marketing copy, or contract? 3. In which jurisdictions does meeting data live, hop by hop? 4. Who can start a recording, and can participants tell? 5. When a customer deletes data, what proves the deletion ran — policy text, or an audited cascade? If the answers arrive quickly and specifically, you're talking to a vendor who expected the questions. Our full set — DPA with a 72-hour breach window, sub-processor list, uptime terms, the GDPR audit — lives at [Privacy & Security](https://intermind.com/features/security), one click deep, where procurement can get properly bored. ## Check it yourself - **[Read the security page](https://intermind.com/features/security)** — infrastructure, encryption, sub-processors, DPA, the audit. - **[Try the live demo](https://intermind.com/demo)** — no signup required, which is itself a data-collection answer. - **[How a meeting actually runs](https://intermind.com/blog/where-one-intermind-meeting-actually-runs)** — the infrastructure map, hop by hop, for the technically curious. --- ## FAQ **Is a video meeting confidential by default?** No platform can promise that as a default — it depends on who processes the audio, what's retained afterwards, and where the data lives. Those are answerable questions: ask for the sub-processor list, the AI retention terms, and the jurisdiction of each hop. A vendor with real answers publishes them. **Does live translation mean an AI is listening to my meeting?** Yes — translation is processing, there's no way around that. What differs between vendors is which machines process it and what they keep. In InterMIND, voice and chat translation run on our own engine on dedicated EU servers, the recap AI is bound to zero-data-retention terms, and no meeting content reaches a US-domiciled model. **Who can record an InterMIND meeting?** Recording is controlled by the host — it's available when the host enables it, never a background default. Access controls cover who can join, present and record. **What happens to meeting data when I delete my account?** Deletion cascades across the database and sweeps every storage blob, recordings included. Retention until that moment is under your control — data is kept until you or your team owner delete it, per the published policy. **Where is InterMIND meeting data processed?** In the EU at every runtime hop: translation engine on our own infrastructure in France, meeting orchestration in Paris, database in Frankfurt, EU-resident storage and analytics. Document translation uses one named EU-disclosed sub-processor (DeepL), listed publicly. --- *Sources: [DeepL — data security](https://www.deepl.com/en/pro-data-security){rel=""nofollow""}, checked August 2026. InterMIND claims in this post are documented on [Privacy & Security](https://intermind.com/features/security) — infrastructure, sub-processor list, DPA and the GDPR audit report — and in [the audit write-up](https://intermind.com/blog/gdpr-audit-what-we-closed).* # Linguee English–Portuguese: what it's actually for (and when you need a different tool) People searching for the Linguee English–Portuguese translator usually want one of two things — and Linguee only does one of them well. If the question is *"how do I say **this word** in this context?"*, that is exactly the job Linguee was built for. If the task is *"translate this paragraph / this document / this conversation"*, it was never built for that — and insisting costs you time. This guide separates the two uses honestly: what Linguee does better than its competitors, where it stops, and which tool takes over from there. --- ## What Linguee is — and why it's good at it Linguee is a **corpus dictionary**: besides the word's translation, it shows dozens of real sentences — contracts, websites, published documents — where that word was actually translated by human translators, side by side with the original. It's from the same company that built DeepL; Linguee's corpus is part of what trained its sibling translator. That's what it does that a plain machine translator doesn't: **nuance**. Is "enforcement" *aplicação*, *execução* or *fiscalização* in Portuguese? Depends on context — and Linguee shows you the context, ten real examples of each option. A machine translator picks for you and hides the alternatives. Use Linguee when: - you're writing in the other language and want the *right* word, not just a translation; - you're learning and want to see a word alive in real sentences; - you need to check how a technical or legal term is actually translated in practice. --- ## Where Linguee stops Three limits, all by design — Linguee is a dictionary, not a translation engine: 1. **It doesn't translate running text.** Paste a paragraph and you don't get the paragraph back translated — you get dictionary entries. For text, its own maker points you at DeepL; Google Translate and others handle it too. 2. **It doesn't translate documents.** PDF, DOCX, slide deck — wrong category of tool. 3. **It has no voice.** No speaking and hearing a translation — not one sentence, let alone a conversation. None of these limits is a flaw. It becomes a problem only when we use the tool from the wrong step: Linguee answers *"which word?"*, not *"what did they say?"*. --- ## The right tool for each job | Job | Right tool | | ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | A word or phrase, with nuance | **Linguee** (or Reverso Context, same model) | | A paragraph, an email, running text | DeepL, Google Translate | | A whole document (PDF, DOCX) | A document translator — InterMIND, for one, [translates files in the meeting across 30 languages](https://intermind.com/features/realtime-translation) | | A spoken phrase, face to face | A phone voice app ([the three kinds of voice translator](https://intermind.com/blog/tradutor-de-voz)) | | **A voice conversation — a call, a meeting, several people** | **Live meeting translation** — everyone speaks and hears their own language | The first row and the last are the two that get confused in practice. The ladder is always the same: word → text → voice → *conversation*. Each rung has its own tool, and the last one is the least known — which is why it gets its own section. --- ## The rung almost nobody knows: the conversation The typical Linguee English–Portuguese user is *working* between the two languages — a foreign client, a supplier, a distributed team, a job interview. Then the voice meeting arrives, and the dictionary can't join the call with you. That step has its own category — the [voice translator](https://intermind.com/blog/tradutor-de-voz) — and the version built for real conversations is live meeting translation. In InterMIND, the call itself is multilingual: you speak Portuguese, the other side hears English; they answer in English, you hear Portuguese — **at the same time, under a second behind, in any combination of 24 languages**, with no English anchoring. Per-language-pair quality is public, measured on real traffic, at [`/benchmark`](https://intermind.com/benchmark) — including the English ↔ Portuguese pair. And if what you like about Linguee is precisely that it *shows instead of promising*: [`/demo`](https://intermind.com/demo) does the same with voice — play some audio or speak, and hear the live translation. --- ## The verdict Linguee answers the question it chose to answer: *which word, in this context*. Keep it open in a tab — we do. Just don't ask it for what it never promised: running text is DeepL's or Google Translate's job; documents are a file translator's job; and a **conversation** — the meeting with the client who doesn't speak your language — is the job of [live meeting translation](https://intermind.com/blog/tradutor-de-voz). Word, text, voice, conversation. Four jobs, four tools — and now you know which is which. — The Mind.com Team --- *Sources: [Linguee — English–Portuguese dictionary](https://www.linguee.com/english-portuguese){rel=""nofollow""}, [DeepL Translator](https://www.deepl.com/en/translator){rel=""nofollow""} — Linguee and DeepL are built by the same company, and Linguee's own pages point to DeepL for running text; checked August 2026.* # Your SOPs Are Translated. Your Meetings About Them Aren't. A quality director quoted in a recent [MasterControl analysis](https://www.mastercontrol.com/uk/gxp-lifeline/electronic-document-management-for-multilingual-compliance/){rel=""nofollow""} said 30% of her documentation-management time goes to keeping translated SOPs (Standard Operating Procedures) in sync across sites. The pains MasterControl describes are real and well-known to anyone running multilingual GxP (Good Practice) operations: - Subtle procedural differences emerging between language versions of a controlled document - Asynchronous updates leaving sites operating from outdated guidance - Training competency hard to verify when materials live in different languages - Approval cycles inflating with every additional language reviewer A controlled document management system — MasterControl, Veeva, IQVIA SmartSolve — is built to solve the **document layer**. Version control, approval routing, controlled distribution, audit trail. That part is well-served. But documents do not sit on shelves. They are **discussed.** And the discussion layer carries a second multilingual problem that no DMS (Document Management System) is designed to fix. --- ## Where Multilingual Breaks Down in Regulated Discussions A SOP gets approved in the DMS. Now what happens around it? **Cross-site CAPA (Corrective and Preventive Action) review.** Eight quality leads from five plants across four countries dial in to walk through a deviation. Three of them are operating in their second or third language. The conversation produces the corrective-action wording. **The wording goes into the audit record. In whose language?** **Inspection readiness.** EU sites prep for an FDA visit. The talking-points walkthrough happens in English over Zoom because that is the lingua franca. The operators who will face the inspectors think and respond in Polish, Spanish, Mandarin. Practice in one language, perform in another — a known driver of inconsistency findings. **Vendor and CMO Q\&A.** A SOP change triggers questions from a contract manufacturer. Twenty back-and-forth clarifications happen on a call. Each one modifies how the procedure is read. **The audit trail of those clarifications exists only in someone's notes — in one language.** **SOP rollout training.** A new procedure goes live across plants. The walkthrough happens once, in English, recorded once. Operators in Brazil, Italy, and Vietnam re-watch with auto-subtitles, miss nuance, and ask the same question three times in three local meetings. In every one of these cases, the document is fine. The DMS is doing its job. **The conversation around the document is where the language tax compounds.** --- ## What InterMIND Does in Those Rooms InterMIND is a video meeting platform with translation built into the core, not bolted on. For regulated discussions, four capabilities matter. ### 1. Live Voice Translation in 24 Languages Each participant picks their own translation language at the start of the call. Speech is translated in under a second, with tone and intent preserved. Live subtitles run alongside the audio in every language in the room. No interpreters, no separate tools. (On-demand file translation in chat covers a wider 30 — see [the per-surface breakdown](https://intermind.com/blog/how-many-languages-do-you-support).) ### 2. Live Shared Protocol — Per-Viewer Language with Diffs This is the capability that maps directly onto the asynchronous-update pain MasterControl describes — and inverts it. The host shares a meeting protocol typed live during the call. **Every participant sees it in their own language, in real time.** When the host edits a line, the change is translated and pushed to every viewer with a **diff view**: previous wording, new wording, difference highlighted. No participant is reading yesterday's version of the procedure. Language versions of the protocol stay synchronized to the second, not to the next translation cycle. ### 3. Real-Time Translation of Chat and Attachments Every chat message in the meeting is translated for every viewer. Files dropped into the chat — RFP annexes, deviation forms, supplier certificates, PDFs of draft procedures — are translated on demand into the requesting participant's language. Translations are cached per language and re-translated only on the changed paragraphs when the source updates. ### 4. Compliance-Grade Transcript Bundle After the meeting you export one bundle: full transcript with timestamps, speaker tags, source language and target language for every utterance, the final protocol, and all language versions of it. Suitable for the audit record. For the data-protection layer underneath all of this — where the meeting runs, how erasure and retention work, what the sub-processor list says — we ran the codebase through a [full GDPR audit and closed each item against the code](https://intermind.com/blog/gdpr-audit-what-we-closed). --- ## What InterMIND Does Not Do InterMIND is not a document management system. It does not store master SOPs, route controlled-document approvals, manage versioning of controlled documents, or replace your e-signature workflow. DocuSign, Adobe Sign, and PandaDoc own legally binding signatures and signature audit trails. **We do not sign documents.** MasterControl, Veeva, and IQVIA SmartSolve own the controlled-document lifecycle. **We do not manage SOPs.** InterMIND lives in the meeting room next to those systems. The DMS holds the source of truth. We make sure the discussion of that source happens without language loss, and that the discussion itself is captured on the audit record in every language that was in the room. --- ## Walkthrough: A Cross-Site CAPA Review Eight participants — Frankfurt (German), Madrid (Spanish), Wrocław (Polish), Mumbai (English), São Paulo (Portuguese), Shanghai (Mandarin), and two Boston (English) leads. The deviation is on a tablet-coating step at the Wrocław plant. | Step | What Happens | | ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Pre-meeting | Each participant logs in and picks their translation language. UI language and translation language can differ — a Spanish speaker can keep an English UI and still receive translation in Spanish. | | Discussion | Each person speaks their native language. Audio is translated to every other language under one second. Live subtitles run in parallel. | | Live protocol | The host (Boston) types the proposed CAPA wording into the shared protocol pane. Every participant sees it in their own language as it is typed. The Wrocław lead suggests an edit; the host accepts; every participant sees the diff in their language. | | Attachments | The host drops the deviation report PDF into chat. Each participant requests the translation in their language; results are cached. | | Close | Transcript bundle exported: original-language audio, per-speaker transcript, translated transcripts, final protocol, and all language versions of the protocol. | The CAPA decision goes into the DMS. The discussion that produced it is on the audit record, in every language that was in the room. --- ## Where to Start InterMIND's Basic plan is free forever — no credit card required. The fastest way to see whether the fit is real for your team is to walk through your own scenario against ours. → [See use cases](https://intermind.com/use-case) — The Mind.com Team **Your DMS keeps the document straight. We keep the conversation straight.** --- *Sources: [MasterControl — Multi-language documentation challenges in European life sciences manufacturing](https://www.mastercontrol.com/uk/gxp-lifeline/electronic-document-management-for-multilingual-compliance/){rel=""nofollow""} — the 30% figure and the four documentation pitfalls cited above; checked August 2026.* # Arabic voice translator: one phrase, face-to-face, or a whole meeting? "Voice translator" — **مترجم صوتي** — is one name for three different products. A phone app that listens and reads the translation back out loud. Earbuds that whisper a translation of the person in front of you. And live meeting translation, where the whole conversation runs in several languages at once and each participant hears it in their own. All three translate speech — but only the last keeps an **instant, multi-person conversation** flowing without turning every exchange into a wait. This guide separates the three, shows how each works, and asks the question that decides your pick: **a phrase, a face-to-face exchange, or a full conversation?** --- ## The three kinds ### Phone-app voice translators Open the app, tap the mic, speak — it shows or speaks the translation. Google Translate, Apple Translate, Microsoft Translator, dozens of travel apps. Genuinely good for one sentence at a time: a menu, a taxi, a front desk. The mechanic is a relay: you speak → it transcribes → it translates → it speaks, then the other person answers into the same device and it relays the other way. It works because the exchanges are short and you pass the phone back and forth. Stretch it to a real discussion and the relay becomes the bottleneck — you take turns operating a translator instead of talking. ### Earbud / device translators AirPods with Live Translation, Pixel Buds, dedicated translator earbuds. You wear them and hear the other person translated in your ear. Nicer than looking at a screen — but **one-to-one, and usually needing matching gear on both sides.** Built for a traveler and a host, not for a room. ### Live meeting translation A different architecture, not a bigger app. The *meeting itself* is multilingual: every participant speaks their own language and hears everyone else in **theirs**, instantly, for the whole meeting. No phone changing hands, no matching earbuds. It's the only kind that survives an instant conversation among several people. --- ## Which one do you need? - **One phrase** — a menu, a direction, a short exchange with a stranger: a phone app is perfect. Don't over-buy. - **A face-to-face exchange with one person, both equipped** — earbuds are the nicest experience. - **A conversation — a call, a meeting, several people, instant back-and-forth** — live meeting translation, because the others turn every exchange into a relay and every extra person into a broken assumption. The usual trap is using a phrase tool for a full conversation: it technically "works," and it makes the meeting twice as slow, because everyone waits on the relay instead of talking. --- ## What "instant" and "live" really require Three things have to be true at once — and this is where most tools quietly fail: - **Per-listener, not per-device.** Every participant hears the room in the language they picked, simultaneously — five people, five languages, one call. - **Sub-second, and continuous.** If translation only arrives after the speaker pauses, that's consecutive interpretation with a synthetic voice — you feel every gap. Truly instant translation keeps pace with the talking. - **No English anchor, no regional gate, no five-language beta.** Arabic ↔ English, Arabic ↔ French, Arabic ↔ Turkish — any mix, translated between participants directly, nothing forcibly routed through English. That's not a bigger phone app. It's the architecture behind [real-time meeting translation](https://intermind.com/blog/real-time-meeting-translation) — the foundational guide to the category. --- ## Where InterMIND fits InterMIND is the third kind: a voice translator for real, instant conversations, built as live meeting translation. - **24 languages live on voice**, chat and shared notes — Arabic among them, in any mix: Arabic ↔ English with a partner in London, Arabic ↔ French with a client in Paris, Arabic ↔ Turkish with a supplier in Istanbul. ([The full end-to-end translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Per-listener audio, sub-second.** Each participant hears the meeting in their own picked language at the same time. (Under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Browser or app** — web, desktop, iOS and Android; guests join from one link, no signup. - **Webinars and conferences included** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — drop a PDF or DOCX into the meeting and each viewer sees it in their language; 30 languages on files. - **Quality you can audit yourself** — per-language-pair scores on real traffic, published at [`/benchmark`](https://intermind.com/benchmark), methodology written down. We don't hide the weaker pairs behind one marketing number — check how your pair stands before you rely on us. A phone app is right for a phrase. Earbuds are right for one person in front of you. When the job is an **instant conversation among several people** — that's what InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and *hear* per-listener translation instead of reading promises about it. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, methodology included. - Start with the category guide: [*Real-time meeting translation*](https://intermind.com/blog/real-time-meeting-translation). - Comparing specific language pairs? See the voice-pair guides for [English ↔ Italian](https://intermind.com/blog/traduttore-inglese-italiano-vocale), [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale), [Romanian ↔ Italian](https://intermind.com/blog/traduttore-rumeno-italiano-vocale) and [Hindi ↔ English](https://intermind.com/blog/hindi-to-english-voice-translator) — or the full guide to [what actually translates a live conversation](https://intermind.com/blog/traduttore-vocale). --- ## FAQ **Can a phone app translate a whole Arabic meeting?** It technically "works" as a relay — speak, wait, pass the phone — but every exchange becomes a wait, and every extra participant breaks the one-device assumption. Phone apps are built for a phrase at a time, not for an instant multi-person conversation. **Do both people need special earbuds for live translation?** Usually yes — earbud translation (AirPods Live Translation, Pixel Buds) is one-to-one and works best when both sides are equipped or share a screen for the other half of the exchange. It is built for a traveler and a host, not for a room. **Which languages can pair with Arabic in a live meeting?** In InterMIND, any of the 24 live languages, in any mix — Arabic ↔ English, Arabic ↔ French, Arabic ↔ Turkish — with each participant hearing the meeting in their own language. Per-pair quality is published at [`/benchmark`](https://intermind.com/benchmark). — The Mind.com Team --- *Sources: [Apple — Live Translation with AirPods](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google — Translate with Pixel Buds](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. Vendors change features and language lists over time — check their pages for the current state.* # Otter.ai alternatives: the 9 best in 2026 (and when to switch) Otter.ai made "AI meeting notes" a category, and for English-first teams it's still a fine default. But the two complaints that send people searching for an Otter alternative in 2026 are consistent: **the minute caps** (300/month free, 1,200 on Pro — heavy meeting loads burn through both) and **the language list** — Otter transcribes in six languages (English, Spanish, French, German, Japanese, Chinese) in a market where competitors claim 30, 58, or 100+. Below are the 9 alternatives worth shortlisting, with pricing and limits verified in July 2026. Full disclosure: **we build one of them** ([InterMIND](https://intermind.com){rel=""nofollow""}), and it's the right pick only for one specific case — meetings that happen in more than one language, live. For everything else we'll point you to the competitor that fits. > Deciding between the two biggest names? We wrote the head-to-head: [Otter vs Fireflies, honestly compared](https://intermind.com/blog/otter-vs-fireflies). And the detailed 1-on-1 against us: [InterMIND vs Otter](https://intermind.com/compare/otter). --- ## The verdict table | Tool | Best for | Free plan | Paid from\* | Languages | | ---------------------------------------------------- | --------------------------------------- | --------------------------- | ---------------------- | ------------------ | | [InterMIND](https://intermind.com/#tool-1) | Live-translated multilingual meetings | Yes (Basic) | per-user, 14-day trial | 24 voice / 30 docs | | [Fireflies.ai](https://intermind.com/#tool-2) | Unlimited transcription, 100+ languages | Unlimited (400 min storage) | $10/seat/mo | 100+ | | [Fathom](https://intermind.com/#tool-3) | Free unlimited recording | Unlimited | \~$16/mo | 38 | | [tl;dv](https://intermind.com/#tool-4) | Sales clips and coaching | Unlimited recording | \~$18/seat/mo | 30+ | | [Notta](https://intermind.com/#tool-5) | Non-English transcription + translation | 120 min/mo | \~$9/mo | 58 | | [MeetGeek](https://intermind.com/#tool-6) | Whole-team auto-capture | 3 h/mo | $9.99/user/mo | 100+ | | [Read AI](https://intermind.com/#tool-7) | Engagement analytics | 5 meetings/mo | $15/user/mo | 16+ | | [Krisp](https://intermind.com/#tool-8) | Noisy calls, no bots | Trial | \~$8/user/mo | 16+ | | [Built-in notetakers](https://intermind.com/#tool-9) | One-platform teams | Included | with suite plan | varies | \*Lowest advertised per-user price on annual billing, July 2026; month-to-month typically runs 1.5–2× higher. --- ## How to choose in 30 seconds - **You hit Otter's minute caps** → [Fireflies](https://intermind.com/#tool-2) (unlimited transcription) or [Fathom](https://intermind.com/#tool-3) (unlimited and free). - **Your recordings aren't in English** → [Notta](https://intermind.com/#tool-5) or [MeetGeek](https://intermind.com/#tool-6) — both transcribe far beyond Otter's six languages. - **Your *meetings* aren't in one language** — the room can't follow each other live → no notetaker fixes that, including Otter. You need [live translation](https://intermind.com/#tool-1). - **Otter is fine but pricey at team scale** → check [your platform's built-in notetaker](https://intermind.com/#tool-9) before paying anyone. --- ## The 9 best Otter.ai alternatives []{#tool-1} ### 1. InterMIND — when your meetings aren't in one language Otter and every tool below share one assumption: the meeting happens in a language everyone understands, and the AI writes it down. InterMIND is for when that assumption breaks. Each participant picks their language on joining and *hears* every speaker in it — in the speaker's own voice, sub-second — while the chat, shared notes, and dropped-in documents translate per viewer too. The post-meeting summary is a by-product, not the product. - **Price:** free Basic plan; Pro and Business per-user with a 14-day trial — [pricing](https://intermind.com/pricing). - **Languages:** 24 live (voice, chat, notes), 30 for documents. [Why we count it that way.](https://intermind.com/blog/how-many-languages-do-you-support) - **Limits:** voice minutes and chat words metered on a rolling 30-day window, shared across the team. - **Best for:** international teams, cross-border sales, any recurring call where someone is limping along in a second language. - **Honest caveat:** single-language meetings don't need it — pick a notetaker below instead. Quality is the thing to verify, not trust: per-language-pair scores are [published monthly](https://intermind.com/benchmark), and [the demo](https://intermind.com/demo) runs the production pipeline on your own voice. []{#tool-2} ### 2. Fireflies.ai — the obvious first candidate The default Otter rival, and on the two pain points that push people off Otter it wins outright: transcription is **unlimited on every plan** (storage is what's capped) and it covers **100+ languages** with auto-detection. - **Price:** free (unlimited transcription, 400 min storage/team); Pro $10/seat/mo annual ($18 monthly), 8,000 min storage; Business $19/$29 unlimited; Enterprise $39 annual. - **Languages:** 100+, with multi-language detection on Business+. - **Best for:** teams that transcribe a lot, in many languages, and want CRM/automation hooks. - **Vs Otter:** no minute anxiety and vastly more languages; the live-transcript editing workspace is weaker. Full head-to-head: [Otter vs Fireflies](https://intermind.com/blog/otter-vs-fireflies). []{#tool-3} ### 3. Fathom — the generous free plan Unlimited free recordings, transcription, and instant summaries. If Otter's caps are the whole problem and your meetings are mostly English, the move takes an afternoon. - **Price:** free unlimited; Premium $16/mo annual; Team $15, Business $25/user/mo annual (2-seat minimum). - **Languages:** 38. - **Best for:** individuals and small teams that want free and low-friction. - **Vs Otter:** more free value; fewer team knowledge-base features. []{#tool-4} ### 4. tl;dv — clips and coaching on a budget Free unlimited recording and transcription, with the paid tiers aimed at sales teams: clips, multi-meeting AI reports, CRM sync, coaching scorecards. - **Price:** free (limited AI summaries); Pro \~$18/seat/mo annual; Business \~$59/seat/mo annual. - **Languages:** 30+. - **Best for:** sales teams that want Gong-lite workflows without Gong pricing. - **Vs Otter:** better sales tooling and languages; the AI layer is what you pay for. []{#tool-5} ### 5. Notta — the language workhorse Notta's pitch is the transcript in another language: 58 transcription languages, translated transcripts on paid plans, bilingual and real-time translation views as add-ons. For *recorded* non-English audio, this is the specialist. - **Price:** free 120 min/mo; Pro from \~$9/mo annual, 1,800 min/mo; Business \~$16.67/seat/mo annual; real-time translation add-on from \~$6/mo. - **Languages:** 58. - **Best for:** individuals transcribing and translating across European and Asian languages. - **Vs Otter:** 58 languages vs 6 — but still after-the-fact; the live meeting stays untranslated. []{#tool-6} ### 6. MeetGeek — auto-capture for teams Joins everything on the calendar, transcribes in 100+ languages, and builds a searchable library with summaries and automations — capture-by-default for a whole team. - **Price:** free 3 h/mo; Pro $9.99/user/mo (20 h/mo); Business $17/user/mo (unlimited). - **Languages:** 100+, auto-detected. - **Best for:** managers who want every call captured without anyone pressing record. - **Vs Otter:** more languages, cheaper entry, set-and-forget; its transcript workspace is post-call focused rather than live. []{#tool-7} ### 7. Read AI — metrics, not just minutes Transcripts plus meeting *analytics*: talk time, engagement, sentiment, and reports that follow up in email and Slack. - **Price:** free 5 meetings/mo (1 h cap); Pro $15/user/mo annual; Enterprise $22.50; Enterprise+ $29.75/user/mo annual. - **Languages:** 16+. - **Best for:** leaders optimizing meeting culture, not just recall. - **Vs Otter:** adds analytics Otter lacks; higher per-seat price for full features. []{#tool-8} ### 8. Krisp — bot-free and noise-proof On-device audio processing: noise cancellation with an AI notetaker on top. No bot joins the call, and it works across any calling app — including phone bridges the bot-based tools can't touch. - **Price:** Core \~$8/user/mo annual (unlimited AI notes); Advanced \~$15/user/mo annual; 7-day trial. - **Languages:** 16+. - **Best for:** noisy environments, call centers, bot-averse counterparties. - **Vs Otter:** cleaner audio in, no bot; lighter search and collaboration features. []{#tool-9} ### 9. Zoom, Teams, and Meet's own notetakers Zoom AI Companion, Copilot in Teams, and Gemini in Meet all generate meeting summaries on paid tiers now — before paying a third party, check whether "good enough" is already in your suite. Where each stops (especially on translation) is documented in our platform guides: [Zoom](https://intermind.com/blog/zoom-live-translation), [Teams](https://intermind.com/blog/teams-live-translation), [Google Meet](https://intermind.com/blog/google-meet-live-translation). - **Price:** bundled with the suite. - **Best for:** teams that live in one platform. - **Vs Otter:** zero setup and zero extra cost; weaker cross-meeting search and no cross-platform story. --- ## The distinction the listicles blur Every Otter alternative above competes on the same axis: **recall** — how well the meeting gets written down. Price per seat, minutes per month, languages per transcript. There's a second axis the category quietly ignores: **comprehension** — whether everyone in the room could follow the meeting *as it happened*. Otter's six languages vs Fireflies' hundred doesn't matter to the participant who couldn't understand the discussion live; a transcript of a meeting you couldn't follow is [documentation of the problem, not a solution to it](https://intermind.com/blog/why-translation-quality-marketing-is-broken). If that's your actual situation, the fix isn't a better notetaker — it's [simultaneous interpretation built into the meeting](https://intermind.com/features/simultaneous-interpretation), which is the job InterMIND does and the reason it sits oddly at the top of this list. --- ## FAQ **What's the best free Otter alternative?** Fathom — unlimited free recordings, transcription, and summaries. tl;dv also records and transcribes free but limits AI summaries tightly. **What's the best Otter alternative for other languages?** For transcribing recordings: Notta (58 languages) or MeetGeek/Fireflies (100+ claimed). For *live* meetings in multiple languages: [InterMIND](https://intermind.com/compare/otter) — it translates the meeting while it happens instead of transcribing it afterward. **Is Fireflies better than Otter?** Different trade: Fireflies gives unlimited transcription and 100+ languages; Otter gives a live transcript you can highlight, comment on, and edit during the call, but meters your minutes and caps languages at six. The full comparison: [Otter vs Fireflies](https://intermind.com/blog/otter-vs-fireflies). **Does Otter.ai translate?** No. Otter transcribes in six languages, one language per conversation. It doesn't translate meetings live or translate transcripts between languages. **Can I just use Zoom's or Teams' built-in AI notes?** Often yes — on paid suite tiers they're decent and free-ish. You give up cross-meeting search, CRM sync, and multi-platform capture. --- ## Compare InterMIND head-to-head Prefer a direct one-on-one to the listicle above? These pages line InterMIND up feature-by-feature against the tools people cross-shop most: | Comparison | The trade-off | | ----------------------------------------------------------------------------- | --------------------------------------------------------- | | [InterMIND vs Otter](https://intermind.com/compare/otter) | Live translation vs six-language transcription | | [InterMIND vs Fireflies](https://intermind.com/compare/fireflies) | Live translation vs unlimited 100+-language transcription | | [InterMIND vs Zoom](https://intermind.com/compare/zoom) | Built-in interpretation vs AI Companion summaries | | [InterMIND vs Microsoft Teams](https://intermind.com/compare/microsoft-teams) | Built-in interpretation vs Copilot summaries | Or skip the shortlist and [start a multilingual meeting](https://intermind.com/meetings) to hear it on your own call. --- ## Test the axis that actually differentiates If languages are why you're leaving Otter, don't take any vendor's language count on faith — including ours: - **[Run the live demo](https://intermind.com/demo)** — InterMIND's production pipeline, your voice, any of 24 languages. - **[Read the benchmark](https://intermind.com/benchmark)** — per-pair quality published monthly, full distribution. - **[InterMIND vs Otter, feature by feature](https://intermind.com/compare/otter)** — the detailed 1-on-1. — The Mind.com Team --- *Sources: vendor pricing pages ([Otter.ai](https://otter.ai/pricing){rel=""nofollow""}, [Fireflies.ai](https://fireflies.ai/pricing){rel=""nofollow""}, [Fathom](https://fathom.ai/pricing){rel=""nofollow""}, [tl;dv](https://tldv.io/app/pricing/){rel=""nofollow""}, [Notta](https://www.notta.ai/en/pricing){rel=""nofollow""}, [MeetGeek](https://meetgeek.ai/pricing){rel=""nofollow""}, [Read AI](https://www.read.ai/pricing){rel=""nofollow""}, [Krisp](https://krisp.ai/pricing/){rel=""nofollow""}), checked July 2026. Prices are annual-billing rates unless noted and may change.* # Otter vs Fireflies (2026): which AI notetaker fits — and when neither does Otter.ai and Fireflies.ai are the two names everyone shortlists for AI meeting notes, and the comparison is genuinely close — they just meter you differently and speak a different number of languages. Here's the honest breakdown, from a vendor of neither: we build [InterMIND](https://intermind.com){rel=""nofollow""}, a live meeting *translator*, and we'll flag the one scenario where both notetakers lose. Everything else below is a straight comparison you can act on. **The short verdict:** Otter has the better *workspace* — live transcript you can highlight and edit, cleaner collaboration. Fireflies has the better *meter* — unlimited transcription on every plan and 100+ languages against Otter's six. Transcript-centric English teams lean Otter; volume and language breadth lean Fireflies. --- ## Side by side | | Otter.ai | Fireflies.ai | | ----------------------------------- | --------------------------------------- | ----------------------------------------------- | | **Free plan** | 300 min/mo, 30 min/conversation | Unlimited transcription, 400 min *storage*/team | | **Entry paid (annual)** | Pro $8.33/user/mo | Pro $10/seat/mo | | **Mid tier (annual)** | Business $19.99/user/mo | Business $19/seat/mo | | **What's metered** | Transcription minutes (1,200/mo on Pro) | Storage (8,000 min on Pro; unlimited Business+) | | **Meeting length cap** | 90 min (Pro), 4 h (Business) | — | | **Transcription languages** | 6 (EN, ES, FR, DE, JA, ZH) | 100+ with auto-detection | | **Multi-language meetings** | One language per conversation | Multi-language mode on Business+ | | **Live transcript workspace** | Highlight, comment, edit live | Post-call focused | | **Automations/CRM** | Business+ | CRM + automation integrations | | **Live translation of the meeting** | No | No | *Pricing verified on vendor pricing pages, July 2026 ([Otter](https://otter.ai/pricing){rel=""nofollow""}, [Fireflies](https://fireflies.ai/pricing){rel=""nofollow""}); annual-billing rates.* ## Choose Otter if… - Your team **works in the transcript** — highlights mid-call, shared notes, edits. Otter's live workspace is the best reason to pick it. - Meetings are **English-first** (or within its six languages) and predictable in volume, so the minute caps don't bite. - You want the lower entry price and don't need automation pipelines. ## Choose Fireflies if… - You transcribe **heavily or unpredictably** — unlimited transcription removes the meter anxiety entirely. - Your recordings span **languages Otter doesn't have** — 100+ with auto-detection is a different league than six. - You want notes flowing into **CRM and automations** without buying the top tier. ## The scenario where both lose Both tools assume the meeting itself worked — that everyone in the room understood the discussion, and the AI just needs to write it down. When your meetings happen across languages — a German engineer, a Brazilian PM, a Japanese client — that assumption is the problem. A transcript (even Fireflies' 100-language one) is delivered *after the call*, to people who couldn't fully follow it *during* the call. That's a different product category: [live simultaneous interpretation built into the meeting](https://intermind.com/features/simultaneous-interpretation). InterMIND translates the voice as it's spoken — each participant hears every speaker in their own language, in the speaker's own voice — and translates the chat, shared notes, and documents alongside, with [per-language-pair quality published monthly](https://intermind.com/benchmark). The meeting notes still get generated; they're just no longer the only thing standing between half your room and the discussion. The 1-on-1s, if you're weighing a switch: [InterMIND vs Otter](https://intermind.com/compare/otter) · [InterMIND vs Fireflies](https://intermind.com/compare/fireflies). The wider field: [the 9 best Fireflies alternatives](https://intermind.com/blog/fireflies-ai-alternatives) and [the 9 best Otter alternatives](https://intermind.com/blog/otter-ai-alternatives). --- ## FAQ **Is Otter or Fireflies cheaper?** Entry level: Otter Pro at $8.33/user/mo annual vs Fireflies Pro at $10. But Otter meters minutes (1,200/mo on Pro) while Fireflies transcribes without limits — at real volume, Fireflies is usually cheaper per transcribed hour. **Which supports more languages, Otter or Fireflies?** Fireflies by a wide margin: 100+ transcription languages with auto-detection vs Otter's six (English, Spanish, French, German, Japanese, Chinese). **Do Otter or Fireflies translate meetings?** Neither translates a live meeting. Fireflies transcribes many languages but doesn't translate between them in real time; Otter neither transcribes beyond its six nor translates. For live translation, that's [a different category](https://intermind.com/blog/best-ai-conference-translation-tools). **Can I use both free plans?** Yes, and plenty of people do: Fathom or Fireflies free for volume capture, Otter free for the occasional transcript-workspace session. Watch the two bots colliding in the same call, though — pick one recorder per meeting. --- ## If language is the real bottleneck - **[Run the live demo](https://intermind.com/demo)** — hear your own voice translated by the production pipeline, in any of 24 languages. - **[Read the benchmark](https://intermind.com/benchmark)** — per-pair quality, monthly, no averaging tricks. — The Mind.com Team --- *Sources: vendor pricing pages ([Otter.ai](https://otter.ai/pricing){rel=""nofollow""}, [Fireflies.ai](https://fireflies.ai/pricing){rel=""nofollow""}), [Otter — supported languages](https://help.otter.ai/hc/en-us/articles/360047247414-Supported-languages){rel=""nofollow""}, [Fireflies — supported languages](https://guide.fireflies.ai/articles/2973706448-learn-about-fireflies-supported-languages){rel=""nofollow""}, checked August 2026. Prices are annual-billing rates unless noted; vendors change plans and language lists over time — check their pages for the current state.* # Speak in your own voice — in a language you don't speak Here is the part of real-time translation that almost everyone gets wrong, and that almost no one talks about: **the voice you hear.** You can have excellent speech recognition and excellent translation, and still end up with a meeting that feels like a machine reading a list. Because the last step — turning the translated text back into sound — is where most tools quietly substitute *you* with a single generic synthetic narrator. Eight people in the room, one robot voice for all of them. You lose who is speaking, the emphasis, the personality. Intelligible, but not a conversation. InterMIND does the last step differently. When you speak, the other participants hear the translation in a voice that's **recognizably yours** — carrying your timbre and your way of speaking — now saying the words in their language. It isn't a flawless impression yet; the point is that it's *you* rather than a stock narrator, and it's getting better. This works for every participant, in both directions, at the same time. This post is the missing chapter of [*Inside the four translation pipelines that run InterMIND*](https://intermind.com/blog/inside-the-translation-pipelines): that piece explained how audio becomes translated audio. This one is about whose voice comes out the other end. --- ## The default everyone ships, and why it's flat If you've used live translation in any of the big meeting platforms, you know the sound. A neutral, evenly-paced voice reads the translation. It's the same voice whether the speaker is your CEO opening a town hall or a colleague cracking a joke. The technology underneath is text-to-speech with one fixed voice model, and the design assumption is that intelligibility is enough. In a real meeting it isn't. Half of what a meeting communicates is *who* is saying it and *how*. Strip the voice and you've turned a discussion into a transcript that happens to be spoken aloud. People stop reacting to each other and start waiting their turn. ## What InterMIND does instead The translation runs as a **cascaded pipeline** — three specialized stages in sequence rather than one model trying to do everything. The first two stages are covered in the [pipelines post](https://intermind.com/blog/inside-the-translation-pipelines); the voice step is the one this post is about: 1. **ASR — speech recognition.** Your words are transcribed in your own language, in your browser, as you speak. (Running it locally saves a round-trip and gives the lowest possible delay before translation can even start.) 2. **MT — translation.** The transcript is grouped into stable sentence fragments — *clauses* — so translation can begin before you've finished the sentence, and each fragment is translated progressively into the listener's language. 3. **Zero-shot TTS — voice synthesis.** Each translated fragment is spoken back out **using a sample of your own voice**, and streamed to the listener. It's that third stage — ASR → MT → **zero-shot TTS** — that produces the effect. "Zero-shot" means the system doesn't need a pre-recorded enrollment or a training session for your voice. It models your voice from the audio of the meeting you're already in. ## The warm-up: how it starts sounding like you so fast There's a chicken-and-egg problem hiding in "use a sample of your own voice." At the very start of a call, the system hasn't *heard* enough of you yet to model your voice well. InterMIND handles this with a progressive warm-up: - **For roughly the first 5–10 seconds**, while it's still gathering enough of your speech, each translated fragment is synthesized using the audio fragment that matches what you *just said* in your source language. The voicing is anchored to your real, immediate speech. - **Once there's a long enough sample** — that 5–10 second mark — the system locks onto it and uses it to voice everything afterward. In practice you don't hear a switch flip. The translation sounds more like you as the conversation gets going — not a perfect double of your voice, but clearly yours rather than a machine's, and improving as the model hears more. The combination of *progressive* translation (clause by clause, not sentence by sentence) and *progressive* voicing is what keeps the whole thing under the latency budget while still sounding human. ## The voice sample is never stored This is the part a security or legal team asks about immediately, so here it is plainly. The voice sample used for synthesis is **ephemeral**. It exists only for the live conference session, in service of voicing the translation, and it is **stored nowhere**. The Mind API and SDK that power the real-time session retain **no data** — everything temporary dies when the conference session ends. It's worth being precise about what this sample *is not*: it is not one of InterMIND's **recording** features. Recording a meeting's video and audio is a separate, deliberate action you take on purpose, with its own controls. The own-voice sample is not a recording — it's a transient input to the speech synthesizer that never outlives the call. This matters beyond privacy hygiene. "Speak in your own voice" is exactly the kind of feature that *sounds* like it should involve storing a voiceprint somewhere. It doesn't. The honest version is the better story: your voice is modeled in the moment and gone when you hang up. ## Why no one else ships this It's not that voice cloning is a secret. It's that doing it **live, per-participant, in both directions, under a one-second budget, across 24 languages, without storing anything** is a different problem than cloning a voice offline for a podcast. The big platforms optimize their translation for caption coverage and a single safe narrator voice — that's the cheap, robust default at scale. Keeping each speaker's own voice means the synthesis stage has to track every participant independently and stay inside the same latency budget the rest of the pipeline lives under. We built the voice engine ourselves, on our own infrastructure, which is what makes that trade-off ours to make. (More on why the engine is our own code: [*What one InterMIND meeting is built from*](https://intermind.com/blog/what-one-intermind-meeting-is-built-from).) ## Where this is going: lip-sync Keeping your voice is one half of a bigger goal. The other half is your **face**. Right now you hear the other person in their own voice, but if you're on camera, their lips still move to the words they actually said — in a language you don't read. The next step is **lip-sync**: re-timing the speaker's mouth to the translated audio, so that on your screen they appear to be speaking *your* language. Put the two together and the whole point of this work comes into focus. Two people who share no common language sit across a video call and see and hear each other as if each were a native speaker of the other's language — same voice, same face, no interpreter in the middle, no robot reading a script. To be clear about status: **voice is live today; lip-sync is on the roadmap, not shipped.** We're calling out the destination because it's why the voice work matters — own-voice translation isn't the feature, it's the first half of "talk to anyone, in any language, as yourself." ## Where to hear it Own-voice translation is **live today, across all 21 voice languages** — the same languages listed in the [docs](https://intermind.com/docs/translation/languages). There's nothing to turn on separately: when translation is enabled in a meeting, participants automatically hear each other in their own voices. We'll be honest about where it stands: today the voice is already recognizably *you*, and the resemblance is something we're actively pushing closer. Go listen and judge for yourself. - **[Try the demo](https://intermind.com/demo)** — runs the live voice pipeline against your audio in any of the 24 languages. - **[See the quality numbers](https://intermind.com/benchmark)** — the same production pipeline, scored monthly against FLORES-200, with the full distribution published per language pair. - **[How it works, in the docs](https://intermind.com/docs/translation/own-voice)** — the short version of this post. A translated meeting should feel like the people who are actually in it talking to each other. Keeping your voice is how it gets there. # We earn only when you earn: how the InterMIND partner model works Our partner program fits in one sentence: **you never pay us before you've been paid. $0 upfront, $0 fixed — ever.** That sentence raises reasonable questions the moment you take it seriously. A share of *which* revenue? Do we see your clients? How many seats do you have to license? What happens in a quarter where you earn nothing? Who counts what, and how do you know the invoice is right? This post answers all of them, in order, with the arithmetic on the table. It exists so that you can understand the whole model before the first call — and so the first call can be about your contracts, not about our mechanics. ## Who this is for The program has two tracks, and they work differently: - **Language service providers (LSPs)** — interpreting agencies delivering remote or on-site interpretation, from national framework holders down to single-interpreter practices. You get the platform as a **white-label delivery tool**: your brand, your billing, your client. We take a revenue share on what you bill — nothing else. - **IT and AV integrators** — companies that build and win tenders in verticals like courts and healthcare, where a translation capability is a line in a larger bid. This is a classic **reseller** track: the entire presale phase is free, and you buy at a wholesale price only after you've won. Most of this post walks through the LSP track, because that's where the mechanics are least familiar. The integrator track is summarized [below](https://intermind.com/#the-integrator-track-reseller). ## What costs you nothing — and stays that way Start with what is *not* billed, because this is where per-seat habits mislead. **Accounts are free and unlimited.** If your operation needs 10 seats to serve several clients at once — or 50, or 200 for every interpreter on your roster — you create them. There is no per-seat license, no seat tier, no seat audit. We don't sell access; access is how you get to the point where both of us earn. **Internal use is free.** Your team's own meetings, sales calls, project calls — all of it, including calls where you try the voice translation yourself. None of this appears on any invoice. **Onboarding and presale are free.** Setting up your workspace, demo sessions for a client you're pitching, test runs before a tender submission — free. If the deal doesn't happen, that cost was ours, not yours. There is no catch hiding here. The reason we can afford this is structural, and it's covered [below](https://intermind.com/#why-we-can-price-this-way). ## What we take a share of We earn a percentage of **your revenue on client contracts delivered through the platform**. Nothing else. The definition matters, so here it is precisely: a *platform contract* is a contract with your client where the platform was used to deliver the service — for example, a court framework where hearings are interpreted through your white-label workspace, or a healthcare contract where appointments run on it. What this means in the negative: - **Not a share of your company's revenue.** Your translation business, your on-site work delivered without the platform, contracts from before the partnership — none of our business, literally. - **Not a fee per minute, per meeting, or per seat.** Usage itself is never the billable unit. A busy month with no revenue attached to it costs you nothing. - **Not a payment from your client.** Your client pays you, on your contract, on your bank account, under your brand. We are not in the transaction and not visible in it. We do not know who your clients are unless you choose to tell us. The rate is agreed per partner and is not public. Two things about it are public: **volume moves it in your favour** — the more annual revenue you deliver through the platform, the lower the marginal rate, in tiered brackets that work like tax brackets (a lower rate applies to revenue above a threshold, never retroactively to the whole sum) — and **the rate depends on who does the interpreting**, which the next-but-one section explains. ## The money flow, step by step Here is a full quarter, end to end: 1. **You win and deliver contracts as usual.** Your client signs with you, pays you, sees your brand. Delivery runs through your platform workspace. 2. **Once a quarter, you declare one number**: your revenue on platform contracts for that quarter. One figure. No client names, no contract copies, no books opened. 3. **We send one invoice**: your declared figure × your rate. A regular invoice, paid by regular bank transfer. 4. **An empty quarter is a €0 invoice.** Didn't win the tender, client paid late, business was slow — you declare zero and owe zero. The principle at the top of this post is not marketing; it is the billing rule. Worked example with a placeholder rate: an agency holds a court framework and receives **€30,000** from the court in a quarter, delivered through the platform. It declares €30,000. At a rate of X%, our invoice is **€30,000 × X%**. That is the entire administrative burden of the relationship: one declared number in, one invoice out, four times a year. ## One number, no audits — how verification actually works A fair question: if you self-declare one figure, what stops anyone from under-declaring? And symmetrically, from your side: what stops *us* from disputing your figure forever? The answer is that the platform meters its own usage. Every session in your workspace, every translated minute, is counted by our infrastructure — we operate it, so we see it. Metered usage times market interpreting rates gives a rough expected revenue. If a declaration lands wildly outside that expectation, it's visible automatically — and it becomes **a conversation, not a court case**. Deliberately, the agreement contains **no audit rights**. We never ask to see your accounting, your client list, or your contracts. The declaration plus our own meter is the entire verification apparatus. This is also why the agreement itself fits on two pages: every clause we didn't add is a week of legal review you don't spend. ## Human interpreting and AI translation carry different rates The platform supports two ways of delivering a translated session, and they are economically different for *you* — so they carry different rates: - **Your interpreter works the session; the platform delivers it.** The interpreting is done by a human you pay. Your margin on that contract is what's left after the interpreter's fee — and our share has to fit inside it. Rate: the lower, tiered X% band. - **The platform's AI does the interpreting; no human interpreter is engaged.** Your largest cost line on that contract doesn't exist, so your margin per contract is a multiple of the human-delivered case. Rate: a higher Y% — and you still keep far more per euro than on human delivery. The higher percentage is not a penalty; it reflects that on these contracts the platform is doing the work you'd otherwise pay an interpreter for. **You do not classify anything.** Real contracts mix modes — a human for rare languages, AI for common ones; a human by day, AI for the night line. Splitting your revenue by mode would be exactly the accounting burden this model exists to avoid. So the split comes from our meter, not from your books: we measure the share of AI-translated minutes across your client sessions in the quarter, and blend the rate accordingly: > **effective rate = X% + (Y% − X%) × AI share of client-session minutes** Your quarterly invoice arrives with a one-page statement showing the inputs: client-session hours, AI-translated hours, the resulting blended rate, and the formula. You see where the number came from before you pay it. Three boundary rules keep the blend honest: 1. **Assistive AI doesn't count as AI delivery.** Captions or document translation running alongside your human interpreter don't move a session into the AI band. Only spoken AI interpreting counts. 2. **Only client sessions count.** Your internal meetings — including ones where your team uses AI translation — never enter the calculation. 3. **A gross mismatch is reviewable in both directions.** The minute-based blend is an approximation; if your revenue structure genuinely diverges from it, that's a conversation and an adjustment, not a dispute clause. []{#the-integrator-track-reseller} ## The integrator track (reseller) For integrators the model is simpler, because your business is bids, not recurring service delivery: - **The presale phase is free** — proof-of-concept, demo environments, and help with the translation section of your tender response. You pay nothing while the bid is open. - **You buy at a wholesale price only after you win**, resell at your price, and keep the margin. - **Your wholesale price is grandfathered for the full duration of your client contract.** Public-sector contracts run for years at fixed prices; your input cost will not move underneath a price you can no longer renegotiate. ## What protects you from us A platform partner's real worry isn't the rate — it's the platform deciding to go around them. Three commitments address that directly: - **Deal registration.** Register a deal you're working, and it's yours. - **No direct sales into your territory.** While the partnership is active, we don't sell directly into your registered vertical and country. - **12-month exclusivity** in your vertical and country is on the table for partners who commit to building it. And one more, less obvious: because your clients contract with *you*, under *your* brand, the client relationship is your asset. A white-label platform makes your business more valuable; it doesn't insert itself between you and what you've built. []{#why-we-can-price-this-way} ## Why we can price this way No trick, just cost structure. We built and operate our own translation engine and run meetings on our own infrastructure. Adding a partner — with all their seats, workspaces, and internal use — adds approximately nothing to our costs; our capacity is a step function we expand ahead of demand, not a per-seat license bill that scales with every account you create. That means giving you free unlimited access costs us nearly nothing, and taking our earnings only as a share of your *actual* revenue is safe for us at any volume. A vendor whose economics are built on per-seat licensing cannot make you this offer without breaking its own P\&L. We can, so we do — and we'd rather take that structural advantage to market as *your* zero-risk entry than as a discount percentage on a price list. ## How to start If you run interpreting contracts — or bid in verticals where translation is a required line — the conversation starts on the [partners page](https://intermind.com/partners). Bring one real contract or one live bid; the first working session is about mapping the model onto it. The agreement that follows is two pages, and you now know everything that's in it. ## FAQ **Do you take a share of my company's whole revenue?** No. The share applies only to revenue on contracts delivered through the platform. Your other business — translation work, on-site delivery without the platform, pre-existing contracts — is entirely outside the model, and we have no visibility into it. **Do you see or contact my clients?** No. Your client contracts with you, pays you, and sees your brand only. We are not part of the transaction and don't know who your clients are unless you tell us. While the partnership is active, we also commit not to sell directly into your registered vertical and country. **How many user accounts do I need to buy?** None — accounts are free and unlimited. If serving several clients at once takes 10 seats, or your whole interpreter roster needs access, you create the accounts and pay nothing for them. There is no per-seat license anywhere in the model. **What happens in a quarter where I earn nothing?** You declare zero and receive a €0 invoice. No minimums, no fixed fees, no "platform maintenance" line. A quarter without revenue on platform contracts is a quarter you pay nothing. **How do you verify my declaration without auditing me?** The platform meters its own usage — sessions and translated minutes in your workspace. Metered usage times market interpreting rates approximates expected revenue, so a wildly divergent declaration is visible automatically. The agreement contains no audit rights: a discrepancy triggers a conversation, not an inspection of your books. **What if a contract mixes human interpreters and AI translation?** You still declare one number. The human/AI split is measured by our meter — the share of AI-translated minutes in your client sessions — and your rate is blended accordingly. The invoice comes with a one-page statement showing the hours, the split, and the formula, so the math is on the table before you pay. **What exactly counts as a "platform contract"?** A contract with your client where the platform was used to deliver the service — hearings interpreted through your workspace, appointments run on it, sessions delivered under your brand on our infrastructure. Internal meetings and unpaid demo use never count, whatever features they use. **Are your rates public?** No — rates are agreed per partner. Two things are fixed in the structure: volume moves the rate in your favour (tiered, like tax brackets, never retroactive), and contracts where the platform's AI does the interpreting carry a higher rate than contracts your interpreters deliver — because on the former, your interpreter cost line doesn't exist. **Can you raise my rate after I've signed contracts against it?** Deals you've already registered and contracts you've signed keep their terms. Any revision applies to new deals only — the same grandfathering logic that fixes an integrator's wholesale price for the duration of a won contract. **Is there a minimum commitment, setup fee, or onboarding fee?** No. $0 upfront and $0 fixed is the billing rule of the program, not a promotional period. The first euro that moves from you to us is a percentage of a euro your client has already paid you. # Real-time meeting translation: how it works, and how to evaluate one **Real-time meeting translation** is a live meeting where each participant speaks, types, and listens in their own language — and the platform translates between them as the meeting happens, not afterwards. No human interpreter in a booth, no "let's just switch to English," no transcript you read the next morning. The category is full of tools that sound like they do this and don't. AI notetakers record and summarise. Caption add-ons subtitle the speaker. General-purpose models translate a block of text when you paste it in. Real-time meeting translation is a narrower, harder thing: every word, every chat message, every shared note, rendered into each listener's language fast enough that the conversation keeps flowing. This is the foundational guide to that category — what the term actually means, what happens under the hood, and the questions worth asking before you sign anything. It's the hub the rest of our writing branches off, so where a topic deserves its own deep dive, we link to it. --- ## What "real-time" actually rules out The hard constraint is latency. A live multilingual conversation works only if the translation arrives fast enough that people don't start talking over it. Past roughly **1.2 seconds end-to-end**, the meeting drifts — participants hesitate, double back, and eventually default to a shared second language. So real-time meeting translation has a sub-second budget that quietly disqualifies most of the tools marketed near it: - **AI notetakers** (think [Fireflies](https://intermind.com/compare/fireflies) or [Otter](https://intermind.com/compare/otter)) are built to transcribe and summarise a meeting — usually English-first, and most usefully *after* it ends. They answer "what did we decide." They do not translate speech live so a German speaker and a Japanese speaker each hear the other in their own language. That's a different job with a different clock. - **General-purpose LLM translation** is good prose translation with no latency contract. Fine for a document; wrong tool for a live audio channel where the model has under a second to respond and can't pause to "think." - **Caption/subtitle plugins** show text of what the speaker said, often in one target language for everyone. That's a captioning feature, not per-participant translation. If a tool can't translate *voice* live, per listener, into each listener's chosen language, it isn't doing real-time meeting translation — whatever the homepage says. --- ## The four things "translation" means in a meeting The second trap is treating translation as one feature. A live meeting actually has at least four translation jobs running at once, and they pull in incompatible directions: 1. **Voice** — audio in, translated audio out, under a second, every viewer in their own language. Constraint: latency. 2. **Chat** — short messages translated as they're sent, with edits that read like edits, not re-translations. 3. **Shared notes** — collaborative typing translated character-by-character, with lists, headings and checkboxes surviving intact. 4. **Documents** — a 40-page PDF dropped into chat, translated as a *file* with its tables, fonts and page breaks preserved. Constraint: fidelity, not speed. No single engine is good at all four — the latency budget that makes voice work is the opposite of the fidelity budget a document needs. Any honest platform runs several pipelines behind one language picker. We pulled ours apart in detail in [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines); the short version is that "one engine for everything" is a marketing simplification, not an architecture. --- ## How a real-time voice pipeline actually works Trace one sentence from a French speaker to a German, a Brazilian and a Japanese listener: 1. **Speech recognition runs in the speaker's browser**, locally — not on a central server. This shaves a network round-trip off the very first step and produces the source transcript with the lowest possible delay. 2. **The transcript fans out to the translation engine over one connection per target language present in the room.** If three people picked German, German shares one stream. If nobody picked Arabic, no Arabic stream opens, and idle streams drop after a few minutes. A four-language meeting costs what four languages cost — not forty. 3. **Each listener gets their own synthesised audio track**, mixed against the original speaker's video. Two people in the same physical room can wear headphones and hear different languages off the same meeting. The part that matters for procurement: **the engine doing the translating is its own thing, on its own infrastructure** — not a general-purpose third-party model the platform is quietly reselling. The sub-second budget rules those out, and so does the data-residency story for anyone regulated. (Where every byte of a meeting physically runs is its own question; we mapped it vendor-by-vendor in [*Where one InterMIND meeting actually runs*](https://intermind.com/blog/where-one-intermind-meeting-actually-runs).) --- ## How to evaluate a real-time meeting translation tool Most of this category competes on a single inflated number — "200+ languages," "99% accurate." Those tell you nothing about the meeting you're about to run. Here's what actually separates one tool from another. ### 1. Per-pair quality on real traffic, not an aggregate "200 languages" means a model emits text in 200 languages. Quality ranges from production-grade on major pairs to unusable on rare ones. Ask for **per-language-pair quality, measured on real traffic, with the distribution** — median, worst 10%, sample size — not one averaged headline. We argued why the whole category dodges this in [*Why translation-quality marketing is broken*](https://intermind.com/blog/why-translation-quality-marketing-is-broken), and we publish our own numbers at [`/benchmark`](https://intermind.com/benchmark): every live pair, every month, scored against [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""} by a [named judge](https://intermind.com/benchmark/methodology). You don't have to take a vendor's word — including ours. And when a meeting runs on one specific pair, check that pair, not the average — the [Hindi to English voice translator](https://intermind.com/blog/hindi-to-english-voice-translator) guide walks through doing exactly that for Hindi ↔ English before relying on it. ### 2. Latency you can feel, not a spec-sheet number Sub-second on voice is the threshold for a conversation that flows. Test it on a real call with real cross-talk, not a demo script. ### 3. Honest language counts, per surface A platform's voice languages, text languages and document languages are rarely the same set, and a single number hides that. Ours, for example: **24 languages live on voice, chat and notes; 30 on documents.** A French participant can request a contract PDF in Estonian even if they can't *listen* to the meeting in Estonian — and we flag that in the picker rather than smoothing it into one figure. The reasoning is in [*How many languages do you support?*](https://intermind.com/blog/how-many-languages-do-you-support). ### 4. Where the data runs For regulated buyers, *where* the meeting is processed is part of the spec, not a footnote. Ask which vendors touch the audio, where they execute, and which are merely reselling a US model. Our full runtime map — and the one post-meeting step that's still US-domiciled, named plainly — is in [*Where one InterMIND meeting actually runs*](https://intermind.com/blog/where-one-intermind-meeting-actually-runs) and [*Multilingual compliance meetings*](https://intermind.com/blog/multilingual-compliance-meetings). ### 5. Live translation vs. a notetaker Be clear which problem you're solving. If you need a record and a summary, an [AI notetaker](https://intermind.com/compare) is the right buy. If you need people who don't share a language to actually *talk*, you need live translation. Some teams want both; few tools do both well. ### 6. What your current platform already does Before buying anything, know exactly what's built into the tool you already pay for — and where it stops. We keep honest, sourced how-tos for each: [Zoom](https://intermind.com/blog/zoom-live-translation), [Microsoft Teams](https://intermind.com/blog/teams-live-translation), and [Google Meet](https://intermind.com/blog/google-meet-live-translation). Each one ends at the same place: captions or a fenced bilingual feature, not a room where everyone hears their own language. --- ## Where InterMIND fits We built InterMIND for the live-translation job specifically: [real-time voice, chat, notes and documents](https://intermind.com/features/realtime-translation) across **24 languages**, on our own engine hosted in the EU, with the translation quality published openly instead of asserted. It's a web app (no install), up to 1080p video, up to 1500 participants — enough for [simultaneous interpretation at conferences and webinars](https://intermind.com/features/simultaneous-interpretation), not just calls — with cloud and local recording. It is not the best tool for "transcribe and summarise my English standup" — that's what notetakers are for, and we say so on the comparison pages rather than pretending otherwise: - [InterMIND vs. Fireflies.ai](https://intermind.com/compare/fireflies) — notetaker vs. live multilingual meeting - [InterMIND vs. Otter.ai](https://intermind.com/compare/otter) — same axis, honestly compared - [All platform comparisons](https://intermind.com/compare) If literal, word-for-word fluency is your worry, [*The false-fluency trap*](https://intermind.com/blog/false-fluency-trap) is the one to read — fast translation that's confidently wrong is worse than slow translation that's right. --- ## Try it yourself - **[Try the live demo](https://intermind.com/demo)** — runs the production voice pipeline on your own audio, in any of the 24 live languages, and scores it against the same judge that scores the public benchmark. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month quality on real traffic. Every pair in the picker, strong or weak, deep-linkable by URL. - **[Read the methodology](https://intermind.com/benchmark/methodology)** — exactly what the numbers measure, what they don't, and who the judge is. Real-time meeting translation is a narrow promise: everyone in their own language, live, fast enough to keep talking. The honest way to evaluate it is to stop reading homepages and start measuring. That's what the links above are for. --- ## FAQ **What is real-time meeting translation?** A live meeting where each participant speaks, types, and listens in their own language, and the platform translates between them as the meeting happens — voice, chat, notes, and documents — fast enough (sub-second on voice) that the conversation keeps flowing. **How is it different from an AI notetaker like Otter or Fireflies?** A notetaker transcribes and summarises a meeting, mostly after it ends. Real-time meeting translation changes what participants *hear during* the call: a German speaker and a Japanese speaker each hear the other in their own language, live. **What latency does live voice translation need?** Past roughly 1.2 seconds end-to-end a conversation drifts — people hesitate and fall back to a shared language. The practical target is sub-second per-listener audio. **How do you evaluate translation quality claims?** Ask for per-language-pair quality measured on real traffic — median, worst 10%, sample size — not one averaged headline. We publish ours monthly at [`/benchmark`](https://intermind.com/benchmark). — The Mind.com Team --- *Sources: [Otter — supported languages](https://help.otter.ai/hc/en-us/articles/360047247414-Supported-languages){rel=""nofollow""}, [Fireflies — supported languages](https://guide.fireflies.ai/articles/2973706448-learn-about-fireflies-supported-languages){rel=""nofollow""}, [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""}, checked August 2026. Vendors change plans and language lists over time — check their pages for the current state.* # Simultaneous interpretation: what it is, how it works, and when AI can do it **Simultaneous interpretation is translation delivered while the speaker is still talking.** Listeners hear the speech in their own language with a delay of a few seconds at most, and the meeting never stops to wait for the interpreter. It's how the UN runs its sessions, how multilingual conferences keep one agenda instead of three, and — since AI joined the field — increasingly how ordinary working meetings run too. This guide explains the whole category in plain language: how the traditional version works and what it costs, how it differs from consecutive interpretation, what "remote simultaneous interpretation" (RSI) changed, what AI can and can't do yet, and how to get simultaneous interpretation into the tools you already use. > Comparing vendors instead? Start with the [buyer's guide to AI translation tools for conferences and meetings](https://intermind.com/blog/best-ai-conference-translation-tools). --- ## What simultaneous interpretation actually involves Professional simultaneous interpreting is one of the most cognitively demanding jobs in language work. The interpreter listens to the source language, converts meaning (not words) on the fly, and speaks the target language — all at once, continuously, while the speaker keeps going. The industry has built serious infrastructure around that difficulty: - **Interpreters work in pairs**, one active and one supporting, swapping roughly every 30 minutes — sustained simultaneous work degrades fast beyond that. - **A soundproof booth** (or interpreting console) isolates the interpreter, with the room's audio in their headphones and their output routed to listeners' receivers. - **One booth and one pair per language.** A three-language event means three booths and six interpreters. - **Booking runs days to weeks ahead**, longer for rare pairs, and billing is per interpreter, per language, per day — plus equipment. None of that is a flaw. It's what delivering professional-grade live interpretation costs. The consequence is simply that simultaneous interpretation became an *event-day* service: budgeted, scheduled, and reserved for the sessions that justify it. ## Simultaneous vs consecutive interpretation The other classic mode is **consecutive interpretation**: the speaker says a few sentences, pauses, and the interpreter renders them in the target language. It needs no equipment and shines in small, controlled settings — a medical consultation, a notarized signing, a site visit. | | Simultaneous | Consecutive | | --------------------- | ----------------------------------------- | --------------------------- | | **Timing** | Parallel with the speaker | Alternates with the speaker | | **Meeting length** | Unchanged | Roughly doubles | | **Equipment** | Booth/console + receivers (or a platform) | None | | **Best at** | Conferences, working meetings, broadcasts | Short two-party exchanges | | **Languages at once** | Many (one channel per language) | Usually one pair | There are niche variants — *whispered* interpretation (chuchotage) for one or two listeners, *liaison* for back-and-forth negotiation — but nearly every real decision is simultaneous vs consecutive, and for anything with an agenda and more than a handful of people, simultaneous wins on time alone. (All the modes and delivery formats, plus how interpretation differs from translation, are laid out in [our definitional guide](https://intermind.com/blog/interpretation-vs-translation).) ## Remote simultaneous interpretation (RSI): the first cost collapse **RSI moves the interpreter out of the on-site booth and into the cloud.** The event's audio streams to interpreters working remotely; their translated audio streams back to each listener's phone, laptop, or headset. Platforms like Interprefy and KUDO built the category, and it took off when events went hybrid. What RSI eliminates: booth rental, receiver hardware, interpreter travel, and part of the lead time. What it keeps: the human interpreters themselves, their per-day rates, and the per-event operational setup. RSI made professional interpretation *cheaper to deliver* — it didn't change *who* does the interpreting or the booking model. (We compare the leading RSI service directly in [InterMIND vs Interprefy](https://intermind.com/compare/interprefy).) ## AI simultaneous interpretation: the second cost collapse The newer shift replaces the interpreting pipeline itself: **speech recognition → machine translation → speech synthesis**, running continuously, per listener, with sub-second latency. No booth, no booking, no per-day rate — interpretation becomes a *feature of the meeting software*, which changes what it can be used for. The weekly sync between Berlin, São Paulo, and Tokyo was never going to book two interpreters per language; with AI it doesn't have to. Three things distinguish AI interpretation done seriously (we build one of these tools, so we'll say plainly what to demand from any of them, [ours included](https://intermind.com/features/simultaneous-interpretation)): 1. **Per-pair quality you can verify.** "Supports 60 languages" says nothing about *your* DE↔EN. Ask for published, per-language-pair quality measured on real traffic — [here's ours, updated monthly](https://intermind.com/benchmark) — or [run a live demo](https://intermind.com/demo) on your own voice and judge. 2. **Whose voice comes out.** Most tools narrate every speaker with one synthetic voice. The better experience keeps each speaker's own voice via zero-shot synthesis — a discussion, not a transcript read aloud. ([How that works.](https://intermind.com/blog/own-voice-translation)) 3. **How much of the meeting is covered.** A meeting is voice *plus* chat, shared notes, and documents. If only the audio is translated, the room still splits into languages the moment someone types. ([The full argument.](https://intermind.com/blog/best-ai-conference-translation-tools)) And the honest limit, stated without hedging: **courts, treaty negotiations, certified depositions, and any setting where the interpretation is the legal record still belong to accredited human interpreters.** AI's territory is the enormous space below that bar — the meetings, webinars, and conferences that never got interpretation at all because booths and day rates priced them out. ## Simultaneous interpretation in Zoom, Teams, and Google Meet The platforms you already use sit at three different points: - **Interpretation channels (Zoom, Teams):** built-in audio channels for interpreters you bring and pay yourself. The platform solves delivery, not interpreting. How-to and limits: [Zoom](https://intermind.com/blog/zoom-live-translation), [Teams](https://intermind.com/blog/teams-live-translation). - **Translated captions:** all three platforms can subtitle a meeting in another language on the right plan. Reading is not hearing — captions work for a webinar, poorly for a discussion. Details per platform: [Zoom](https://intermind.com/blog/zoom-live-translation), [Teams](https://intermind.com/blog/teams-live-translation), [Google Meet](https://intermind.com/blog/google-meet-live-translation). - **Built-in AI speech translation:** early and partial (Meet's Gemini speech translation is the furthest along). Language pairs and plan gating decide whether your meeting qualifies. If you need every participant to *hear* the meeting in their own language, both directions, without booking anyone — that's the job [AI simultaneous interpretation](https://intermind.com/features/simultaneous-interpretation) does in the browser, and the platform comparisons above show exactly where each built-in option stops. ## Choosing, in one pass - **Certified, high-stakes, on the record** → accredited human interpreters, on-site or via RSI. - **Large formal event, human quality, hybrid audience** → an RSI platform ([vs Interprefy](https://intermind.com/compare/interprefy)); AI event delivery is the budget alternative ([vs Wordly](https://intermind.com/compare/wordly)). - **Working meetings, webinars, everyday multilingual calls** → AI interpretation built into the meeting: [every listener picks a language](https://intermind.com/features/simultaneous-interpretation), the meeting runs at full speed, and the chat, notes, and documents come back translated too. - **Occasional light need on one platform** → try that platform's captions first ([Zoom](https://intermind.com/blog/zoom-live-translation) / [Teams](https://intermind.com/blog/teams-live-translation) / [Meet](https://intermind.com/blog/google-meet-live-translation)) and upgrade when reading stops being enough. --- ## FAQ **What does simultaneous interpretation mean?** Interpretation delivered in parallel with the speaker — listeners hear the translation while the speech is still happening, typically a few seconds behind, so the meeting doesn't pause for translation. **How many interpreters does simultaneous interpretation need?** Two per language pair, swapping every \~30 minutes, per industry standard. A three-language conference typically staffs six interpreters plus equipment. **What does simultaneous interpretation cost?** Human: per interpreter, per language, per day — a one-day two-language event commonly runs into thousands of dollars with equipment. AI: a software subscription; the marginal cost of an interpreted meeting is zero. ([Pricing.](https://intermind.com/pricing)) **Is simultaneous interpretation possible in Zoom?** Yes, two ways: Zoom's interpretation channels with human interpreters you book yourself, or AI translation. Zoom's own AI features are captions-first — the full picture is in [our Zoom guide](https://intermind.com/blog/zoom-live-translation). **Can AI replace simultaneous interpreters?** In certified high-stakes settings — no, and vendors who claim otherwise should worry you. In everyday meetings, webinars, and conferences, AI now does the job well enough to be verified rather than assumed: check [per-pair published quality](https://intermind.com/benchmark) or [test it live](https://intermind.com/demo). --- ## Hear it, don't read about it Simultaneous interpretation is an audio experience, and no article settles whether the quality clears your bar. - **[Run the live demo](https://intermind.com/demo)** — speak, and hear yourself in another language, in your own voice, on the production pipeline. - **[Read the benchmark](https://intermind.com/benchmark)** — monthly per-language-pair scores on real traffic, full distribution. - **[The feature, in detail](https://intermind.com/features/simultaneous-interpretation)** — how AI simultaneous interpretation works in an InterMIND meeting. — The Mind.com Team --- *Sources: [Chmiel — Boothmates forever? On teamwork in a simultaneous interpreting booth](https://akjournals.com/abstract/journals/084/9/2/article-p261.xml){rel=""nofollow""} (two interpreters per booth, \~30-minute turns), [Interprefy — platform](https://www.interprefy.com/){rel=""nofollow""}, [KUDO AI Speech Translator](https://kudo.ai/solutions/kudo-ai-speech-translator/){rel=""nofollow""}, [Zoom — Enabling and configuring translated captions](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0059081){rel=""nofollow""}, [Google Meet — Speech Translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""}, checked August 2026. Vendors change plans and language lists over time — check their pages for the current state.* # Language translator: how a voice translator actually works, and where it stops "Language translator" is a word that covers four quite different tools. A **text translator** (DeepL, Google Translate) turns typed text from one language into another. A **phone voice translator** listens, then reads a translation back. **Earbud translators** whisper a translation into your ear. And **live meeting translation** runs a whole conversation in several languages at once, each participant hearing it in their own. They all get called the same thing, and they are not interchangeable. This guide separates them, explains how the voice ones work under the hood, and names the line where each one stops. --- ## The four kinds, briefly ### Text translator The workhorse for anything written: paste a paragraph, get it back in another language. Built for documents, emails, a website. It doesn't touch speech, and it's one block at a time — not a conversation. ### Phone voice translator Open an app, speak, and it plays back (or shows) the translation, then the other person answers into the same phone. Perfect for a phrase — a menu, a direction, a hotel desk. The mechanic is a **relay**: speak → transcribe → translate → speak back, one turn at a time. It works because the exchanges are short and you pass the device back and forth. ### Earbud translator AirPods with Live Translation, Pixel Buds, dedicated translator earbuds. You wear them and hear the other person translated in your ear. Nicer than staring at a screen — but it's **one-to-one, and needs matching gear on both sides.** Built for a traveler and a local, not for a room of people. ### Live meeting translation The different one. Here the *room* is multilingual: everyone speaks their own language and hears everyone else in **their** language, live, for the whole meeting. No passing a phone, no matching earbuds. It's the only one of the four that holds a real conversation together. --- ## How a voice translator actually works Under the hood, a spoken-language translator is a short pipeline: 1. **Speech recognition** turns what you said into text. 2. **Translation** converts that text into the target language. 3. **Speech synthesis** (for the voice kind) speaks the result. Two things decide whether it feels live or clunky: - **When it translates.** If it waits for you to finish, then translates, you get consecutive interpretation with a synthetic voice — you feel every gap. If it keeps pace as you talk, it feels live. - **Who it translates for.** A phone app translates *for the two people holding it*. Live meeting translation translates the *whole room for everyone*, each into their own language, simultaneously — that's a harder problem and a different architecture. The full version of how the live kind works is in our foundational guide, [*Real-time meeting translation: how it works, and how to evaluate one*](https://intermind.com/blog/real-time-meeting-translation), and the engineering detail is in [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines). --- ## Which one do you actually need? - **Something written?** — a text translator. - **A phrase, one exchange with a stranger?** — a phone voice translator. Don't over-buy. - **Face-to-face with one person, both of you equipped?** — earbuds. - **A conversation — a call, a meeting, several people, back-and-forth?** — live meeting translation, because the others turn every exchange into a relay and every extra person into a broken assumption. --- ## Where InterMIND fits InterMIND is the fourth kind: a language translator for live conversations, built as real-time meeting translation. - **24 languages live on voice**, chat and shared notes — any mix, no English anchor, no regional gate, no five-language beta. ([The full end-to-end translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Per-listener audio, sub-second.** Each participant hears the meeting in their own picked language at the same time — five people, five languages, one call. - **Webinars and conferences included** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — drop a PDF or DOCX into the meeting and each viewer gets it in their language, 30 languages on files. - **Quality you can audit.** Per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark), with the methodology written down. A text translator is right for a document. A phone app is right for a phrase. Earbuds are right for one person in front of you. When it's a *conversation*, that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and *hear* per-listener translation instead of reading it. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, methodology included. - Start with the category guide: [*Real-time meeting translation*](https://intermind.com/blog/real-time-meeting-translation). - Comparing specific language pairs? See the voice-pair guides for [English ↔ Italian](https://intermind.com/blog/traduttore-inglese-italiano-vocale), [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale), [Romanian ↔ Italian](https://intermind.com/blog/traduttore-rumeno-italiano-vocale) and [Hindi ↔ English](https://intermind.com/blog/hindi-to-english-voice-translator) — or the [Arabic voice translator](https://intermind.com/blog/mutarjim-sawti) guide. --- ## FAQ **How does a voice translator work?** It's a three-step pipeline: speech recognition turns what you said into text, machine translation converts the text, and speech synthesis speaks the result. Whether it feels live depends on *when* it translates (as you talk vs. after you pause) and *who* it translates for (two people holding a phone vs. a whole room). **What's the difference between a phone translator app and live meeting translation?** A phone app is a relay — speak, wait, pass the device — built for one phrase at a time. Live meeting translation runs the whole conversation in several languages at once, each participant hearing every speaker in their own language, simultaneously. **Do translator earbuds work for a group conversation?** Earbud translation (AirPods Live Translation, Pixel Buds) is one-to-one and generally needs matching gear or a shared screen on the other side. For a call or meeting with several people it breaks down — that's the job of live meeting translation. — The Mind.com Team --- *Sources: [Apple — Live Translation with AirPods](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google — Translate with Pixel Buds](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. Vendors change features and language lists over time — check their pages for the current state.* # Microsoft Teams live translation: how it works, and where it stops If you searched "Microsoft Teams live translation," the honest short answer is: **yes, Teams can translate a live meeting — three different ways — and which one you get depends on whose license the organizer has and what kind of meeting you booked.** Microsoft has shipped more live-translation machinery than any other conferencing vendor, including a genuinely ambitious AI interpreter. It has also fenced each piece differently, and the fences are where buying decisions actually happen. This post explains how each works, how to turn it on, and exactly where the ceiling is. > This is the platform how-to companion to our foundational guide, [*Real-time meeting translation: how it works, and how to evaluate one*](https://intermind.com/blog/real-time-meeting-translation). For the Google Meet and Zoom versions of this post, see [Meet](https://intermind.com/blog/google-meet-live-translation) and [Zoom](https://intermind.com/blog/zoom-live-translation). New to interpretation as a category? Start with [*Simultaneous interpretation: booth, RSI, or AI*](https://intermind.com/blog/simultaneous-interpretation-guide). --- ## The three features, and why they're not the same thing ### 1. Live translated captions (text) The established feature. Teams transcribes the meeting live and translates the caption stream into each participant's chosen language — **every attendee picks their own caption language independently**. Microsoft's support page lists **31 translation languages** for meetings (its own licensing page says 40, and the town-hall page says "over 50" — we'd cite the 31-language list it actually enumerates). Works in meetings, webinars, and town halls — though in town halls the organizer pre-selects a pool of six languages, ten with Premium. The gate: **the organizer needs a Teams Premium or Microsoft 365 Copilot license** for attendees to get translated captions. Plain English captions are free; translation is the paid layer. And captions aren't saved — when the meeting ends, the translation evaporates. ### 2. Interpreter agent (audio, AI) The Ignite 2024 headliner, now generally available: an AI agent that translates **spoken audio into spoken audio**, per participant — you pick the language you want to *listen* in, and Interpreter can even **simulate the speaker's own voice** in the translation. Calls got it in January 2026; a turn-based consecutive mode for two-language meetings rolled out in May 2026. Of the big three platforms, this is architecturally the closest thing to per-listener simultaneous interpretation. Which is exactly why its fences matter — they're in the limits section, and they're decisive. ### 3. Language interpretation (humans, channels) The classic: the organizer pre-configures up to **16 language pairs** and invites **human interpreters** — people you source, brief, and pay — who each get an audio channel. Attendees pick a channel and balance original vs. interpreter volume with a slider. No premium license documented, but the production burden is yours, and the constraints are real: set up before the meeting starts, interpreters must join from the desktop app, no web support, no breakout rooms, no end-to-end-encrypted meetings. --- ## How to turn each one on **Translated captions:** organizer has Premium or Copilot → any attendee turns on live captions, opens the caption settings, and picks a language under *Translate to*. **Interpreter agent:** you (the listener) need a **Microsoft 365 Copilot license**. In a scheduled meeting, turn on Interpreter for yourself, pick the language under *Listen to meeting in*, and optionally let it simulate voices (your admin controls whether that's on by default). **Language interpretation:** organizer enables it *while scheduling*, assigns interpreters per language pair, and attendees pick their channel once the meeting starts. It cannot be added to a meeting that's already running. The friction isn't the toggles. It's what the toggles can and can't do. --- ## The limits that actually decide if it fits As of mid-2026, by Microsoft's own documentation: - **The Interpreter agent speaks 10 languages** — English, Spanish, Portuguese, Japanese, Mandarin, Italian, German, French, Korean, and (since April 2026) Traditional Chinese. Under the hood it **pivots through English text**: speech is recognized, converted to English, translated, then synthesized. Every non-English pair pays the relay toll. - **It's metered and licensed per user.** Interpreter comes with a **Microsoft 365 Copilot license** and includes **20 hours per user per month**; beyond that, access is "subject to available capacity." Multilingual meetings as a metered utility. - **It skips the meetings where you'd want it most.** No ad-hoc/instant meetings, **no webinars, no town halls**, no Teams Free, no Teams Rooms on Android. The big multilingual broadcast formats get translated captions only. - **Nothing survives the meeting.** Recordings capture **original audio only** — no interpretation track. Translated transcripts exist only live; afterwards, only the spoken-language original remains. Translated captions aren't saved either. - **Microsoft's own quality caveat:** Interpreter is "not optimized for rapid exchanges or overlapping dialogue" — that is, for the way real working meetings sound. No latency figure is published. - **Human interpretation is a production.** Pre-configured only, desktop-only interpreters, no web attendees, no breakout rooms, no E2EE — and the interpreters bill you per day, per language, [a cost structure of its own](https://intermind.com/features/simultaneous-interpretation). None of this makes Teams' stack bad — it's the most complete of the big platforms. But it makes it **a per-seat-licensed, 10-language, 20-hour-a-month interpreter for scheduled internal meetings, with captions for everything else.** --- ## The meeting without the license matrix If those limits are the problem, the fix isn't a better toggle — it's a meeting that's multilingual for everyone in it by default: - [Start or join a multilingual meeting](https://intermind.com/meetings) — every participant in their own language, with no per-listener license and no 20-hour meter. - [How the real-time translation stack works](https://intermind.com/features/realtime-translation) — 24 languages, translated directly between pairs, so a French↔Japanese meeting never routes through English. --- ## The structural ceiling, in one sentence **Teams will interpret a scheduled meeting for the colleagues whose seats carry Copilot licenses — in 10 languages, through an English relay, 20 hours a month, leaving no translated record.** A genuinely multilingual organization needs the *meeting* to be multilingual — for everyone in it, including the webinar audience, the guest without your license, and the recording you keep. If your multilingual meetings are internal, scheduled, and inside the 10 languages — and your organization is already paying for Copilot — Interpreter is a real answer, built into a tool you already have. If your meetings include webinars, externals, language pairs that shouldn't route through English, or anything you need a record of, you've hit the ceiling. --- ## When you've outgrown the license matrix This is the job we built InterMIND for: **the meeting itself runs in every participant's language — no per-listener license, no meter, no English relay.** Concretely, where Teams gates and meters: - **24 languages live on voice**, chat and shared notes, translated **directly between pairs** — a French↔Japanese meeting never routes through English. ([The full end-to-end stack](https://intermind.com/features/realtime-translation) is its own page.) - **Everyone in the meeting gets it.** Translation is how the platform works, not a per-seat entitlement — guests and externals included, no 20-hour meter. - **Webinars, town halls and conferences included** — up to 1,500 participants, each picking their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). The full event workflow — Q\&A, recap and slide decks included — is at [events & webinars](https://intermind.com/use-case/events). - **The record survives.** Multilingual transcript, recording, and translated documents bundled after the meeting — not captions that evaporate. - **Quality you can audit.** Per-language-pair scores on real traffic, published monthly at [`/benchmark`](https://intermind.com/benchmark) — not asserted, measured. We're not claiming Teams is bad — Interpreter is the most serious translation feature any incumbent has shipped. We're claiming the fences around it describe a different meeting than the one multilingual teams actually run. The feature-by-feature version is at [InterMIND vs. Microsoft Teams](https://intermind.com/compare/microsoft-teams). --- ## Try the other side of the ceiling - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and hear per-listener translation with no license matrix. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, with the methodology written down. - [InterMIND vs. Microsoft Teams](https://intermind.com/compare/microsoft-teams) — honest, feature-by-feature. - Just need a record, not live translation? The notetaker side of the market: [the 9 best Otter.ai alternatives](https://intermind.com/blog/otter-ai-alternatives). Teams live translation is real, ambitious, and precisely fenced: captions for the many, an English-relay interpreter for the licensed few, human channels for the events that justify staffing. Knowing where the fences are is the whole decision. --- ## FAQ **Does Microsoft Teams translate speech or only captions?** Both, behind different gates. Live translated captions cover 31 languages as text. The Interpreter agent speaks translations as synthetic audio in 10 languages — but it pivots through English text, and each user needs the right license. **What license do you need for Teams live translation?** Translated captions need the meeting organizer on Teams Premium or Microsoft 365 Copilot — then all participants get them. Interpreter is metered per user: 20 hours per user per month with a Copilot license. **Does Teams Interpreter translate between any two languages?** No. It routes through English text, and it doesn't work in ad-hoc calls, webinars, town halls, Teams Free, or on Teams Rooms on Android. Microsoft also notes it isn't optimized for rapid exchanges or overlapping dialogue. **Are translated captions saved in the Teams transcript or recording?** No. Captions aren't saved, translated transcripts exist only live, and recordings keep the original audio only. — The Mind.com Team --- *Sources: [Microsoft — Use live captions in Teams meetings](https://support.microsoft.com/en-us/office/use-live-captions-in-microsoft-teams-meetings-4be2d304-f675-4b57-8347-cbd000a21260){rel=""nofollow""}, [Microsoft Learn — Meeting transcription and captions](https://learn.microsoft.com/en-us/microsoftteams/meeting-transcription-captions){rel=""nofollow""}, [Microsoft — Interpreter in Teams meetings and calls](https://support.microsoft.com/en-us/office/interpreter-in-microsoft-teams-meetings-and-calls-c7efe2bb-535d-42ab-a5c4-d2d91619b46d){rel=""nofollow""}, [Microsoft Learn — Interpreter agent in Teams](https://learn.microsoft.com/en-us/microsoftteams/interpreter-agent-teams){rel=""nofollow""}, [Microsoft — Use language interpretation in Teams meetings](https://support.microsoft.com/en-us/office/use-language-interpretation-in-microsoft-teams-meetings-b9fdde0f-1896-48ba-8540-efc99f5f4b2e){rel=""nofollow""}, [Microsoft Learn — Teams Premium licensing](https://learn.microsoft.com/en-us/microsoftteams/teams-add-on-licensing/licensing-enhance-teams){rel=""nofollow""}. Microsoft expands language lists and licensing over time; check the pages for the current state. All facts checked August 2026.* # How much of your meeting still exists in your language a week later? Run this test on your last cross-language meeting: **one week later, what can each participant still open and read in their own language?** Not "was the meeting translated." Almost every platform can now answer yes to that, with an asterisk. The test is what remains. The decision your Portuguese colleague wants to quote back in a dispute. The action item your German teammate needs to re-read before the follow-up. The three weeks of context a new hire from Bogotá has to absorb. The number procurement asks you to confirm from the call two Tuesdays ago. A meeting lasts forty-five minutes. Its outputs get used for months. And on today's major platforms, translation is built for the forty-five minutes — the moment the call ends, the translated layer disappears, and the record that remains speaks one language. This post is the audit: what each vendor's own documentation says survives the call, and in whose language. Then the architectural reason the answer comes out the way it does — and what it looks like when a meeting is built to outlive the call instead. --- ## A meeting's real lifespan The industry measures meeting translation on two axes: how fast (latency) and how wide (language count). Both describe the live call. We've argued elsewhere that a third axis decides more purchases than either — [how much of the meeting comes back in your language](https://intermind.com/blog/real-time-meeting-translation) — and the week-later test is that axis stretched to its natural length. Because the call is the shortest-lived thing the meeting produces. What actually circulates afterwards: - **Decisions** — cited in later meetings, escalations, and audits, often word-by-word. - **Action items** — re-read by the people executing them, days later. - **History** — the searchable record a newcomer reads to catch up, and the reference the team searches when memory disagrees. - **Documents** — the contract, the spec, the deck that the meeting was about. If translation exists only during the call, then every one of those artifacts exists in one language — usually English — and everyone else works from memory or from a second-hand summary. The organization's *memory* is monolingual even when its meetings weren't. A participant who heard the meeting in Spanish spends the following week working with a record they can only read in English. --- ## What survives the call, vendor by vendor Everything below is from vendor public documentation, checked August 2026. Links in the footer. ### Microsoft Teams Teams documents the live layer thoroughly — translated captions in 31 languages, an AI Interpreter agent in 10 — and is equally explicit about the afterlife: - **Translated captions aren't saved.** Microsoft's caption documentation states captions are not stored after the meeting; the translation exists only while the meeting runs. - **Recordings capture original audio only** — the Interpreter agent's translated audio is not in the recording. - **Translated transcripts exist only live.** Afterwards, the transcript that remains is in the spoken language. The full walkthrough of the live features is in our [Teams live translation](https://intermind.com/blog/teams-live-translation) post; the summary here is one sentence in Microsoft's own docs: nothing translated survives the meeting. ### Google Meet Meet's Gemini-powered speech translation translates spoken audio between English and five languages, one pair per meeting. On persistence, Google's documentation is direct: - **Speech translation is not available in recordings** (or live streams) — it's a live-only beta, capped at 90 minutes per meeting. - **No audio is saved.** Google states plainly that speech-translation audio isn't stored — a clean privacy answer, and also a complete answer to the week-later test. - Translated captions, the older text feature, are likewise an on-screen live view. Details and limits in [Google Meet live translation](https://intermind.com/blog/google-meet-live-translation). ### Zoom Zoom's live stack is the broadest of the three — translated captions in 36 languages, a five-language Voice translator, human interpretation channels — and its documentation describes all three as live views: - **Translated captions** are rendered on screen, per participant, while the meeting runs. Zoom's translated-captions documentation does not state that the translated stream is saved after the meeting (checked August 2026). - **The Voice translator** speaks synthetic translated audio live, in meetings only, on the desktop app. Its documentation likewise does not state that translated audio persists anywhere (checked August 2026). - What Zoom's recording stack keeps is the meeting *as spoken* — the original audio and its transcript. The live features are mapped in [Zoom live translation](https://intermind.com/blog/zoom-live-translation). ### The AI notetakers Otter, Fireflies and their category are the mirror image of the platforms: built precisely for the afterlife — transcripts, summaries, search — which makes their language model the interesting part: - **The transcript persists in the language the meeting was held in.** [Otter](https://intermind.com/compare/otter) transcribes in six languages (English, Spanish, French, German, Japanese, Chinese); [Fireflies](https://intermind.com/compare/fireflies) claims 100+ with auto-detection. Either way, one meeting produces one transcript in one language — the spoken one. - A meeting held in English produces an English record, however many languages its *readers* think in. The notetaker solves "we forgot what was said," not "half of us read this at half speed." The category comparison is in [Otter alternatives](https://intermind.com/blog/otter-ai-alternatives) and [Fireflies alternatives](https://intermind.com/blog/fireflies-ai-alternatives). --- ## The week-later audit, in one table | Platform | Translation during the call | What its docs say survives the call | A week later, in *your* language | | ----------------- | -------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | | Microsoft Teams | Translated captions (31 languages, organizer needs Premium/Copilot); Interpreter agent (10 languages, metered) | Captions not saved; recordings keep original audio only; translated transcripts live-only | Nothing translated | | Google Meet | Speech translation (English ↔ 5 languages, one pair per meeting, 90-min beta); translated captions | Speech translation unavailable in recordings; no audio saved | Nothing translated | | Zoom | Translated captions (36 languages); Voice translator (5 languages, add-on, desktop, meetings only) | Recording and transcript of the meeting as spoken; docs don't state translated output is saved | Nothing translated stated in documentation | | Otter / Fireflies | None (notetakers don't translate the live meeting) | Transcript + summary — in the spoken language | The record persists; the language doesn't change | | InterMIND | Voice, chat, notes and documents translated per participant, live | The meeting persists as a channel: chat history readable in each member's language, a recap delivered to each guest in their own language, documents translated on request | The history, in the language you think in | The pattern in the middle column isn't an accident, and it isn't vendors forgetting a feature. It's architecture. --- ## Why translated moments can't become translated history A caption is a *view* — text rendered over a live stream for as long as the stream exists. When the call ends, the stream ends, and a view of nothing is nothing. The same holds for translated audio synthesized on the fly: it's played, not kept. Translation implemented as an overlay dies with the thing it overlays. That's not a bug in Zoom, Teams or Meet; it's the honest consequence of where translation sits in their architecture — a layer on the call. For a meeting's history to exist in every reader's language, translation has to be a property of the *place that holds the history*, not of the call that produced it. The chat message, the recap, the document need to live somewhere that renders them per reader — this week, next month, for the member who joins in October. That's the difference we mean when the InterMIND homepage says "[a space for everyday communication](https://intermind.com) " rather than "video calls with translation." It's also why we treat a meeting becoming a channel as the product working as intended, not as a leftover artifact. ## What a space that outlives the call looks like Concretely, in InterMIND: - **The meeting persists as a channel.** The call is an event inside the space; when it ends, the space — members, chat, files, history — stays. - **Chat history reads in each member's language.** Messages are [translated per viewer](https://intermind.com/features/multilingual-chat), and that applies to the history, not just to messages arriving live. - **Every guest gets the recap in their own language.** After the call, the [AI recap](https://intermind.com/features/recap) — summary, decisions, action items — is delivered per participant, each in theirs. - **Documents translate on request** — [30 languages](https://intermind.com/features/document-translation), inside the same space as the meeting they belong to. - **Voice comes back as a recap, honestly labeled.** We removed live transcription deliberately; the persistent trace of what was *said* is the recap, not a word-for-word multilingual replay. If your compliance case needs a verbatim record, that's what recording — under the host's control — is for. - **You choose the jurisdiction the space lives in.** History that persists is history that's stored; where it's stored is a first-class setting, with an [EU data path](https://intermind.com/features/security) and zero data retention on AI processing. And because quality claims about translation age badly, the translation itself is [benchmarked publicly, per language pair, every month](https://intermind.com/benchmark) — the history is only worth keeping if the translation in it holds up. ## Run the audit on your own stack Five questions, answerable from any vendor's documentation in an afternoon: 1. **Where do translated captions go when the meeting ends?** If the docs don't say they're saved, they aren't. 2. **What language is the recording's transcript in?** "The spoken one" means your multilingual meeting has a monolingual record. 3. **Can a participant read last month's chat in their own language?** Not "can they machine-translate an export" — can they open it and read it. 4. **What does the participant who missed the meeting receive, and in whose language?** 5. **Who can verify the translation quality of whatever *does* persist?** A record nobody can trust is a record nobody uses. If the answers are "nowhere, English, no, an English summary, nobody" — you don't have a multilingual meeting tool. You have a monolingual archive with multilingual moments. --- ## Try the week-later test with a real meeting - **[Try the live demo](https://intermind.com/demo)** — no signup; hear per-participant voice translation, then look at what the room leaves behind. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, methodology included. - **[How one InterMIND meeting is built](https://intermind.com/blog/what-one-intermind-meeting-is-built-from)** — the architecture behind the space, for the technically curious. --- ## FAQ **Are Zoom translated captions saved after the meeting?** Zoom's translated-captions documentation describes a live, per-participant caption view and does not state that the translated stream is saved after the meeting (checked August 2026). What Zoom's recording features keep is the meeting as spoken — original audio and its transcript. **Are Microsoft Teams translated captions or Interpreter audio saved?** No — and here Microsoft is explicit: captions aren't saved, translated transcripts exist only during the meeting, and recordings capture original audio only, without the Interpreter agent's translated track (checked August 2026). **Does Google Meet speech translation work in recordings?** No. Google documents speech translation as a live beta — not available in recordings or live streams, capped at 90 minutes per meeting — and states that no audio is saved (checked August 2026). **Can you get a meeting transcript in another language?** From the major platforms, the transcript that persists is in the meeting's spoken language. AI notetakers persist a transcript too — also in the spoken language. Translating a transcript afterwards is an export-and-translate workflow, outside the meeting tool. InterMIND takes a different route: the persistent record — chat history, recap, documents — is rendered in each member's language inside the space itself. **What does "a space that outlives the call" mean?** It means the meeting doesn't end when the call does: it persists as a channel holding the chat, files, recaps and history, and every member reads that history in their own language. The call is an event inside the space, not the container of everything. **Why does history in your language matter if you attended the meeting live?** Because the work happens after: decisions get cited, action items get re-read, disputes get settled by the record, and new teammates catch up from it. If all of that exists only in English, everyone who thinks in another language does the *follow-through* at a disadvantage — which is where the meeting's actual value lives. --- *Sources: [Microsoft — Use live captions in Teams meetings](https://support.microsoft.com/en-us/office/use-live-captions-in-microsoft-teams-meetings-4be2d304-f675-4b57-8347-cbd000a21260){rel=""nofollow""}, [Microsoft — Interpreter in Teams meetings and calls](https://support.microsoft.com/en-us/office/interpreter-in-microsoft-teams-meetings-and-calls-c7efe2bb-535d-42ab-a5c4-d2d91619b46d){rel=""nofollow""}, [Microsoft Learn — Meeting transcription and captions](https://learn.microsoft.com/en-us/microsoftteams/meeting-transcription-captions){rel=""nofollow""}, [Google Meet — Learn about Speech Translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""}, [Google Meet — Use translated captions](https://support.google.com/meet/answer/10964115){rel=""nofollow""}, [Zoom — Enabling and configuring translated captions](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0059081){rel=""nofollow""}, [Zoom — Using the Voice translator for meetings](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0084896){rel=""nofollow""}, [Otter — supported languages](https://help.otter.ai/hc/en-us/articles/360047247414-Supported-languages){rel=""nofollow""}, [Fireflies — supported languages](https://guide.fireflies.ai/articles/2973706448-learn-about-fireflies-supported-languages){rel=""nofollow""}. Vendors change plans and language lists over time; check their pages for the current state. All facts checked August 2026.* # Voice translator: phrase, face-to-face, or live conversation "Voice translator" names three products at once. A phone app that listens and reads a translation back. Earbuds that whisper a translation of the person in front of you. And live meeting translation, where a whole conversation runs in several languages instantly, each participant hearing it in their own. They all translate speech. Only the last keeps an *instant, multi-person conversation* flowing without turning every exchange into a wait. This guide separates the three, shows how each works, and names the question that decides which you need: **a phrase, a face-to-face exchange, or a conversation?** --- ## The three kinds ### Phone-app voice translators Open an app, tap the mic, speak, and it plays back — or shows — the translation. Google Translate, Apple Translate, Microsoft Translator, dozens of travel apps. Genuinely good for one sentence at a time: a menu, a taxi, a hotel desk. The mechanic is a relay: you speak → it transcribes → it translates → it speaks back, then the other person answers into the same device and it relays the other way. It works because the exchanges are short and you pass the phone back and forth. Stretch it to a real back-and-forth and the relay is the bottleneck — you take turns operating a translator instead of talking. ### Earbud / device translators AirPods with Live Translation, Pixel Buds, dedicated translator earbuds. You wear them and hear the other person translated in your ear. Nicer than a screen — but **one-to-one, and it needs matching gear on both sides.** Built for a traveler and a local, not for a room. ### Live meeting translation A different architecture, not a better app. The *room* is multilingual: everyone speaks their own language and hears everyone else in **their** language, instantly, for the whole meeting. No phone to pass, no matching earbuds. It's the only one that survives an instant conversation among several people. --- ## Which one do you need? - **A phrase** — a menu, a direction, one exchange with a stranger. A phone app is perfect. Don't over-buy. - **A face-to-face exchange with one person, both equipped** — earbuds are the nicest experience. - **A conversation — a call, a meeting, several people, instant back-and-forth** — live meeting translation, because the others turn every exchange into a relay and every extra person into a broken assumption. The trap is using a phrase tool for a conversation: it technically "works," and it makes the conversation twice as slow, because everyone waits on the relay instead of talking. --- ## What "instant" and "live" really require Three things have to be true at once — and this is where most tools quietly fail: - **Per-listener, not per-device.** Every participant hears the room in their own picked language, simultaneously — five people, five languages, one call. The *whole room* is translated for *everyone*, each to their own language. - **Sub-second, and continuous.** If translation only arrives after the speaker pauses, it's consecutive interpretation with a synthetic voice — you feel every gap. Truly instant translation keeps pace with the talking. - **No English anchor, no regional gate, no five-language beta.** Any mix of languages, translated between participants directly — not everything routed through English, not five languages in a beta. That's not a bigger phone app. It's the architecture behind [real-time meeting translation](https://intermind.com/blog/real-time-meeting-translation) — the foundational guide to the category. --- ## Where InterMIND fits InterMIND is the third kind: a voice translator for instant, real conversations, built as live meeting translation. - **24 languages live on voice**, chat and shared notes — any mix, no English anchor, no regional gate, no five-language beta. ([The full end-to-end translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Per-listener audio, sub-second.** Each participant hears the meeting in their own picked language at the same time. (Under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Webinars and conferences included** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — drop a PDF or DOCX into the meeting and each viewer gets it in their language, 30 languages on files. - **Quality you can audit** — per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark), methodology written down. A phone app is right for a phrase. Earbuds are right for one person in front of you. When it's an instant *conversation* among several people, that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and *hear* per-listener translation instead of reading it. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, methodology included. - Start with the category guide: [*Real-time meeting translation*](https://intermind.com/blog/real-time-meeting-translation). - Comparing specific language pairs? See the voice-pair guides for [English ↔ Italian](https://intermind.com/blog/traduttore-inglese-italiano-vocale), [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale), [Romanian ↔ Italian](https://intermind.com/blog/traduttore-rumeno-italiano-vocale) and [Hindi ↔ English](https://intermind.com/blog/hindi-to-english-voice-translator) — or the [Arabic voice translator](https://intermind.com/blog/mutarjim-sawti) guide. — The Mind.com Team --- *Sources: [Apple — Live Translation with AirPods](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google — Translate with Pixel Buds](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. Vendors change features and language lists over time — check their pages for the current state.* # Spanish voice translator: one phrase, one counterpart, or a whole meeting? Spanish is where voice translation gets real. Over twenty countries speak it natively; US–Latin America business runs on the English–Spanish pair daily — sales calls, supplier reviews, nearshored engineering standups, clinics, courts. So when you search for a Spanish voice translator, the stakes are usually higher than a menu. And the search returns three different products wearing one name. A phone app that listens, then reads the translation back. Earbuds that whisper a translation of the person in front of you. And live meeting translation, where a whole call runs in Spanish and English (and Portuguese, and anything else) at once, and each participant hears it in their own language. All three "translate Spanish." Only one of them holds up when the exchange is a *conversation*. ## The three products behind one search ### 1. Phone-app voice translators Open the app, tap the mic, speak; it transcribes, translates, and plays back. Google Translate, Apple Translate, Microsoft Translator all handle English↔Spanish well — it's among the most heavily used pairs in every one of them. For a phrase, this is the right tool, and it's free. The mechanic is a relay: you speak → it processes → it answers → your counterpart replies into the same phone. Two people passing one device works for a taxi or a hotel desk. Stretch it to a supplier negotiation — three people, interruptions, someone thinking out loud — and the relay becomes the meeting. You're not conversing; you're taking turns operating a translator. ### 2. Earbud and device translators The 2025–2026 hardware wave: AirPods with Apple's Live Translation, Pixel Buds, dedicated translator earpieces. The person across the table speaks Spanish; you hear English in your ear. It feels closer to conversation because you're not staring at a screen. But note the shape: **one-to-one, one direction at a time, on your device.** For reciprocity, your counterpart needs compatible hardware and setup. It's designed for a traveler and a shopkeeper — not for the Tuesday call where Monterrey, Bogotá and Chicago are all on the line. Past two people, or past the people who happen to own the hardware, the model breaks. ### 3. Live meeting translation The third kind is a different architecture, not a bigger app. In live meeting translation the *room* is multilingual: each participant speaks their own language and hears everyone else **in theirs, simultaneously**, for the whole meeting. Nobody passes a phone; nobody pairs earbuds. Translation is a property of the call itself. It's the only one of the three built for the way Spanish is actually used at work: several people, both directions at once, nobody wanting to perform their second language in front of colleagues. The category guide — what "real time" must mean, what to ask before adopting one — is here: [real-time meeting translation](https://intermind.com/blog/real-time-meeting-translation). ## The Spanish-specific part Two things make the Spanish pair unforgiving for translation tools: - **Region.** Mexican, Rioplatense, Caribbean, Castilian — vocabulary and pacing differ enough that phrase tools tuned on one register wobble on another. A tool that can't handle *ahorita*, *vos querés* and *ordenador* in the same afternoon isn't done with Spanish. That's a quality question, and quality claims deserve receipts: we publish [per-language-pair scores on real traffic, monthly](https://intermind.com/benchmark) — including the Spanish pairs — instead of asserting accuracy in a press release. - **The pair isn't always with English.** Latin American business runs Spanish↔Portuguese constantly (Mexico–Brazil, Colombia–Brazil), and a workday can need Spanish↔German or Spanish↔Japanese with no English speaker on the call. Tools that anchor every translation on English handle exactly one shape of Spanish workday. InterMIND translates directly between any of its [24 voice languages](https://intermind.com/features/realtime-translation) — Spanish↔Portuguese included, no English in the middle. ## Which one do you actually need? One question decides it: **is this a phrase, or is this a conversation?** - **A phrase** — directions, a menu, one question to a stranger: a phone app, free, done. - **One counterpart, face to face, hardware in hand** — earbuds are the most pleasant version of the relay. - **A conversation — a call, a meeting, several people, both directions**: you need live meeting translation, because the other two turn every exchange into waiting, and every extra participant into a broken assumption. The trap is using a phrase tool for a conversation. It technically works — at half the speed and half the humanity, with everyone waiting on the relay instead of talking. ## Where InterMIND fits InterMIND is the third kind, built as meeting translation from the start: - **24 languages live in voice, chat and shared notes** — Spanish to and from any of them, each participant choosing their own listening language, simultaneously, at sub-second delay. - **Nothing to install, nothing to pay for participants** — guests [join by link, free, no account](https://intermind.com/blog/free-for-participants); the host's plan covers the room. - **Documents too** — drop a PDF or DOCX into the meeting and each reader gets it in their language (30 languages for files). - **The meeting persists** — chat history, the per-guest recap and the documents stay readable in each member's language, [a week later included](https://intermind.com/blog/the-meeting-a-week-later). A phone app is right for a phrase. Earbuds are right for one counterpart. When it's a conversation — Monterrey, Bogotá and Chicago on one call, in three accents of two or three languages — that's the job InterMIND was built for. ## Hear it instead of reading about it - **[Try the live demo](https://intermind.com/demo)** — no signup; speak English or Spanish and hear the room translated per listener. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair monthly quality on real traffic, Spanish pairs included, methodology public. - New to the category? Start with [real-time meeting translation: how it works](https://intermind.com/blog/real-time-meeting-translation). --- ## FAQ **What is the best Spanish voice translator?** Match the tool to the exchange. For phrases, the free phone apps (Google Translate, Apple Translate, Microsoft Translator) handle English↔Spanish well. For one face-to-face counterpart, translator earbuds are the most comfortable relay. For a conversation with several people — a call or meeting — you need live meeting translation, where each participant hears the room in their own language simultaneously. **Is there a Spanish voice translator that works both directions at once?** Phone apps and earbuds relay one direction at a time. Live meeting translation runs both directions continuously: the Spanish speaker hears English colleagues in Spanish while they hear her in English, with no turn-taking. That's the architecture InterMIND uses, across 24 languages. **Does Spanish voice translation work between Spanish and Portuguese, without English?** Depends on the architecture. Tools that anchor translation on English cover Spanish↔English shapes only. InterMIND translates directly between any two of its 24 voice languages — Spanish↔Portuguese included — with no English relay. **How good is machine translation for Spanish meetings?** Don't take any vendor's word for it, ours included — ask for measurements. We publish per-language-pair translation scores on real production traffic every month at [/benchmark](https://intermind.com/benchmark), with the methodology written down, so you can check the Spanish pairs yourself before relying on them. --- *Sources: [Google Translate — translate by speech](https://support.google.com/translate/answer/6142474){rel=""nofollow""}, [Apple — Live Translation with AirPods](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google Pixel Buds — real-time translation](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}. Vendors change features and language lists over time — check their pages for the current state. InterMIND numbers are verified against the production product. All facts checked August 2026.* # Online voice translator, in real time: what those words actually promise Search for a voice translator and the qualifiers do the selling: *online*, *free*, *real-time*, *instant*. Underneath the adjectives sit three different products wearing one name. A phone app that listens, then reads the translation back. Earbuds that whisper a translation of the person in front of you. And live meeting translation, where a whole call runs in several languages at once and everyone hears it in their own. All three "translate voice." Only the last one keeps a *conversation* moving without turning every exchange into a wait. This guide unpacks the qualifiers, separates the three products, and names the one question that decides which you need: **is this a phrase, or is this a conversation?** --- ## The three products behind one search ### 1. Phone-app voice translators The familiar one: open an app, tap the mic, speak, and it plays back — or displays — the translation. Google Translate, Apple Translate, Microsoft Translator and dozens of travel apps do this. They're genuinely good at what they're for: a menu, a taxi, a hotel desk, one sentence at a time. The mechanic is a relay. You speak → it transcribes → it translates → it speaks back. Then the other person answers into the same phone and it relays the other way. It works because the exchanges are short and you're both willing to pass one device back and forth. Stretch it to a real back-and-forth — three people, interruptions, someone thinking out loud — and the relay is the bottleneck. You're not having a conversation; you're taking turns operating a translator. ### 2. Earbud / device translators The 2025–2026 hardware wave: AirPods with Apple's Live Translation, Pixel Buds, dedicated translator earbuds. You wear them, the other person speaks, and you hear a translation in your ear. It feels closer to a conversation because you're not staring at a screen. But look at the shape: it's **one-to-one, one direction at a time.** It translates the person in front of you, into your ear, on your device. For it to be mutual, the other side needs the same hardware and the same setup. It's built for a traveler and a shopkeeper, not for five people on a call speaking five languages. The moment the "room" grows past two people, or the others don't own compatible gear, the model breaks. ### 3. Live meeting translation The third kind is a different architecture, not a better app. In a live meeting translator, the *room* is multilingual: every participant speaks their own language, and every participant hears everyone else **in their own**, at the same time, for the whole meeting. Nobody passes a phone. Nobody wears paired earbuds. Translation is a property of the call, not of somebody's device. It's the only one of the three that survives a real conversation — several people, all speaking, no shared hardware, no English in the middle. --- ## What "online", "free" and "real-time" actually filter for - **Online** usually means "nothing to install" — a page that translates voice right in the browser. That exists, and it's also where meeting translation lives: a browser call is "online" by definition. - **Free** is true for phrases: phone apps handle a sentence at no cost. For an ongoing *conversation*, the real cost isn't the app's price — it's the time everyone loses waiting on the relay. - **Real-time** is the most-sold and least-delivered promise. An app that waits for you to finish speaking before it starts translating isn't real-time — it's consecutive interpretation with a synthetic voice. Real real-time keeps pace with the speaker, under a second behind. If what you mean by those filters is "a conversation that flows," the filters you actually want are different ones: per listener, continuous, and no shared device — the checklist below. --- ## How to tell which one you need One question settles it: - **Is it a phrase?** — a menu, a direction, a quick word with a stranger. A phone app is perfect. Don't spend more. - **Is it face-to-face with one person, and you own the device?** — earbuds are the nicest experience. - **Is it a conversation — a call, a meeting, several people, back-and-forth?** — you need live meeting translation, because the other two turn every exchange into a relay and every extra person into a broken assumption. The trap is using a phrase tool for a conversation. It technically "works" — and makes the conversation twice as slow and half as human, because everyone is waiting on the relay instead of talking. --- ## What "live" has to mean For a voice translator to hold a conversation together, three things must be true at once — and this is where most tools quietly fail: - **Per listener, not per device.** Every participant hears the room in the language they picked, simultaneously — five people, five languages, one call. Not "I translate what they say for you"; the *whole room* is translated for *everyone*, each in their own language. - **Under a second, and continuous.** If translation only arrives after the speaker stops, that's consecutive interpretation with a synthetic voice — you feel every pause. Real live translation keeps pace with the speaker. - **No English anchoring, no regional lockouts, no five-language beta.** Any language combination, translated directly between participants — not everything routed through English, not five languages in beta. That isn't a bigger phone app. It's the architecture behind [real-time meeting translation](https://intermind.com/blog/real-time-meeting-translation) — the category's foundational guide, if you want the full picture of how it works and what to ask before you pick one. --- ## Where InterMIND fits InterMIND is the third kind: a voice translator for real conversations, built as live meeting translation rather than a phrase relay. - **24 live languages on voice**, chat and shared notes — any combination, no English anchoring, no regional lockouts, no five-language beta. ([The full end-to-end translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Per-listener audio, under a second.** Every participant hears the meeting in their chosen language at the same moment. (How it works under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Webinars and conferences included** — up to 1,500 participants, each picking the language they listen in. That's [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation), not an add-on. - **Documents too** — drop a PDF or DOCX into the meeting and every reader receives it in their own language, 30 languages on files. - **Quality you can check.** We publish per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark), with the methodology written down — not accuracy claims in a press release. A phone app is fine for a phrase. Earbuds are fine for one person in front of you. When it's a *conversation* — several people, several languages, no shared hardware — that's the job InterMIND was built for. --- ## Try the conversation side - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and *hear* per-listener translation instead of reading about it. - **[See the benchmark](https://intermind.com/benchmark)** — translation quality per pair, per month, on real traffic, methodology included. - Voice notes arriving in another language? See [how to transcribe (and understand) WhatsApp audios](https://intermind.com/blog/transcrever-audio-whatsapp). - After a dictionary with examples, not voice? See [Linguee English–Portuguese: what it's actually for](https://intermind.com/blog/linguee-tradutor-ingles-portugues). - New to the category? Start with [*Real-time meeting translation: how it works and how to evaluate it*](https://intermind.com/blog/real-time-meeting-translation). - Comparing specific language pairs? See the voice-pair guides for [English ↔ Italian](https://intermind.com/blog/traduttore-inglese-italiano-vocale), [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale), [Romanian ↔ Italian](https://intermind.com/blog/traduttore-rumeno-italiano-vocale) and [Hindi ↔ English](https://intermind.com/blog/hindi-to-english-voice-translator) — or the [Arabic voice translator](https://intermind.com/blog/mutarjim-sawti) guide. "Voice translator" is three products. Match the tool to the shape of the exchange — phrase, face-to-face, or conversation — and the choice makes itself. — The Mind.com Team --- *Sources: [Google Translate — translate by speech](https://support.google.com/translate/answer/6142474){rel=""nofollow""}, [Apple — AirPods Live Translation](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google Pixel Buds — real-time translation](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. Vendors change features and language lists over time — check their pages for the current state. InterMIND numbers verified against the shipped product.* # French–Italian voice translator: for a phrase, and for a conversation If you need to translate spoken French into Italian, or Italian into French, there are two very different situations hiding behind one search — and the right tool is different for each. **A phrase** — asking directions in Milan, reading a menu in Lyon, one exchange with a shopkeeper. Any phone voice translator does this well: speak, it plays back, done. Don't over-buy. **A conversation** — a call between a French team and an Italian team, a negotiation, an interview, anything with back-and-forth and more than two people. Here the phone-app model breaks, because it's a relay: you speak, it translates, it plays back, then the other person answers into the same device and it relays the other way. Two people passing a phone can just about manage. A real conversation can't. --- ## Why "voice translator" behaves differently in a conversation The phone-app and earbud models share one assumption: **two people, one device or one pair of earbuds, taking turns.** For French ↔ Italian that's fine when it's a traveler and a local. It falls apart the moment the exchange is a meeting: - More than two people — a couple of French speakers, a couple of Italian speakers — and there's no single device to pass. - Everyone talking, interrupting, thinking out loud — the relay adds a beat to every turn, and the conversation stops flowing. - No shared hardware — you can't hand a Zoom call a pair of earbuds. A conversation needs the translation to be a property of the *call*, not of one person's phone: every French speaker hears the Italian speakers in French, every Italian speaker hears the French speakers in Italian, at the same time, live. That's live meeting translation — a different architecture, covered in full in [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). --- ## What good French–Italian live translation looks like - **Both directions at once.** French → Italian *and* Italian → French, simultaneously, without anyone switching a mode or passing a device. - **Per-listener.** Each person hears the room in their own language, sub-second — not "the translator reads it back after you stop talking." - **Direct, not through English.** French ↔ Italian translated between the two languages, not French → English → Italian, which drops nuance twice. - **Quality you can check.** For a Romance pair like French–Italian the translation is strong; we publish per-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark) rather than asserting accuracy. --- ## Where InterMIND fits InterMIND is a voice translator built for the conversation case: French and Italian speakers on the same call, each hearing the other in their own language, live. - **French and Italian are both live voice languages** — part of [24 languages on voice](https://intermind.com/features/realtime-translation), chat and shared notes, any mix, no English anchor, no regional gate. - **Per-listener audio, sub-second.** (How it works under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Meetings, webinars and conferences** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — share a PDF or DOCX and each side reads it in their language. For a menu in Nice or a taxi in Turin, a phone app is all you need. For a French–Italian *conversation*, that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run the live voice pipeline on your own audio and *hear* French ↔ Italian translated per listener. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair quality on real traffic, methodology included. - The full picture: [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). — The Mind.com Team --- *Sources: the phone-app and earbud categories described above are documented by their vendors — [Google Translate — translate by speech](https://support.google.com/translate/answer/6142474){rel=""nofollow""}, [Apple — AirPods Live Translation](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google Pixel Buds — real-time translation](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. InterMIND numbers verified against the shipped product.* # English–Italian voice translator: for a phrase, and for a conversation If you need to translate spoken English into Italian, or Italian into English, two very different situations hide behind one search — and the right tool differs for each. **A phrase** — asking directions in London, catching an airport announcement, one exchange with a receptionist. Any phone voice translator does this well: speak, it plays back, done. Don't over-buy. **A conversation** — a call with an English-speaking client, a job interview, a meeting with an international team, an online class. This is the single most common language pair in work, and it's exactly where the phone-app model breaks, because it's a relay: you speak, it translates, it plays back, then the other person answers into the same device and it relays the other way. Two people passing a phone can just about manage. A real meeting can't. --- ## Why "voice translator" behaves differently in a conversation The phone-app and earbud models share one assumption: **two people, one device or one pair of earbuds, taking turns.** For English ↔ Italian that's fine when it's a traveler and a local. It falls apart the moment the exchange is a meeting: - More than two people — English-speaking colleagues on the call, you and an Italian colleague on the other side — and there's no single device to pass. - Everyone talking, interrupting, thinking out loud — the relay adds a beat to every turn, and the conversation stops flowing. - No shared hardware — you can't hand a Zoom or Teams call a pair of earbuds. A conversation needs the translation to be a property of the *call*, not of one person's phone: the English speakers hear you in English, you hear them in Italian, at the same time, live. That's live meeting translation — a different architecture, covered in full in [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). --- ## What good English–Italian live translation looks like - **Both directions at once.** English → Italian *and* Italian → English, simultaneously, without anyone switching a mode or passing a device. - **Per-listener.** Each person hears the room in their own language, sub-second — not "the translator reads it back after you stop talking." You speak Italian and stay focused on the content; your passive English stops being the meeting's bottleneck. - **Continuous, not turn-based.** A work call doesn't wait: the translation has to follow the rhythm of people interrupting and resuming, not impose its own. - **Quality you can check.** English ↔ Italian is one of the strongest pairs; we publish per-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark) rather than asserting accuracy. --- ## Where InterMIND fits InterMIND is a voice translator built for the conversation case: English speakers and Italian speakers on the same call, each hearing the other in their own language, live. - **English and Italian are both live voice languages** — part of [24 languages on voice](https://intermind.com/features/realtime-translation), chat and shared notes, in any mix: a three-way call can run English ↔ Italian ↔ Spanish with nobody switching anything. - **Per-listener audio, sub-second.** (Under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Meetings, webinars and conferences** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — share a PDF or DOCX and each side reads it in their own language. For a taxi in Manchester or a menu in New York, a phone app is all you need. For an English–Italian *conversation* — the client call, the interview, the team meeting — that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio and *hear* English ↔ Italian translated per listener. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair quality on real traffic, methodology included. - The full picture: [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). - Other pairs: [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale), [Romanian ↔ Italian](https://intermind.com/blog/traduttore-rumeno-italiano-vocale). — The Mind.com Team --- *Sources: the phone-app and earbud categories described above are documented by their vendors — [Google Translate — translate by speech](https://support.google.com/translate/answer/6142474){rel=""nofollow""}, [Apple — AirPods Live Translation](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google Pixel Buds — real-time translation](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. InterMIND numbers verified against the shipped product.* # Italian–Russian voice translator: for a phrase, and for a conversation If you need to translate spoken Italian into Russian, or Russian into Italian, two very different situations hide behind one search — and the right tool differs for each. **A phrase** — asking directions, reading a menu, one exchange with a shopkeeper. Any phone voice translator does this well: speak, it plays back, done. Don't over-buy. **A conversation** — a call between an Italian team and a Russian-speaking team, a negotiation, an interview, anything with back-and-forth and more than two people. Here the phone-app model breaks, because it's a relay: you speak, it translates, it plays back, then the other person answers into the same device and it relays the other way. Two people passing a phone can just about manage. A real conversation can't. --- ## Why "voice translator" behaves differently in a conversation The phone-app and earbud models share one assumption: **two people, one device or one pair of earbuds, taking turns.** For Italian ↔ Russian that's fine when it's a traveler and a local. It falls apart the moment the exchange is a meeting: - More than two people — a couple of Italian speakers, a couple of Russian speakers — and there's no single device to pass. - Everyone talking, interrupting, thinking out loud — the relay adds a beat to every turn, and the conversation stops flowing. - No shared hardware — you can't hand a video call a pair of earbuds. A conversation needs the translation to be a property of the *call*, not of one person's phone: every Italian speaker hears the Russian speakers in Italian, every Russian speaker hears the Italian speakers in Russian, at the same time, live. That's live meeting translation — a different architecture, covered in full in [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). --- ## What good Italian–Russian live translation looks like - **Both directions at once.** Italian → Russian *and* Russian → Italian, simultaneously, without anyone switching a mode or passing a device. - **Per-listener.** Each person hears the room in their own language, sub-second — not "the translator reads it back after you stop talking." - **Direct, not through English.** Italian ↔ Russian translated between the two languages, not Italian → English → Russian, which drops nuance twice — and matters more for a cross-family pair than for two Romance languages. - **Quality you can check.** We publish per-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark) rather than asserting accuracy. --- ## Where InterMIND fits InterMIND is a voice translator built for the conversation case: Italian and Russian speakers on the same call, each hearing the other in their own language, live. - **Italian and Russian are both live voice languages** — part of [24 languages on voice](https://intermind.com/features/realtime-translation), chat and shared notes, any mix, no English anchor, no regional gate. - **Per-listener audio, sub-second.** (How it works under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Meetings, webinars and conferences** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — share a PDF or DOCX and each side reads it in their language. For a menu or a taxi, a phone app is all you need. For an Italian–Russian *conversation*, that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run the live voice pipeline on your own audio and *hear* Italian ↔ Russian translated per listener. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair quality on real traffic, methodology included. - The full picture: [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). — The Mind.com Team # Romanian–Italian voice translator: for a phrase, and for a conversation If you need to translate spoken Romanian into Italian, or Italian into Romanian, two very different situations hide behind one search — and the right tool differs for each. **A phrase** — a street direction, one line of a form, a quick exchange at a counter. Any phone voice translator does this well: speak, it plays back, done. Don't over-buy. **A conversation** — a family video call between Italy and Romania, a job interview, a meeting with a supplier in Bucharest, an important appointment where the details matter. Romanian–Italian is one of the most lived language pairs in Italy — and it's exactly where the phone-app model breaks, because it's a relay: you speak, it translates, it plays back, then the other person answers into the same device and it relays the other way. Two people passing a phone can just about manage. A real conversation can't. --- ## Why "voice translator" behaves differently in a conversation The phone-app and earbud models share one assumption: **two people, one device or one pair of earbuds, taking turns.** For Romanian ↔ Italian that's fine when the exchange is short and in person. It falls apart the moment it's a real call: - More than two people — relatives in Romania on one side, family in Italy on the other — and there's no single device to pass. - Everyone talking, interrupting, thinking out loud — the relay adds a beat to every turn, and the conversation stops flowing. - No shared hardware — you can't hand a video call a pair of earbuds. A conversation needs the translation to be a property of the *call*, not of one person's phone: Romanian speakers hear the Italian side in Romanian, Italian speakers hear the Romanian side in Italian, at the same time, live. That's live meeting translation — a different architecture, covered in full in [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). --- ## What good Romanian–Italian live translation looks like - **Both directions at once.** Romanian → Italian *and* Italian → Romanian, simultaneously, without anyone switching a mode or passing a device. - **Per-listener.** Each person hears the call in their own language, sub-second — not "the translator reads it back after you stop talking." - **Direct, not through English.** Romanian ↔ Italian translated between the two languages, not Romanian → English → Italian, which drops nuance twice — and for two close Romance languages, nuance is half the meaning. - **Quality you can check.** We publish per-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark) rather than asserting accuracy — see how Romanian ↔ Italian stands before you rely on us. --- ## Where InterMIND fits InterMIND is a voice translator built for the conversation case: Romanian speakers and Italian speakers on the same call, each hearing the other in their own language, live. - **Romanian and Italian are both live voice languages** — part of [24 languages on voice](https://intermind.com/features/realtime-translation), chat and shared notes, in any mix, no English anchor. - **Per-listener audio, sub-second.** (Under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Guests join from one link, in the browser** — the person on the other side installs nothing and signs up for nothing: they open the link and speak their own language. - **Documents too** — share a PDF or DOCX and each side reads it in their own language. For one phrase at a counter, a phone app is all you need. For a Romanian–Italian *conversation* — the family video call, the interview, the negotiation — that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio and *hear* Romanian ↔ Italian translated per listener. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair quality on real traffic, methodology included. - The full picture: [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). - Other pairs: [English ↔ Italian](https://intermind.com/blog/traduttore-inglese-italiano-vocale), [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale). — The Mind.com Team # Spanish–Italian voice translator: for a phrase, and for a conversation If you need to translate spoken Spanish into Italian, or Italian into Spanish, two very different situations hide behind one search — and the right tool differs for each. **A phrase** — asking directions in Madrid, reading a menu in Rome, one exchange with a shopkeeper. Any phone voice translator does this well: speak, it plays back, done. Don't over-buy. **A conversation** — a call between a Spanish team and an Italian team, a negotiation, an interview, anything with back-and-forth and more than two people. Here the phone-app model breaks, because it's a relay: you speak, it translates, it plays back, then the other person answers into the same device and it relays the other way. Two people passing a phone can just about manage. A real conversation can't. --- ## Why "voice translator" behaves differently in a conversation The phone-app and earbud models share one assumption: **two people, one device or one pair of earbuds, taking turns.** For Spanish ↔ Italian that's fine when it's a traveler and a local. It falls apart the moment the exchange is a meeting: - More than two people — a couple of Spanish speakers, a couple of Italian speakers — and there's no single device to pass. - Everyone talking, interrupting, thinking out loud — the relay adds a beat to every turn, and the conversation stops flowing. - No shared hardware — you can't hand a video call a pair of earbuds. A conversation needs the translation to be a property of the *call*, not of one person's phone: every Spanish speaker hears the Italian speakers in Spanish, every Italian speaker hears the Spanish speakers in Italian, at the same time, live. That's live meeting translation — a different architecture, covered in full in [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). --- ## What good Spanish–Italian live translation looks like - **Both directions at once.** Spanish → Italian *and* Italian → Spanish, simultaneously, without anyone switching a mode or passing a device. - **Per-listener.** Each person hears the room in their own language, sub-second — not "the translator reads it back after you stop talking." - **Direct, not through English.** Spanish ↔ Italian translated between the two languages, not Spanish → English → Italian, which drops nuance twice. - **Quality you can check.** For a Romance pair like Spanish–Italian the translation is strong; we publish per-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark) rather than asserting accuracy. --- ## Where InterMIND fits InterMIND is a voice translator built for the conversation case: Spanish and Italian speakers on the same call, each hearing the other in their own language, live. - **Spanish and Italian are both live voice languages** — part of [24 languages on voice](https://intermind.com/features/realtime-translation), chat and shared notes, any mix, no English anchor, no regional gate. - **Per-listener audio, sub-second.** (How it works under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Meetings, webinars and conferences** — up to 1,500 participants, each on their own listening language: [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation). - **Documents too** — share a PDF or DOCX and each side reads it in their language. For a menu in Seville or a taxi in Milan, a phone app is all you need. For a Spanish–Italian *conversation*, that's the job InterMIND was built for. --- ## Try it - **[Try the live demo](https://intermind.com/demo)** — run the live voice pipeline on your own audio and *hear* Spanish ↔ Italian translated per listener. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair quality on real traffic, methodology included. - The full picture: [*Voice translator: what actually translates a live conversation*](https://intermind.com/blog/traduttore-vocale). — The Mind.com Team # Voice translator: what actually translates a live conversation Search "voice translator" and you get three products wearing one name. A phone app that listens, then reads the translation back to you. A pair of earbuds that whisper a translation of the person in front of you. And live meeting translation, where a whole call runs in several languages at once and everyone hears it in their own. They all "translate voice." Only the last one keeps a *conversation* moving without turning every exchange into a two-step relay. This guide separates the three, shows how each actually works, and names the one question that decides which you need: **is this a phrase, or is this a conversation?** --- ## The three things "voice translator" means ### 1. Phone-app voice translators The familiar one: open an app, tap the mic, speak, and it plays back — or displays — the translation. Google Translate, Apple Translate, Microsoft Translator and dozens of travel apps do this. They're genuinely good at what they're for: a menu, a taxi, a hotel desk, one sentence at a time. The mechanic is a relay. You speak → it transcribes → it translates → it speaks back. Then the other person answers into the same phone and it relays the other way. It works because the exchanges are short and you're both willing to pass one device back and forth. Stretch it to a real back-and-forth — three people, interruptions, someone thinking out loud — and the relay is the bottleneck. You're not having a conversation; you're taking turns operating a translator. ### 2. Earbud / device translators The 2025–2026 hardware wave: AirPods with Apple's Live Translation, Pixel Buds, dedicated translator earbuds. You wear them, the other person speaks, and you hear a translation in your ear. It feels closer to a conversation because you're not staring at a screen. But look at the shape: it's **one-to-one, and one-directional at a time.** It translates the person in front of you, into your ear, on your device. For it to be mutual, they need the same gear and the same setup. It's built for a traveler and a shopkeeper, not for five people on a call who each speak a different language. The moment the "room" has more than two people, or the other people aren't holding a compatible device, the model breaks. ### 3. Live meeting translation The third thing is a different architecture, not a better app. In a live meeting translator, the *room* is multilingual: each participant speaks their own language, and each participant hears everyone else in **their** language, at the same time, for as long as the meeting runs. No one passes a phone. No one wears matching earbuds. The translation is a property of the call, not of any one person's device. This is the only one of the three that survives a real conversation — several people, all talking, no shared hardware, no English in the middle. --- ## How to tell which one you actually need One question does it: - **Is it a phrase?** — a menu, a direction, one sentence to a stranger. A phone app is perfect. Don't over-buy. - **Is it a face-to-face exchange with one person, and you both have the gear?** — earbuds are the nicest experience. - **Is it a conversation — a call, a meeting, several people, back-and-forth?** — you need live meeting translation, because the other two turn every exchange into a relay and every extra person into a broken assumption. The trap is using a phrase tool for a conversation. It technically "works" — and it makes the conversation twice as slow and half as human, because everyone is waiting on the relay instead of talking. --- ## What "live" really has to mean For a voice translator to hold a conversation together, three things have to be true at once — and this is where most tools quietly fail: - **Per-listener, not per-device.** Every participant hears the room in their own picked language, simultaneously — five people, five languages, one call. Not "you translate what's said to you"; the *whole room* is translated for *everyone*, each to their own language. - **Sub-second, and continuous.** If translation only arrives after the speaker pauses, it's consecutive interpretation with a synthetic voice — you feel every gap. Real live translation keeps pace with the talking. - **No English anchor, no regional gate, no five-language beta.** Any mix of languages, translated between the participants directly — not everything routed through English, not five languages in a beta. That's not a bigger phone app. It's the architecture behind [real-time meeting translation](https://intermind.com/blog/real-time-meeting-translation) — the foundational guide to the category, if you want the full picture of how it works and what to ask before buying one. --- ## Where InterMIND sits InterMIND is the third kind: a voice translator for real conversations, built as live meeting translation rather than a phrase relay. - **24 languages live on voice**, chat and shared notes — any mix, no English anchor, no regional gate, no five-language beta. ([The full end-to-end translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Per-listener audio, sub-second.** Each participant hears the meeting in their own picked language at the same time. (How that works under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Webinars and conferences included** — up to 1,500 participants, each picking their own listening language. That's [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation), not an add-on. - **Documents too** — drop a PDF or DOCX into the meeting and each viewer gets it in their language, 30 languages on files. - **Quality you can audit.** We publish per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark), with the methodology written down — not accuracy claims in a press release. A phone app is right for a phrase. Earbuds are right for one person in front of you. When it's a *conversation* — several people, several languages, no shared hardware — that's the job InterMIND was built for. --- ## Try the conversation side of it - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and *hear* per-listener translation instead of reading it. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, methodology included. - New to the category? Start with [*Real-time meeting translation: how it works, and how to evaluate one*](https://intermind.com/blog/real-time-meeting-translation). - Comparing specific language pairs? See the voice-pair guides for [English ↔ Italian](https://intermind.com/blog/traduttore-inglese-italiano-vocale), [French ↔ Italian](https://intermind.com/blog/traduttore-francese-italiano-vocale), [Spanish ↔ Italian](https://intermind.com/blog/traduttore-spagnolo-italiano-vocale), [Italian ↔ Russian](https://intermind.com/blog/traduttore-italiano-russo-vocale), [Romanian ↔ Italian](https://intermind.com/blog/traduttore-rumeno-italiano-vocale) and [Hindi ↔ English](https://intermind.com/blog/hindi-to-english-voice-translator) — plus the [Arabic voice translator](https://intermind.com/blog/mutarjim-sawti) guide. "Voice translator" is three products. Match the tool to the shape of the exchange — phrase, face-to-face, or conversation — and the choice makes itself. — The Mind.com Team --- *Sources: [Google Translate — translate by speech](https://support.google.com/translate/answer/6142474){rel=""nofollow""}, [Apple — AirPods Live Translation](https://support.apple.com/en-us/123185){rel=""nofollow""}, [Google Pixel Buds — real-time translation](https://support.google.com/googlepixelbuds/answer/7573100){rel=""nofollow""}, checked August 2026. Vendors change features and language lists over time — check their pages for the current state. InterMIND numbers verified against the shipped product.* # How to transcribe WhatsApp voice messages — and what to do when the audio is in another language The three-minute voice note always lands at the worst time: you're in a meeting, somewhere loud, or just can't listen right now. The good news: for most cases **you don't need any extra app** — WhatsApp transcribes voice messages by itself, right on the device. The less-good news: transcription fails in predictable situations, and there's a whole case it can't solve — when the audio is in a language you don't speak. This guide covers all three scenarios, from the simplest to the most overlooked. --- ## 1. The built-in path: WhatsApp's own transcription Since late 2024 WhatsApp transcribes voice messages natively — as of August 2026 the transcript languages are English, Portuguese, Spanish and Russian. Two details matter: - **Transcription happens on your device.** The audio isn't shipped to a transcription server — processing is local, and the chat's end-to-end encryption stays intact. That's the reason to prefer the native feature over any third-party app. - **It's opt-in and per message.** You enable it once, then transcribe only the voice notes you want. The path (exact menu names vary slightly by version): 1. **Settings → Chats → Voice message transcripts** — switch it on and pick a language. The device downloads a small language pack, once. 2. In the chat, **long-press the voice note → Transcribe**. The text appears under the message itself. If the option isn't there, update the app — the feature rolled out in stages and needs a reasonably recent WhatsApp (and, on iPhone, a recent iOS). --- ## 2. Where native transcription fails — and what you can actually do On-device transcription is good, but it isn't magic. The failure modes are predictable: - **Noisy audio, background music, or a far-away microphone.** The local model is compact; it degrades faster than a server-side transcriber. There's no button that fixes it — ask for the audio again, or listen at 1.5×. - **Mixed languages in one message** ("and then he said *isso não foi o que combinamos*…") — the transcript follows the configured language and mangles the other-language stretch. - **Heavy slang, proper names, technical terms.** The text comes out with holes exactly where the important words were. - **WhatsApp Web / Desktop:** transcription is a phone feature; in the browser it may simply not be there. Transcribe on the phone. What **not** to do: forward the voice note to some unknown "WhatsApp transcriber app." Forwarding takes the message out of the encrypted chat and hands it to a third-party server with a privacy policy you haven't read. For a personal or work audio that's rarely a sensible trade — especially with the local path sitting there for free. --- ## 3. The case transcription can't solve: the audio is in another language Here's the fastest-growing scenario — and the one no "Transcribe" button fixes: the supplier in China, the customer in the US, the relative who emigrated. The voice note arrives, transcription even works… **and the text is still in a language you can't read.** You can paste the transcript into a text translator, of course. It works once. But look at what the flow became: listen → transcribe → copy → paste → translate → write a reply → translate it back → record. Every loop is friction, and working conversations run ten loops a day. The real problem stopped being "turn audio into text" — it's **holding a conversation across languages**. That problem has its own category — the [voice translator](https://intermind.com/blog/tradutor-de-voz) — and one criterion decides the choice: is this a one-off phrase, or a conversation that will continue? - **One-off phrase:** the transcribe-copy-translate loop is fine, and any translation app will do. - **Ongoing conversation:** the chain of voice notes is the symptom — you're *conversing* by messages because talking live was impossible. That's exactly the case for live meeting translation: in InterMIND, each side speaks their own language on a call and hears the other in theirs, **across 24 languages, with per-listener audio under a second behind** — nothing to transcribe, copy or paste. And the meeting ends with a recap note for whoever missed it. Want to feel the difference without scheduling anything: open [`/demo`](https://intermind.com/demo), play a stretch of audio (it can be that exact voice note) near the microphone, and *hear* the live translation instead of reading it. --- ## The practical summary | Situation | Best path | | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | | Can't listen right now | WhatsApp's native transcription (on-device, encryption intact) | | Noisy audio / slang / mixed languages | Listen at 1.5× — local transcription will fail | | Personal audio + third-party app | Avoid: the audio leaves the encrypted chat | | Audio in another language, one-off | Transcribe → translate the text | | Ongoing cross-language conversation | [Live meeting voice translator](https://intermind.com/blog/tradutor-de-voz) — everyone speaks and hears their own language | WhatsApp transcription is one of those features already sitting in your pocket that half of us never switched on. Switch it on. And when the transcript is still unreadable because the language is the problem — the problem changed its name, and so does the tool: [it's a conversation, not a phrase](https://intermind.com/blog/tradutor-de-voz). --- ## FAQ **How do I turn on WhatsApp voice message transcripts?** Settings → Chats → Voice message transcripts — switch it on and pick a language (the device downloads a small language pack once). Then in any chat, long-press a voice note → Transcribe. **Is the voice note sent to a server for transcription?** No — WhatsApp's transcription runs on your device, and the chat's end-to-end encryption stays intact. That's the reason to prefer it over any third-party "WhatsApp transcriber" app, which requires taking the audio out of the encrypted chat. **Which languages can WhatsApp transcribe?** English, Portuguese, Spanish and Russian (as of August 2026, per WhatsApp's help page). A voice note in another language — or mixing two languages — comes out mangled. **What if the voice message is in a language I don't speak?** Transcription doesn't solve that — the text is still unreadable. For a one-off, transcribe and paste into a text translator. For an ongoing conversation, that's the [voice translator](https://intermind.com/blog/tradutor-de-voz) category: a live call where each side speaks and hears their own language. — The Mind.com Team --- *Sources: [WhatsApp — How to turn voice message transcripts on or off](https://faq.whatsapp.com/241617298315321){rel=""nofollow""}, [WhatsApp blog — voice message transcripts announcement](https://blog.whatsapp.com/making-voice-messages-easier-with-transcripts){rel=""nofollow""}, checked August 2026. WhatsApp changes features and language lists over time — check their pages for the current state.* # What one InterMIND meeting is built from Almost every product is built from the same default stack — the big proprietary SaaS defaults everyone reaches for. They're the frictionless path. At every layer where your meeting data actually lives, we took a different one: our own code, or open-source we could self-host. This is the companion to [*Where one InterMIND meeting actually runs*](https://intermind.com/blog/where-one-intermind-meeting-actually-runs), which mapped the *geography* — where each service executes and what data passes through it. This post answers what a security team asks next: *what is this thing built from — and can we read it, audit it, and replace it?* Not where it runs — what it's made of, layer by layer. --- ## The defaults, and what they cost Every product is a stack of choices. For most products, most of those choices are made by default: Google Analytics, Firebase, the Google Translate API, Auth0, React. They're the frictionless path, and for most teams that's a reasonable call. The trade-off is that each one puts a part of your stack behind a vendor you can't read, can't audit, and can't leave without a rewrite. We made a different choice at every layer where your meeting data actually lives: **our own code, or open-source software we could self-host.** Where a layer doesn't touch the content of your meeting, we stay pragmatic and say so. Here's the whole picture. --- ## The spine: the engine is our code, not a third party's Start with the layer that matters most, because most of your meeting flows through it. Real-time transport and voice/chat translation both run on **`mind-sdk` + the Mind API — our own engine, on OVH France.** The default way to build a translated meeting is to bolt a real-time SaaS (LiveKit) onto a translation API (DeepL, Google); we run neither in the live path. No third-party translation model is in the loop. (We do use DeepL — but only for documents dropped into chat, not the live voice/chat path; see the runtime map. We covered the pipeline mechanics in [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) Here's the part that isn't in the runtime map: the SDK your meeting runs on is **open source under the BSD 3-Clause license** — the `mind-sdk` client is public at [gitlab.com/mindlabs/api/sdk](https://gitlab.com/mindlabs/api/sdk){rel=""nofollow""}, copyright MindMeeting OÜ, our Estonian IP entity. It speaks to the Mind API at [api.mind.com](https://api.mind.com){rel=""nofollow""}, which we run ourselves on OVH France. This isn't a "look, but don't touch" source-available arrangement. BSD 3-Clause is a permissive, OSI-approved license. Your security team can clone the SDK, read exactly how your audio and text are captured, framed, and streamed, and audit that integration against your own requirements. The server-side engine it talks to is ours — not a third party's black box — and a fully self-hostable engine for a tenant who needs it is on our roadmap, not something we offer today. We'll update this post the moment it ships. --- ## Layer by layer: the default vs. what we run | Layer | The usual default | What we run | Why it matters to you | | ------------------------------------------------- | -------------------------------------------- | ----------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | **Real-time + translation engine** (voice + chat) | LiveKit + a translation API (DeepL / Google) | `mind-sdk` (BSD-3-Clause client) + our Mind API, OVH France | The heaviest data flow is our own engine, not a third-party model — and its client SDK is open and auditable | | **Frontend framework** | React (Meta) / Next.js | Vue + Nuxt | Community-governed OSS — no single corporation owns the framework your UI rides on | | **Product analytics** | Google Analytics | PostHog | Open-source, EU cloud, proxied first-party through our own domain — usage data doesn't flow into a third-party advertising platform | | **Fonts** | Google Fonts CDN | Self-hosted (`@nuxt/fonts`) | No third-party font callout from the page your users load — a recurring GDPR finding, avoided | | **Authentication** | Auth0 / Clerk / Firebase Auth | Self-run OIDC, federated to your Google / Microsoft | No auth middleman holds your sessions — you bring your own identity provider | | **Document translation** | Google Translate | DeepL (Cologne) | Specialized EU vendor, German processing | | **Content / docs** | Contentful / Sanity (headless CMS) | Nuxt Content (git-tracked markdown) | The words on our site live in our repo, not a vendor's database | | **Application database** | Firestore / DynamoDB (proprietary) | Postgres (on Neon) | Open standard — portable to any Postgres host, no proprietary query API to rewrite | | **Object storage** | Proprietary blob APIs | Tigris (S3-compatible) | Open protocol — recordings and exports are portable to any S3 store | | **CRM / sales** | Salesforce / HubSpot | Pipedrive (Estonian) | Customer and deal records sit in an EU-domiciled CRM, not a US sales platform | Two threads run through that table. **Open-source** where the tool processes your data — so it can be audited, and in principle self-hosted. **Open standards** (Postgres, the S3 API, OIDC) where we depend on infrastructure — so nothing is locked to one vendor's pricing or compliance posture. Postgres can move to any Postgres host; storage can move to any S3 store; auth federates to the identity provider you already run. The last row sits on a third axis: the CRM holding customer records is **EU-domiciled** (Pipedrive, Estonian) rather than a US sales platform — not open-source, but not under US jurisdiction either. A couple of these deserve a sentence more. PostHog is open-source and self-hostable; we run it on PostHog's EU cloud and proxy it first-party through our own origin, so the events aren't silently dropped by ad-blockers *and* don't transit a third-party analytics domain. Authentication never goes to a third-party auth SaaS that would sit between you and your sessions — we run the OIDC flow ourselves and federate to your existing Google or Microsoft identity. And the fonts on every page are served from our own domain; the only place Google Fonts appears in our codebase is an offline brand-asset script, never the app your users load. --- ## Where we're pragmatic — said out loud We don't pretend the whole stack is hand-rolled or non-US. It isn't, and a post that claimed otherwise would be contradicted by our own runtime map. The plumbing — hosting and SSR (**Vercel**), the meeting server's compute (**Fly.io**), payments (**Stripe**), transactional email (**Resend**) — runs on US-domiciled SaaS. Stripe and Resend handle billing and invites and never see meeting content. Vercel and Fly are rented compute: our own code runs on them, and the meeting server on Fly does handle the live session and the transcript our digest reads — but that's our code on their machines, not a vendor product ingesting your meeting. All of it executes in the EU at runtime (the subject of [the runtime map](https://intermind.com/blog/where-one-intermind-meeting-actually-runs)). That's a deliberate, bounded trade-off: own and open-source the data plane; use the best available SaaS for the control plane. Naming it is the point — "sovereignty" means little if the exceptions aren't on the table next to the wins. --- ## The post-meeting AI steps, and the plan No proprietary US-domiciled model touches meeting-derived content. The language-model steps that run *after* the call — the **AI digest** (topics, decisions, action items), the **post-meeting summary**, and the **AI note-editor** — all sit on EU processors. The digest and the editor's generative actions run on **EU-hosted Mistral with zero-data-retention** (reached through Vercel's AI Gateway, pinned to the Mistral provider). The summary and the editor's *translate* action run on **our own EU engine** on OVH — the same one behind live voice and chat. Real-time voice, chat, notes, and documents never went near a general-purpose LLM in the first place. The only US model still in the loop judges our public [translation benchmark](https://intermind.com/benchmark) — scoring machine translations of fixed [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""} reference sentences, never anyone's meeting. We're still going further on the EU-Mistral steps: an **owner-controlled opt-out** to turn the digest off entirely, and a **self-hosted, open-weights summarization model on OVH** (Kimi-class) to replace the external Mistral. The point of the open-weights route isn't whose lab trained the weights — it's that open weights can run on infrastructure we control, which keeps it on the same open-and-self-hostable axis as the rest of the data plane. Both are on the roadmap, not shipped; we'll update this post when they land. --- ## Why this matters beyond us This isn't engineering for its own sake. The reason to build a stack this way shows up on your side of the contract: 1. **Auditable.** The SDK your meeting runs on is open-source code your security team can read, and the engine behind it is our own — not a third party's black box. 2. **Portable.** Open standards at every data layer — Postgres, S3, OIDC — mean no proprietary lock-in. What can be moved isn't tied to a single vendor. 3. **Self-hostable.** The open-standard data layers — Postgres, S3, OIDC — already run on infrastructure you control; a fully self-hosted translation engine is on the roadmap for the tenant who needs it. This is the picture on 2026-06-07. We'll update it when the stack changes — a vendor swap, a layer rebuilt, the digest model replaced. The current configuration is verifiable in our open `vercel.json`, our `nuxt.config.ts`, and the BSD-3-Clause `mind-sdk` repository linked above. If a layer here looks wrong, or your security review needs an answer this map doesn't give, write us. We'd rather correct a missing detail than have you find it in a code audit. --- *Sources: the [mind-sdk repository](https://gitlab.com/mindlabs/api/sdk){rel=""nofollow""} (BSD-3-Clause), [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""}; stack facts verified against the deployed configuration (`vercel.json`, `nuxt.config.ts`) and the shipped code, checked August 2026.* # Where one InterMIND meeting actually runs Every serious enterprise procurement conversation eventually reaches the same question: *"Where does this data go?"* The DPO wants a sub-processor list. The CIO wants to know which vendors are US-domiciled. Legal wants a diagram with arrows. We'd rather you have the full picture than ship it piece by piece over email. So here is the data path of one meeting — every external service it touches, where each one executes, and what data flows through it. Verified against the actual deployment configuration on 2026-05-28. Every path that touches meeting content is EU at runtime — including the post-meeting AI steps, which we used to flag as the one gap and have since moved onto EU processors. We say plainly where the one remaining US model sits, and why it never sees your meeting. *This post maps **where** your meeting runs. Its companion, [What one InterMIND meeting is built from](https://intermind.com/blog/what-one-intermind-meeting-is-built-from), maps **what it's built from** — which layers are our own code, which are open-source, and where we're pragmatic about proprietary SaaS.* --- ## What "where it runs" actually means Two things get conflated in sovereignty conversations and they are not the same thing: 1. **Runtime / data path.** Where the bytes of your meeting are physically processed during the request. This is what data-residency regulations and most DPAs are actually about. 2. **Vendor corporate domicile.** Where the SaaS vendor is legally incorporated. This is what CLOUD-Act discussions are about — the theoretical reach of a US compulsion against the vendor's parent entity, regardless of where the workload runs. Almost every "is this EU?" question is really one of these two, asked imprecisely. We answer them separately for every vendor below. --- ## The data path of one meeting Trace one call from join to follow-up email: 1. **Browser opens the meeting page.** SSR runs on Vercel, pinned to `fra1` (Frankfurt). All request/response data — session cookies, API payloads, server-rendered HTML — is processed in EU at runtime. 2. **WebSocket connects to our meeting server** in Paris (`cdg`). Meeting orchestration, presence, signalling — all EU. 3. **Speech recognition runs in the speaker's browser.** Local. Never leaves the device until the resulting transcript is sent for translation. (We covered why in [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) 4. **Voice and chat translation hit our own engine on OVH France.** This is **`mind-sdk` + the Mind API** — our code, our hosts, in France. No third-party model is in the loop. Sub-second budget, per-language WebSocket pool, EU-resident at every hop. 5. **A document dropped into chat** (PDF, DOCX, PPTX, XLSX) goes server-side from the Paris ws-server to **DeepL** in Cologne. German company, German processing. Voice and chat do not touch DeepL. 6. **Application data** — users, teams, messages, meeting metadata — lives in **Neon Postgres on AWS Frankfurt** (`eu-central-1`). Snapshots in the same region. 7. **Recordings, attachments, exports** are stored on **Tigris**, S3-compatible storage on Fly. Edge-replicated; the bucket is configurable to multi-region EU for tenants who need it pinned tighter. 8. **Errors and performance traces** go to **Sentry's EU instance** (`de.sentry.io`). The US org was retired in May. 9. **Product analytics** go to **PostHog EU** (`eu.i.posthog.com`). 10. **Transactional email** (magic links, invites, receipts) goes via **Resend** out of `eu-west-1` (Ireland). Everything above is EU at runtime. The translation engine — the part most of your data actually flows through — is also our own code, not a third party's. The client SDK it runs on is open-source (BSD-3-Clause) and auditable today; self-hosting the engine itself is on the roadmap for a customer who needs it. --- ## The vendor map | Vendor | What it does | Runtime location | | ------------------------------- | ----------------------------------------- | ------------------------------------------------------ | | **OVH** (`mind-sdk` + Mind API) | Voice + chat translation engine | France | | **Fly.io** | Meeting WebSocket orchestration | Paris (`cdg`) | | **Vercel** (Nuxt + Nitro APIs) | App shell, server APIs, SSR | Frankfurt (`fra1`) | | **Neon** | Application Postgres | AWS Frankfurt (`eu-central-1`) | | **Tigris** | Object storage (recordings, attachments) | Edge-replicated; EU-pinnable | | **DeepL** | Document translation (PDF/DOCX/PPTX/XLSX) | Cologne | | **Sentry** | Error tracking | `de.sentry.io` (EU) | | **PostHog** | Product analytics | `eu.i.posthog.com` | | **Resend** | Transactional email | Ireland (`eu-west-1`) | | **Stripe** | Payments | Ireland (Stripe Payments Europe Ltd.) for EU customers | The two heaviest data flows by volume — voice/chat translation through our own engine on OVH and document translation through DeepL — also happen to be the two vendors whose parent entity is in the EU. That covers the bulk of meeting content. The full sub-processor list with corporate-domicile detail goes into the DPA as standard practice; the table above is the runtime view, which is what most data-residency clauses are about. --- ## The post-meeting AI steps, named plainly After the call ends we run a few language-model steps on what was said: the **AI digest** (topics, decisions, action items, open questions), the **post-meeting summary**, and the **AI note-editor** (translate a note, or fix / extend / simplify it). These are the only places a general-purpose model touches meeting-derived content — and all of them now run on EU processors: - The **digest** runs on **EU-hosted Mistral** (`mistral-large-3`, `mistral-medium-3.5` fallback), reached through Vercel's AI Gateway pinned to the Mistral provider with **zero-data-retention** — the request fails rather than falling back to a non-ZDR or US host. - The **summary** and the editor's **translate** action go through **our own EU engine** on OVH — the same one that translates live voice and chat — so the summary never leaves the data plane the meeting already lived in. - The editor's **generative** actions (fix, extend, reduce, simplify, summarize) can't run on a translation engine, so they use the **same EU Mistral + zero-data-retention** path as the digest. Real-time voice, real-time chat, notes, and document translation never go through any of these — they were EU-resident from the start. The one place a US-domiciled model is still in the loop touches **no meeting data**: our public [translation-quality benchmark](https://intermind.com/benchmark) uses a frontier model as an automated judge, scoring machine translations of [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""} reference sentences. That's a fixed, public dataset — not anyone's meeting. We're still going further on the EU-Mistral steps: a planned **owner-controlled opt-out** to disable the digest entirely, and a **self-hosted open-weights model (Kimi-class) on OVH** to replace the external Mistral for summarization tasks that don't need a frontier reasoner. Both are on the roadmap, not shipped; we'll update this post when they land. --- ## What this means for your DPA For most EU buyers — German Mittelstand, regulated industries running standard GDPR DPAs — the picture above answers the data-residency question directly: every runtime hop your meeting takes is in the EU. Vendor-domicile gets disclosed in the sub-processor list per normal practice; nothing surprising there. The residency map is one piece; for the rest of the data-protection obligations — erasure, retention, portability, consent — we ran the codebase through a full audit and [closed each item against the code](https://intermind.com/blog/gdpr-audit-what-we-closed). For French *souveraineté numérique* and SecNumCloud-grade procurement, vendor corporate domicile is itself part of the criterion, not just runtime location. That's a different conversation — an alternative deployment topology that keeps every component under European-jurisdiction vendors. We don't run that by default; we'll spin it up for a tenant that needs it and where the contract justifies the build. For US-domestic and most APAC buyers, the inverse is usually true — they want low latency from their region, which is a different problem. Today we run single-region in `fra1`. If your traffic justifies a US edge, we'll plan that with you. --- ## What this post commits us to This is the picture on 2026-05-28. We'll update it when the stack changes — vendor swap, region migration, a new external service. The current configuration is verifiable in our open `vercel.json`, the `mind-sdk` + Mind API engine running at OVH France, and every vendor's own dashboard. If something here looks wrong, or your DPO needs an answer this map doesn't give, write us. We'd rather correct a missing detail than have you discover it in a contract review. --- *Sources: runtime regions and model routing verified against the deployed configuration (`vercel.json`, `fly.toml`) and the shipped code; [Vercel AI Gateway](https://vercel.com/docs/ai-gateway){rel=""nofollow""} (the digest's provider-pinning path), [FLORES-200](https://github.com/facebookresearch/flores){rel=""nofollow""}; checked August 2026.* # Why translation-quality marketing is broken — and what we publish instead Open any live-translation vendor's site. You will see the same kinds of numbers: - "200+ languages" - "6,000+ language pairs" - "World's first" / "Highest accuracy" - "99% accurate" Now try to find — on any of those vendor pages — what those numbers mean for a meeting you are about to run. Per-language quality. Reproducible methodology. Sample size. Score over time. Honest disclosure of where the model is weak. You will not find it. Not in the marketing copy, and rarely in the docs. This is the equilibrium of the category. It exists because of three things: 1. **Most vendors do not own their translation engine.** They route through OpenAI, Google, DeepL, Microsoft, or some combination. Publishing per-pair quality data would be benchmarking someone else's model — there is no marketing value in that. 2. **Honest quality data is hard to put on a billboard.** A single score is noisy. A distribution is more useful but harder to compress. A `last-six-months trend` is more useful still, and even harder. 3. **Procurement has not pushed back yet.** Buyers accept the marketing numbers at face value, and so the equilibrium holds. The equilibrium will not hold. The next class of buyer — pharma, legal, financial, audit, public sector — is going to ask harder questions than "how many languages." We built [`/benchmark`](https://intermind.com/benchmark) because we think they should not have to take a vendor's word for it. --- ## What the marketing numbers don't tell you **"200+ languages"** means a vendor has a model that emits text in 200 languages. Quality across those languages ranges from production-grade for major pairs (EN↔DE, EN↔ES, EN↔FR) to barely usable for low-resource pairs. Without a per-pair breakdown, you cannot tell which side of that line your meeting will land on. **"6,000+ language pairs"** is `N × N` combinatorics on 80 source languages. Saying you support 6,000 pairs is the easy part. Saying any specific pair is good enough for a CAPA review, a contract negotiation, or an earnings call — that is the part not in the brochure. **"99% accurate"**, without specifying what was measured, against what reference, on what sample, by what judge — is content-free. Translation quality has no universal scalar. It has a distribution that depends on language pair, content domain, audio quality (for voice), latency budget, and what "good enough" means for the specific use case. --- ## What a buyer actually needs to know The questions that show up in real DPA reviews and procurement evaluations: 1. **Per-pair quality** — how does this perform on DE↔EN, EN↔AR, JA↔KO, specifically? 2. **Sample size** — how many runs is your reported number based on? Ten? Ten thousand? 3. **Methodology** — who is judging the translations, against what reference, with what rubric? 4. **Distribution, not average** — what does the worst-case 10% look like? The best 10%? The median? 5. **Drift over time** — has a given pair gotten better or worse since you last published a number? 6. **What you don't measure** — what does your benchmark explicitly not capture? None of these are unanswerable. They are just not on anyone's marketing page. --- ## What we publish [`/benchmark`](https://intermind.com/benchmark) is our answer. The methodology is at [`/benchmark/methodology`](https://intermind.com/benchmark/methodology) — written before we knew you'd be reading this. Three things separate it from category norms. ### 1. Real traffic, not a curated suite Every score in the public benchmark comes from a real [`/demo`](https://intermind.com/demo) test run. We do not pre-select pairs that perform well. The same pipeline that serves a buyer's demo is the one being measured. ### 2. The judge is named Primary: `google/gemini-3.5-flash`. Fallback: `anthropic/claude-sonnet-5`. Both via Vercel AI Gateway. The judge is part of the methodology — disclosed by name. When we change the judge (we have, as models retire), historical rows carry the judge that actually scored them; old scores never get silently re-scored. ### 3. The distribution is the data, not the average Every published row shows median, p10, p90, min, max, and sample size — not a single number. A single number for a translation pair is noise. The shape of the distribution is the signal. --- ## Practices the category hasn't adopted - **Low-score pairs are not hidden.** The public index is gated on `≥ 10 distinct IPs, ≥ 10 runs, median ≥ 60` — but anyone can deep-link to any pair directly and see the real numbers, including the pairs that are doing badly this month. - **Known issues are documented.** When the chat-test harness was broken for a few weeks earlier in 2026, that period is suppressed from the index and noted in writing on the methodology page. History does not get silently rewritten. - **What we deliberately do NOT claim** is a full section on the methodology page. We say where the LLM judge itself is imperfect. We say what we do not measure (latency, cost, user satisfaction, ASR-side errors before translation even runs). We disclose that our own automated smoke tests are part of the traffic. --- ## A filter for the next vendor evaluation If you are evaluating any multilingual meeting platform — ours or another — the methodology is the page worth reading. The numbers themselves are the easy part. A practical filter for any vendor in this category: - **Ask for per-language-pair, per-month quality data on real traffic.** Not a curated benchmark. Not an aggregate. - **Ask what their judge is, what they explicitly do not measure, and what has changed in the last six months.** - **Ask what happens when a pair's score drops** — do they tell anyone, or do they fix it silently? If the vendor has all three answers in writing, evaluate them seriously. If they don't, you are buying marketing — not translation quality. --- ## Try it yourself - **[Try the live demo](https://intermind.com/demo)** — runs the production translation pipeline on your audio, scores it against the same judge that scores the public benchmark, and shows you the output. - **[See the benchmark](https://intermind.com/benchmark)** — every published language pair, every month, with the full distribution. - **[Read the methodology](https://intermind.com/benchmark/methodology)** — how the numbers are computed, what they include, what they do not. You will not need to take our word for any of it. That is the point. --- *Sources: the judge identifiers, gating thresholds and suppression rules above are verified against the shipped code and published at [/benchmark/methodology](https://intermind.com/benchmark/methodology); the judges are served via [Vercel AI Gateway](https://vercel.com/docs/ai-gateway){rel=""nofollow""}; the marketing claims quoted at the top are category-typical phrasings, not quotes of a named vendor; checked August 2026.* # Zoom alternatives for multilingual meetings (2026) Search "Zoom alternatives" and you get the same comparison twenty times: gallery view sizes, meeting length caps on the free tier, whiteboards, breakout rooms. Useful if Zoom's video features are your problem. Useless if your actual problem is that **your meetings happen in more than one language** — because on those lists, every alternative is the same product: a room that assumes everyone shares a language. This is the other comparison. If you're leaving Zoom (or deciding not to) because of how it handles a Spanish-speaking customer, a German counsel, or a team spread across São Paulo, Warsaw and Jakarta, three axes decide the purchase — and none of them is video quality: 1. **Live: does every participant get the meeting in their own language — voice, not just captions — and simultaneously?** 2. **Economics: who has to hold a license or add-on for that to happen?** 3. **Afterwards: what still exists in each participant's language once the call ends?** Here's Zoom's own baseline on those axes, then the alternatives people actually evaluate — Microsoft Teams, Google Meet, Cisco Webex — then us, with a table to compare all five. Every competitor fact below comes from the vendor's public documentation, checked August 2026, linked in the footer. --- ## The baseline: what you'd be leaving Zoom's translation stack in 2026 is three separate features: - **Translated captions** — text, rendered per participant, in 36 languages. Included on Business Plus and Enterprise-tier Workplace plans; other paid plans need the Translated Captions add-on. - **Voice translator** — synthetic translated audio in five languages (English, Chinese, French, Japanese, Spanish), via the Live Translation add-on or ZoomMate, desktop app only, meetings only — no webinars. - **Language Interpretation** — audio channels for human interpreters you hire, brief and pay yourself. On persistence: Zoom's documentation for translated captions and the Voice translator describes live views and does not state that translated output is saved after the meeting (checked August 2026). The recording keeps the meeting as spoken. The full walkthrough is in [Zoom live translation](https://intermind.com/blog/zoom-live-translation). If that stack fits your meetings, staying is a defensible choice. The question is what the alternatives change. ## Microsoft Teams Teams ships the most machinery of the big three: translated captions in 31 languages, plus the AI **Interpreter agent** — synthetic translated speech in 10 languages. The documented fences: - Translated captions require the **organizer to hold a Teams Premium or Microsoft 365 Copilot license**. - The Interpreter agent is **licensed per listener** (Microsoft 365 Copilot) and **metered — 20 hours per user per month**, in scheduled meetings only: no instant meetings, no webinars, no town halls. - Nothing translated survives: captions aren't saved, translated transcripts exist only live, recordings keep **original audio only**. So as a Zoom alternative for multilingual work, Teams changes the language list and the licensing model — from Zoom's add-ons to per-seat Copilot licenses with a monthly meter. Details: [Teams live translation](https://intermind.com/blog/teams-live-translation). ## Google Meet Meet's Gemini-powered **speech translation** translates spoken audio in near real time, in a voice resembling the speaker's. The documented fences: - It runs **between English and five languages** (French, German, Italian, Portuguese, Spanish) — every pair includes English. - **One language pair per meeting.** A room with English, Spanish and German speakers picks one pair; that's the meeting. - It's a **beta capped at 90 minutes**, unavailable in recordings and live streams — and no audio is saved. - Translated captions (text) cover more languages on Business Standard+ Workspace editions. As a Zoom alternative, Meet trades Zoom's five-language add-on for a five-language English-anchored pair — a different shape of the same ceiling. Details: [Google Meet live translation](https://intermind.com/blog/google-meet-live-translation). ## Cisco Webex Webex is the enterprise-hardware answer on most alternative lists. On the language axes, its documentation describes: - **Translated captions as a paid add-on** — 16 spoken languages transcribed, translated into 120+ caption languages. Text, not voice. - **Human-interpreter audio channels** for meetings and webinars — you bring the interpreters. - **Message translation** in Webex spaces, and AI Assistant meeting summaries. Voice translation by machine isn't part of the documented stack — spoken translation means interpreter channels you staff. As a Zoom alternative it's a captions-plus-interpreters model at enterprise scale (up to 1,000 attendees documented). Feature-by-feature: [InterMIND vs Webex](https://intermind.com/compare/webex). ## A note on the notetakers Otter and Fireflies appear on Zoom-alternative lists, but they answer a different question — they join your meeting and write it down. [Otter](https://intermind.com/compare/otter) transcribes in six languages, [Fireflies](https://intermind.com/compare/fireflies) claims 100+; either way the live meeting stays untranslated and the transcript persists in the one language the meeting was held in. If your problem is multilingual *meetings*, a notetaker is a supplement, not an alternative. We've compared that category separately: [Otter alternatives](https://intermind.com/blog/otter-ai-alternatives), [Fireflies alternatives](https://intermind.com/blog/fireflies-ai-alternatives). ## InterMIND We build for exactly the case the lists skip, so here is our side of the same three axes — stated the same way: - **Live:** every participant picks their language on joining and **hears every speaker in it, simultaneously** — 24 languages for voice, chat and shared notes, in the speaker's own voice. There's no English anchor and no per-meeting pair: a Portuguese–Polish–Indonesian meeting is just a meeting. Documents translate too (30 languages). The numbers differ per pipeline on purpose — [there is no single honest answer](https://intermind.com/blog/how-many-languages-do-you-support). - **Economics:** the host's plan covers the room. **Participants pay nothing and need no account** — guests join by link. No per-listener license, no add-on, no meter. ([Why participation is free](https://intermind.com/blog/free-for-participants).) - **Afterwards:** the meeting persists as a channel — chat history readable in each member's language, a recap delivered to every guest in their own language, documents in the same space. We wrote up the whole axis: [how much of your meeting still exists a week later](https://intermind.com/blog/the-meeting-a-week-later). And because "our translation is great" is exactly the claim you shouldn't take from any vendor, quality is published monthly, per language pair, on real traffic: [/benchmark](https://intermind.com/benchmark). --- ## The comparison, in one table | | Live voice translation | Languages | Per participant, simultaneously? | Who needs a license for translation | What survives the call, translated | | ------------------- | ---------------------------------------------------------------------- | ---------------------------------- | ------------------------------------------------- | ------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------- | | **Zoom** | Synthetic audio, 5 languages (add-on, desktop, meetings only) | 36 caption / 5 voice | Captions: yes. Voice: within 5 languages | Host plan tier or add-on; Voice via Live Translation add-on / ZoomMate | Not stated in documentation | | **Microsoft Teams** | Interpreter agent, 10 languages, scheduled meetings only | 31 caption / 10 voice | Captions: yes. Interpreter: per licensed listener | Organizer Premium/Copilot for captions; each listener Copilot for Interpreter (20 h/month) | Nothing — captions not saved, recordings original audio only | | **Google Meet** | Gemini speech translation, English ↔ 5 languages, one pair per meeting | 5 voice (English-anchored) | No — one pair for the whole meeting | Google AI Pro/Ultra or qualifying Workspace editions | Nothing — unavailable in recordings, no audio saved | | **Cisco Webex** | Human interpreter channels (you staff them) | 16 spoken → 120+ caption (add-on) | Captions: yes. Voice: per interpreter channel | Captions add-on; interpreters hired separately | Recording as spoken; summaries translatable | | **InterMIND** | Yes — every participant hears every speaker in their chosen language | 24 voice/chat/notes · 30 documents | Yes — each participant independently | Host's plan only; participants free, no account | Channel history, per-guest recap, documents — in each member's language | --- ## How to choose (including when not to switch) - **Your meetings are effectively English-only.** Stay where you are — a multilingual meeting product solves a problem you don't have. - **One presenter, many readers** — webinars where attendees follow along in text: Zoom's 36-language translated captions on a Business Plus plan are a real answer, in a tool you already run. - **A Microsoft shop where the multilingual few hold Copilot seats**: Teams' Interpreter agent covers 10 languages for exactly those seats, 20 hours a month, in scheduled internal meetings. - **Your meetings route through English anyway, with one other language**: Meet's speech translation covers English ↔ five languages, one pair at a time, while the beta caps suit you. - **Big staffed events with budget for interpreters**: Webex's interpreter channels at 1,000-attendee scale are built for that production model. - **The meeting itself is multilingual** — people interrupting each other in three languages, guests without your licenses, a team that needs to *reread* decisions in their own language next week: that's the case we built [InterMIND](https://intermind.com/meetings) for. ## See it instead of reading about it - **[Try the live demo](https://intermind.com/demo)** — no signup; speak, and hear the room translated per listener in any of 24 languages. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, monthly translation quality on real traffic, with the methodology public. - **[Compare feature-by-feature](https://intermind.com/compare/zoom)** — InterMIND vs Zoom, with every competitor cell sourced. --- ## FAQ **What is the best Zoom alternative for multilingual meetings?** It depends on which axis is broken for you. If captions are enough, Zoom's own 36-language translated captions or Teams' 31 are the field. If participants need to *hear* the meeting in their own language, the documented options narrow fast: Teams' Interpreter (10 languages, per-listener Copilot license, 20 h/month), Meet's speech translation (English ↔ 5, one pair per meeting), or InterMIND (24 languages, per participant, host's plan covers everyone). **Which Zoom alternatives translate voice, not just captions?** Three documented machine options in 2026: Zoom's own Voice translator (5 languages, add-on), Microsoft Teams' Interpreter agent (10 languages, per-listener license), Google Meet's speech translation (English ↔ 5, one pair). Webex handles voice via human interpreter channels. InterMIND translates voice per participant in 24 languages as the room's default behavior. **Is Microsoft Teams or Zoom better for translated meetings?** Compare the fences, not the brands. Zoom: 36 caption languages, voice in 5 via a paid add-on, desktop only. Teams: 31 caption languages behind an organizer Premium/Copilot license, voice in 10 behind a per-listener Copilot license with a 20-hour monthly meter, scheduled meetings only. Neither saves anything translated after the call — that's documented on both sides. **Are there free Zoom alternatives with live translation?** For participants, InterMIND is free by design — guests join by link with no account, and the [live demo](https://intermind.com/demo) requires no signup either; the host's plan covers the room. On the incumbent platforms, translation sits behind paid tiers, add-ons or per-seat AI licenses on the host or listener side (documented per vendor above). **Does any Zoom alternative keep the meeting translated after it ends?** Among the platforms above, the vendors' documentation answers no — Teams states captions aren't saved and recordings keep original audio; Google states speech translation is unavailable in recordings; Zoom's docs don't state translated output persists. InterMIND is built the other way around: the meeting persists as a channel whose chat, recap and documents each member reads in their own language. The full audit: [the meeting, a week later](https://intermind.com/blog/the-meeting-a-week-later). --- *Sources: [Zoom — Enabling and configuring translated captions](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0059081){rel=""nofollow""}, [Zoom — Using the Voice translator for meetings](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0084896){rel=""nofollow""}, [Zoom — Using Language Interpretation](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0064768){rel=""nofollow""}, [Microsoft — Use live captions in Teams meetings](https://support.microsoft.com/en-us/office/use-live-captions-in-microsoft-teams-meetings-4be2d304-f675-4b57-8347-cbd000a21260){rel=""nofollow""}, [Microsoft — Interpreter in Teams meetings and calls](https://support.microsoft.com/en-us/office/interpreter-in-microsoft-teams-meetings-and-calls-c7efe2bb-535d-42ab-a5c4-d2d91619b46d){rel=""nofollow""}, [Microsoft Learn — Teams Premium licensing](https://learn.microsoft.com/en-us/microsoftteams/teams-add-on-licensing/licensing-enhance-teams){rel=""nofollow""}, [Google Meet — Learn about Speech Translation](https://support.google.com/meet/answer/16221730){rel=""nofollow""}, [Google Meet — Use translated captions](https://support.google.com/meet/answer/10964115){rel=""nofollow""}, [Webex — Show real-time translation and transcription in meetings and webinars](https://help.webex.com/en-us/article/nqzpeei/Show-real-time-translation-and-transcription-in-meetings-and-webinars){rel=""nofollow""}, [Otter — supported languages](https://help.otter.ai/hc/en-us/articles/360047247414-Supported-languages){rel=""nofollow""}, [Fireflies — supported languages](https://guide.fireflies.ai/articles/2973706448-learn-about-fireflies-supported-languages){rel=""nofollow""}. Vendors change plans, prices and language lists over time; check their pages for the current state. All facts checked August 2026.* # Zoom interpreter setup: how Language Interpretation works, and when it's the wrong tool Zoom has a real, built-in interpreter feature — **Language Interpretation** — and it's the right answer for a specific meeting shape: a formal session where you've *hired professional interpreters* and want Zoom to route their audio. What it isn't is translation: Zoom provides the channels, and finding, booking, and paying the humans who speak into them is entirely your job. This is the practical guide: what the feature requires, how to set it up properly, the limits that surprise people mid-meeting, and an honest line on when you should use it versus [AI simultaneous interpretation](https://intermind.com/features/simultaneous-interpretation) instead. > For the full picture of everything Zoom can do about language — translated captions and the AI voice translator included — see [Zoom live translation: how it works, and where it stops](https://intermind.com/blog/zoom-live-translation). --- ## What Zoom's interpreter feature actually is Language Interpretation creates **separate audio channels per language** inside a meeting or webinar. Your interpreters listen to the floor audio and speak into their assigned channel; each participant picks a channel and hears the interpreter, with the original speaker ducked underneath (floor audio returns to full volume about 8 seconds after the interpreter stops). Participants can also mute the original audio entirely. It's simultaneous interpretation *delivery* — the same job an on-site booth-and-receivers rig does, minus the hardware. The interpreting itself is still done by people you bring. ## Requirements before you start - **A paid plan:** Pro, Business, Education, or Enterprise, with Language Interpretation enabled in account settings. - **A scheduled meeting with an automatically generated meeting ID.** Personal Meeting IDs don't support it, and **you can't add interpretation to an instant meeting** — it's configured at scheduling time. - **Your own interpreters**, pre-assigned by email when you schedule (up to 20 interpreters per session); the host can also assign someone manually once the meeting is running. - **Desktop or web app** for anyone interpreting or managing channels — and interpreters must join with computer audio. ## Setting it up, step by step 1. **Enable the feature:** account Settings → Meeting → In Meeting (Advanced) → Language Interpretation. Add any custom languages you need. 2. **Schedule the meeting** with a generated meeting ID and tick *Enable language interpretation*; enter each interpreter's email and language pair. 3. **Start the meeting**, click *Interpretation*, confirm (or reassign) interpreters, and click *Start*. 4. **Participants pick a channel** from the Interpretation menu and optionally mute the original audio. Mobile participants can listen to a channel, but only listen — managing and interpreting need desktop. 5. **Interpreters swap** by the usual professional cadence (pairs, \~30-minute turns) — Zoom supports interpretation relay for indirect pairs. ## The limits that bite mid-meeting - **Breakout rooms lose interpretation** — channels exist in the main session only. A multilingual workshop that splits into groups goes back to a common language the moment it splits. - **Recordings capture one channel.** A local recording keeps whatever audio the recording participant could hear; the cloud record is not a per-language multitrack. The "record" of a two-language meeting is one language plus fragments. - **Instant calls are out** — interpretation is a scheduled-meeting feature by design. - **The interpreters are your problem:** sourcing, vetting, booking days ahead, and paying per-day professional rates — two per language, per the industry standard. Zoom solves routing, not staffing. (What that staffing actually involves: [our plain-language guide to simultaneous interpretation](https://intermind.com/blog/simultaneous-interpretation-guide).) - **Captions ≠ this feature.** Zoom's translated captions are a separate, text-only capability with its own plan gating — covered in [the Zoom translation guide](https://intermind.com/blog/zoom-live-translation). None of these are bugs. They're the shape of a feature built for formal, staffed, scheduled events. ## When Zoom's interpreter channels are the right call Use Language Interpretation when the event is **worth professional humans**: a board meeting with a certified interpreter, an AGM, a press briefing, a formal training session with contracted interpreters. If you're already paying interpreters, Zoom routes them competently and your attendees never install anything new. ## When it's the wrong tool The mismatch shows up on the meetings that were never going to book interpreters — the weekly sync with the Berlin and São Paulo teams, the ad-hoc sales call, the support escalation. For those, the calculus changed: **AI simultaneous interpretation** does the interpreting itself, on demand, with no scheduling and no per-day rates. The shape of the trade, honestly: professional humans still win where a mistranslated clause is a liability — [we say so explicitly](https://intermind.com/blog/simultaneous-interpretation-guide). Everywhere below that bar, [an AI interpreter built into the meeting](https://intermind.com/features/simultaneous-interpretation) gives you what Zoom's channels can't: - **Every participant speaks and hears** — not one stage feeding an audience, but 24 languages, each listener picking their own, every speaker heard in their own voice. - **Nothing to staff or schedule** — interpretation is just *on*, for instant calls too. - **The rest of the meeting is translated** — chat, shared notes, and documents come back in each viewer's language, and the record isn't one arbitrary audio channel. - **Quality you can check** — [published per language pair, monthly](https://intermind.com/benchmark), not claimed. The direct comparison, feature by feature: [InterMIND vs Zoom](https://intermind.com/compare/zoom). --- ## FAQ **How do I get an interpreter in Zoom?** Zoom doesn't provide interpreters — enable Language Interpretation (paid plans), schedule the meeting with interpretation on, and assign interpreters you've hired yourself by email. Zoom routes their audio into per-language channels. **Does Zoom have AI interpretation?** Zoom's AI translation features are captions-first (translated captions on eligible plans, plus a voice-translator beta with narrow language coverage). Its Language Interpretation feature is human-powered by design. For built-in AI simultaneous interpretation, you're looking at [a different category of tool](https://intermind.com/blog/best-ai-conference-translation-tools). **How many languages does Zoom interpretation support?** The standard channel list covers the major conference languages, and hosts can add custom languages when enabling the feature — up to 20 interpreters can staff a session. The constraint in practice isn't the channel count; it's hiring interpreters for each pair. **Can I use Zoom interpretation in breakout rooms?** No — interpretation runs in the main session only. If your format depends on multilingual small groups, that's a structural blocker. **What does Zoom interpretation cost?** The feature is included on Pro/Business/Education/Enterprise plans; the real cost is the interpreters — professional simultaneous work is billed per interpreter, per language, per day, typically two interpreters per language. ([The full cost breakdown.](https://intermind.com/blog/simultaneous-interpretation-guide)) --- ## Hear the alternative before you book anyone If the meeting you're trying to fix is a working call rather than a staffed event, test the AI route first — it takes two minutes and costs nothing: - **[Run the live demo](https://intermind.com/demo)** — speak, and hear yourself in another language, in your own voice. - **[Read the benchmark](https://intermind.com/benchmark)** — per-pair quality on real traffic, updated monthly. - **[InterMIND vs Zoom](https://intermind.com/compare/zoom)** — where each platform actually stands, feature by feature. — The Mind.com Team --- *Sources: [Zoom support — Language Interpretation](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0064768){rel=""nofollow""}, checked July 2026.* # Zoom live translation: how it works, and where it stops If you searched "Zoom live translation," the honest short answer is: **yes, Zoom can translate a live meeting — three different ways — and which one you get depends on your plan, your region, and whether you're willing to hire the interpreters yourself.** All three are real. None of them is a room where a German, a Japanese and a Brazilian participant each *hear* the meeting in their own language for as long as it runs. This post explains how each works, how to turn it on, and exactly where the ceiling is. > This is the platform how-to companion to our foundational guide, [*Real-time meeting translation: how it works, and how to evaluate one*](https://intermind.com/blog/real-time-meeting-translation). For the Google Meet and Microsoft Teams versions of this post, see [Meet](https://intermind.com/blog/google-meet-live-translation) and [Teams](https://intermind.com/blog/teams-live-translation). New to interpretation as a category? Start with [*Simultaneous interpretation: booth, RSI, or AI*](https://intermind.com/blog/simultaneous-interpretation-guide). --- ## The three features, and why they're not the same thing Zoom has shipped live translation in three distinct forms. They get conflated constantly, so separate them first: ### 1. Translated captions (text) The workhorse. Zoom's automated captions run live transcription; translated captions translate that text stream into each participant's chosen language. Two genuinely good things here: **the language list is long** — 36 fully supported languages, plus Greek, Norwegian and Welsh as target-only — and **every participant picks their own caption language independently**, without asking the host. It works in meetings *and* webinars. The catch is the gate: translated captions come with **Business Plus and Enterprise** Workplace plans. On Pro or base Business you need the paid **Translated Captions add-on** — this is the one AI feature Zoom did *not* fold into the AI Companion bundle that ships with every paid plan. And it's a reading experience. Captions translate what anyone *reads*, not what anyone *hears*. ### 2. Voice translator for meetings (audio) The 2025–2026 headline feature, announced with AI Companion 3.0 at Zoomtopia. This one speaks: it takes the translated-caption stream and renders it as **synthetic translated audio**, per participant, with a slider to balance the original voice against the translated one. Each participant picks their own speaking and listening language. It shipped through 2025–2026 as a tightly fenced beta and left beta exactly the way Zoom's docs predicted: as a purchase. It now requires the **Live Translation add-on or ZoomMate** (where advanced AI features draw AI Credit costs) — the details are in the limits section below. ### 3. Language Interpretation (humans, channels) The oldest feature, and the one event teams know: the host designates up to 20 **human interpreters**, Zoom opens an audio channel per language, and listeners pick a channel — original audio underneath at low volume. Zoom supplies the plumbing here, not the interpreters. You book those, brief them, and pay them yourself. It's available on Pro, Business, Education and Enterprise accounts, with nine default channel languages plus custom ones. (The full setup walkthrough, limits included: [*Zoom interpreter setup*](https://intermind.com/blog/zoom-interpreter).) --- ## How to turn each one on **Translated captions:** an admin enables automated captions and translated captions in account settings (Business Plus/Enterprise or the add-on). In the meeting, each participant clicks the captions control and picks their language — no host involvement needed. **Voice translator:** an admin enables it in account settings; the account (Basic through Enterprise) needs the Live Translation add-on or ZoomMate, and participants need the desktop app 7.0.0+. In the meeting, you pick your speaking language and the language you want to hear, and adjust the original/translated audio balance. **Language Interpretation:** the host schedules a meeting or webinar with interpretation enabled, assigns interpreters in advance (or in-meeting from the desktop app), and participants pick their audio channel once it starts. The friction isn't the toggles. It's what the toggles can and can't do. --- ## The limits that actually decide if it fits As of August 2026, by Zoom's own support documentation: - **Captions translate text, not the room.** The long language list and per-participant choice are real — but everyone is reading subtitles while the meeting's audio stays monolingual. Zoom also notes plainly that "translated captions may not be accurate," and dialects don't translate between each other (no French (France) ↔ French (Canada)). - **The voice translator speaks five languages and costs extra.** The full list: **English, Chinese, French, Japanese, Spanish**. Anything outside those five — Hindi ↔ English, for instance — isn't covered at any price; for that pair specifically, see the [Hindi to English voice translator](https://intermind.com/blog/hindi-to-english-voice-translator) guide. It's out of beta and priced as the Live Translation add-on or via ZoomMate with AI Credit costs. **Meetings only — no webinars.** Desktop app only. And on long speech it stops being simultaneous: extended turns are translated *after the speaker pauses* — consecutive interpretation with a synthetic voice. - **Language Interpretation needs your humans and your planning.** Scheduled meetings only — no Personal Meeting ID, no instant meetings, no breakout rooms — and the interpreters are professionals you source, book and pay per day, per language. That's not a flaw; it's a different product category, with [a cost structure of its own](https://intermind.com/features/simultaneous-interpretation). - **No machine voice translation in webinars at all.** A multilingual webinar on Zoom today means translated captions for readers, or human interpreter channels you staff yourself. None of this makes Zoom's features bad. The caption stack in particular is broad and genuinely per-participant. But it makes Zoom's live translation **text-first, with spoken translation either a five-language paid add-on or staffed by people you hire.** --- ## The structural ceiling, in one sentence **Zoom translates what your meeting *reads* — broadly and well. What your meeting *hears* is either a five-language paid add-on, or human interpreters you bring yourself.** A genuinely multilingual meeting needs every participant *hearing* the room in their own language, for the whole meeting, wherever they are — and that's an architecture, not a longer language list. If your meetings are caption-friendly — presentations, webinars where attendees read along — Zoom's translated captions on a Business Plus plan may be all you need, in a tool you already have. If your meetings are *conversations* — people interrupting, deciding, thinking out loud in three languages — you've hit the ceiling. --- ## When you've outgrown captions and quotas This is the job we built InterMIND for: not subtitles under a monolingual call, but **every participant hearing the meeting in their own language, simultaneously, live.** Concretely, where Zoom's stack reads and its add-on speaks five languages: - **24 languages live on voice**, chat and shared notes — any mix, no English anchor, no regional gate, no five-language cap. ([The full end-to-end translation stack](https://intermind.com/features/realtime-translation) is its own page.) - **Per-listener audio, sub-second.** Each participant hears the meeting in their own picked language at the same time — five people, five languages, one room. (How that works under the hood: [*Inside the four translation pipelines*](https://intermind.com/blog/inside-the-translation-pipelines).) - **Webinars and conferences included** — up to 1,500 participants, each picking their own listening language. That's [simultaneous interpretation without the booth](https://intermind.com/features/simultaneous-interpretation), not an add-on — the full event workflow is at [events & webinars](https://intermind.com/use-case/events). - **Documents too** — drop a PDF or DOCX into the meeting and each viewer gets it in their language, 30 languages on files. - **Quality you can audit.** We publish per-language-pair scores on real traffic at [`/benchmark`](https://intermind.com/benchmark) — Zoom's accuracy claims are made in its own press releases; ours are reproducible. We're not claiming Zoom is bad — for caption-first meetings its translation stack is one of the broadest shipping. We're claiming it's a different shape of meeting. The feature-by-feature version is at [InterMIND vs. Zoom](https://intermind.com/compare/zoom). --- ## Try the other side of the ceiling - **[Try the live demo](https://intermind.com/demo)** — run our live voice pipeline on your own audio, in any of 24 languages, and *hear* per-listener translation instead of reading it. - **[See the benchmark](https://intermind.com/benchmark)** — per-pair, per-month translation quality on real traffic, with the methodology written down. - [InterMIND vs. Zoom](https://intermind.com/compare/zoom) — honest, feature-by-feature. - Just need a record, not live translation? The notetaker side of the market: [the 9 best Fireflies.ai alternatives](https://intermind.com/blog/fireflies-ai-alternatives). Zoom live translation is real, useful, and clearly fenced: captions for most, a five-language paid add-on for some, human interpreters for the rest. Knowing where the fences are is the whole decision. --- ## FAQ **Does Zoom translate audio or only captions?** Three features, three answers. Translated captions render the meeting as text in each participant's chosen language. The Voice translator speaks synthetic translated audio in five languages. Language Interpretation opens audio channels for human interpreters you hire yourself. **How many languages do Zoom translated captions support?** 36 fully supported languages, plus Greek, Norwegian and Welsh as translation targets only. Dialects don't translate between each other — no French (France) ↔ French (Canada). **What does the Zoom Voice translator require?** A Zoom account (Basic through Enterprise) with the Live Translation add-on or ZoomMate (AI Credit costs apply), and the desktop app 7.0.0 or higher. It supports five languages — English, Chinese, French, Japanese, Spanish — in meetings only, and long speech is translated after the speaker pauses. **Do translated captions cost extra on Zoom?** They're included when the host is on a Business Plus or Enterprise-tier Workplace plan; on other paid plans you need the Translated Captions add-on — it's not part of the AI Companion bundle. — The Mind.com Team --- *Sources: [Zoom — Enabling and configuring translated captions](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0059081){rel=""nofollow""}, [Zoom — Viewing captions in another language](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0060844){rel=""nofollow""}, [Zoom — Using the Voice translator for meetings](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0084896){rel=""nofollow""}, [Zoom — Enabling or disabling Voice translator for meetings](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0084897){rel=""nofollow""}, [Zoom — Using Language Interpretation in your meeting or webinar](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0064768){rel=""nofollow""}, [Zoom — Enabling Language Interpretation](https://support.zoom.com/hc/en/article?id=zm_kb&sysparm_article=KB0058373){rel=""nofollow""}, [Zoomtopia 2025 announcements](https://news.zoom.com/zoomtopia2025/){rel=""nofollow""}. Zoom expands plans and language lists over time; check the support pages for the current state. All facts checked August 2026.*