Interpretation vs translation: what's the difference — and every type of interpretation explained (2026)
Interpretation is the conversion of spoken (or signed) language into another language, live, while the communication is happening. Translation is the conversion of written text from one language to another, done after the text exists. The American Translators Association compresses it into one line: translators do the writing, interpreters do the talking.
The two words get used interchangeably in everyday speech — "we need a translator for the meeting" almost always means interpreter — but they name different professions, different skills, different tools, and different buying decisions. This guide gives you the working definitions, a side-by-side comparison, every type of interpretation with the situations each one fits, and where AI now sits in the picture.
What is interpretation?
Interpretation is real-time language conversion of speech. An interpreter listens to a speaker in one language and renders the meaning — not the word-for-word text — in another language, either while the speaker is still talking (simultaneous) or in the pauses between their sentences (consecutive). The output is ephemeral: it exists in the moment, for the people in the room or on the call, and is gone when the meeting ends.
Because it happens live, interpretation allows no second draft. The interpreter carries everything — vocabulary, tone, cultural context, the speaker's intent — in working memory, at the speaker's pace. That is why the United Nations, where sessions run in six official languages rendered simultaneously into the other five, treats conference interpreting as one of the most demanding language professions, and why professional simultaneous interpreters work in pairs and swap roughly every 30 minutes.
What is translation?
Translation is the conversion of written text. A translator works on a document that already exists — a contract, a website, a subtitle file, a medical record — and produces a new written document in the target language. The work happens after the fact, with time to research terminology, consult reference material, use translation-memory tools, and revise. The output is durable: a text that can be reviewed, certified, and reused.
That review step is the practical dividing line. A translated contract can be checked word by word before anyone signs it. An interpreted negotiation cannot — which is why the two professions certify differently and price differently, and why "can I get it in writing?" is a translation question even when the meeting itself was interpreted.
Interpretation vs translation: the difference at a glance
| Interpretation | Translation | |
|---|---|---|
| Medium | Spoken or signed language | Written text |
| When it happens | Live, in real time | After the source text exists |
| Direction | Often both ways in one session | Usually one way per assignment |
| Deliverable | Ephemeral — heard, then gone | Durable — a document you can review |
| Time to work | Seconds (simultaneous) to a pause (consecutive) | Hours to weeks, with revision |
| Tools | Booth/console, receivers, or a platform | CAT tools, translation memory, glossaries |
| Fidelity target | Meaning and intent, at speaking pace | Full accuracy, reviewable word by word |
| Typical pricing | Per interpreter, per day or hour | Per word or per page |
Two consequences follow directly from the table. First, the skills barely overlap: an outstanding contract translator may be unable to interpret a meeting, and vice versa — hiring one to do the other's job is a category error, not a compromise. Second, machine assistance entered the two fields at different speeds: machine translation of text has decades of history, while machine interpretation — live speech, both directions, at conversation pace — only became practical recently. We cover that shift below.
Types of interpretation
"Types of interpretation" mixes two independent axes that are worth separating: the mode (how the interpreting itself is timed and delivered to the listener) and the delivery format (where the interpreter is and how the audio reaches you).
The modes
| Mode | How it works | Typical setting |
|---|---|---|
| Simultaneous | Interpreter renders the speech while the speaker is still talking; listeners hear it seconds behind | Conferences, multilingual meetings, broadcasts, the UN |
| Consecutive | Speaker pauses every few sentences; interpreter renders the segment | Medical consultations, depositions, site visits, interviews |
| Whispered (chuchotage) | Simultaneous interpretation whispered directly to one or two listeners, no equipment | A delegate or executive in a room running in another language |
| Liaison (bilateral) | Interpreter alternates between both languages in a back-and-forth exchange | Negotiations, escorted visits, small two-party meetings |
| Sight translation | A written document read aloud in another language on the spot | Forms and letters inside a medical or legal appointment |
Simultaneous interpretation is the mode that scales: the meeting runs at full speed, any number of languages can run in parallel (one channel per language), and listeners choose their channel. Its historical cost — soundproof booths, two interpreters per language, day rates — is what made it an event-day service rather than an everyday one. We wrote a full plain-language guide to it: Simultaneous interpretation: what it is, how it works, and when AI can do it.
Consecutive interpretation trades time for simplicity: no equipment, but the meeting takes roughly twice as long, because everything is said twice. It shines where precision per sentence matters more than pace — a diagnosis, a sworn statement, a contract clause read aloud.
Whispered and liaison interpretation are small-scale variants of the two modes above — worth knowing by name because agencies quote them separately, but nearly every real decision reduces to simultaneous vs consecutive.
The delivery formats
| Format | What it is | What it changes |
|---|---|---|
| On-site | Interpreter physically present (booth or beside the listener) | The traditional baseline: highest setup cost, hardware in the room |
| OPI — over-the-phone | Interpreter joins by phone, on demand | Minutes-based pricing, audio only, no visual context |
| VRI — video remote | Interpreter joins by video call | Adds visual context (critical for sign language); common in hospitals |
| RSI — remote simultaneous | Interpreters work from a hub or home via a platform; audio streams to each listener's device | Removes booths, receivers, and travel from simultaneous interpreting |
| AI interpretation | A speech pipeline (recognition → translation → synthesis) interprets continuously, per listener | Removes the booking and the per-day rate; interpretation becomes a software feature |
The formats are a cost story. On-site simultaneous is the most expensive configuration in language services: per interpreter, per language, per day, plus equipment. OPI and VRI made consecutive interpretation available on demand — a hospital doesn't schedule a Tagalog interpreter for a possible emergency; it dials one. RSI did the same restructuring for simultaneous: the interpreters are still professionals at day rates, but the booth, the receivers, and the travel are gone. And AI interpretation removes the remaining constraint — the booking itself — which changes which meetings get interpreted at all. The weekly sync between Berlin, São Paulo, and Tokyo was never going to book two interpreters per language pair; with AI it doesn't have to.
Which type do you need?
Match the situation, not the terminology:
- A conference, summit, or multilingual event → simultaneous. Human (on-site or RSI) for headline stakes and budget to match; AI for the sessions, breakouts, and events that a booth budget would exclude. What that looks like in practice: events and conferences.
- A medical appointment or patient conversation → consecutive, usually via VRI or OPI on demand; in several jurisdictions qualified interpreters in healthcare are a legal requirement, not a courtesy. For multilingual telehealth and staff meetings, AI interpretation covers the everyday layer: healthcare.
- A working meeting, webinar, or recurring cross-border call → this is the space human interpretation never reached (nobody books a booth for a stand-up) and where AI simultaneous interpretation is the native fit: every participant picks a language, the meeting runs at full speed.
- A court hearing, sworn deposition, or certified proceeding → an accredited human interpreter, full stop. Where the interpretation is the legal record, machine output doesn't qualify — and vendors claiming otherwise should worry you.
- A document — contract, report, website, patent → that's translation, not interpretation. Different professional, different certification, priced per word. (If the document comes up inside a live meeting, that's the one place the two jobs meet — see below.)
Where AI fits: machine interpretation
AI entered translation and interpretation asymmetrically. Machine translation of text is decades old and everywhere. Machine interpretation — live speech in, live speech out, both directions, at conversation pace — became practical only when speech recognition, machine translation, and voice synthesis each got fast enough to chain with sub-second latency.
We build one of these systems, so here is the honest shape of the category rather than a pitch:
- What AI interpretation does today: continuous simultaneous interpretation of a meeting, per listener, with no booking and no per-day rate. In InterMIND's case that is 24 languages of live voice, with each speaker's own voice preserved via zero-shot synthesis rather than one synthetic narrator (how that works) — plus the parts of a meeting a human interpreter was never asked to cover: chat messages, shared notes, and documents dropped into the call, translated inline.
- What to demand from any vendor, ours included: published per-language-pair quality, not a language count. "Supports 60 languages" says nothing about your DE↔EN. Ours is measured monthly on FLORES-200 and published in full at /benchmark — or run the live demo and judge with your own ears.
- What AI interpretation does not do: replace accredited interpreters in certified settings. Courts, treaty negotiations, and sworn proceedings belong to humans. AI's territory is the enormous space below that bar — the meetings that never got interpretation because booths and day rates priced them out.
The major meeting platforms sit at intermediate points: Zoom and Teams offer interpretation channels for human interpreters you source and pay yourself, all three offer translated captions on qualifying plans, and Google Meet ships Gemini speech translation on qualifying plans. If you're comparing tools rather than learning the category, the buyer's guide does that comparison with sources: best AI translation tools for conferences and meetings.
FAQ
What is the difference between interpretation and translation? Interpretation converts spoken (or signed) language live, while the communication happens; translation converts written text, after the text exists. An interpreter works in real time with no revision pass; a translator works on a document with time to research and revise. They are separate professions with separate certifications.
What are the main types of interpretation? By mode: simultaneous (rendered while the speaker talks), consecutive (rendered in the speaker's pauses), whispered/chuchotage (simultaneous, whispered to one or two listeners), and liaison (back-and-forth in small exchanges). By delivery: on-site, over-the-phone (OPI), video remote (VRI), remote simultaneous (RSI), and AI interpretation.
What is simultaneous interpretation? Interpretation delivered in parallel with the speaker — listeners hear their language a few seconds behind, and the meeting never pauses for translation. It's how the UN runs sessions in six languages and how multilingual conferences keep one agenda. Full guide: simultaneous interpretation explained.
Is an interpreter the same as a translator? No. An interpreter works with live speech; a translator works with written text. The skills barely overlap — one profession trains working memory and real-time delivery, the other trains research, writing, and revision. Someone who does both exists, but each role is hired, certified, and priced separately.
Which is harder, interpretation or translation? They are hard differently. Interpretation compresses everything into real time — no dictionary, no second draft, meaning rendered at the speaker's pace, which is why simultaneous interpreters work in pairs and rotate every ~30 minutes. Translation demands a different rigor: full accuracy on a text that will be reviewed word by word, sometimes certified and legally binding.
Can AI do interpretation? For everyday meetings, webinars, and events — yes: AI systems now interpret live speech continuously, per listener, in both directions. Quality varies by vendor and language pair, so verify rather than assume: check published per-pair benchmarks (ours is here) or test it live on your own voice. For certified settings — courts, sworn proceedings — accredited human interpreters remain the only qualifying option.
Do Zoom, Teams, or Google Meet include interpretation? Partially. Zoom and Teams provide interpretation channels for human interpreters you hire yourself, and all three platforms offer translated captions on qualifying plans; Meet additionally ships Gemini speech translation on qualifying plans. None of it is turnkey full-meeting interpretation — the per-platform breakdowns are here: Zoom, Teams, Meet.
See the difference, live
Definitions explain the category; hearing it settles it.
- Run the live demo — speak, and hear yourself interpreted into another language, in your own voice, on the production pipeline.
- Read the benchmark — per-language-pair quality, measured monthly, published in full.
- AI simultaneous interpretation, in detail — how live interpretation works inside an InterMIND meeting.
— The Mind.com Team
Sources: ATA — Translator vs. Interpreter, United Nations DGACM — Interpretation, Chmiel — Boothmates forever? On teamwork in a simultaneous interpreting booth (two interpreters per booth, ~30-minute turns), Zoom — Using language interpretation in meetings and webinars, Microsoft Teams — Use language interpretation in meetings, Google Meet — Speech Translation, checked August 2026. Vendors change plans and language lists over time — check their pages for the current state.