Why translation-quality marketing is broken — and what we publish instead

Every translation vendor publishes language counts. None publishes verifiable per-pair quality on real traffic. Why that gap matters in your next procurement evaluation — and what we publish instead.

The Mind.com Team

Why translation-quality marketing is broken — and what we publish instead

Why translation-quality marketing is broken — and what we publish instead

Open any live-translation vendor's site. You will see the same kinds of numbers:

  • "200+ languages"
  • "6,000+ language pairs"
  • "World's first" / "Highest accuracy"
  • "99% accurate"

Now try to find — on any of those vendor pages — what those numbers mean for a meeting you are about to run. Per-language quality. Reproducible methodology. Sample size. Score over time. Honest disclosure of where the model is weak.

You will not find it. Not in the marketing copy, and rarely in the docs.

This is the equilibrium of the category. It exists because of three things:

  1. Most vendors do not own their translation engine. They route through OpenAI, Google, DeepL, Microsoft, or some combination. Publishing per-pair quality data would be benchmarking someone else's model — there is no marketing value in that.
  2. Honest quality data is hard to put on a billboard. A single score is noisy. A distribution is more useful but harder to compress. A last-six-months trend is more useful still, and even harder.
  3. Procurement has not pushed back yet. Buyers accept the marketing numbers at face value, and so the equilibrium holds.

The equilibrium will not hold. The next class of buyer — pharma, legal, financial, audit, public sector — is going to ask harder questions than "how many languages." We built /benchmark because we think they should not have to take a vendor's word for it.


What the marketing numbers don't tell you

"200+ languages" means a vendor has a model that emits text in 200 languages. Quality across those languages ranges from production-grade for major pairs (EN↔DE, EN↔ES, EN↔FR) to barely usable for low-resource pairs. Without a per-pair breakdown, you cannot tell which side of that line your meeting will land on.

"6,000+ language pairs" is N × N combinatorics on 80 source languages. Saying you support 6,000 pairs is the easy part. Saying any specific pair is good enough for a CAPA review, a contract negotiation, or an earnings call — that is the part not in the brochure.

"99% accurate", without specifying what was measured, against what reference, on what sample, by what judge — is content-free. Translation quality has no universal scalar. It has a distribution that depends on language pair, content domain, audio quality (for voice), latency budget, and what "good enough" means for the specific use case.


What a buyer actually needs to know

The questions that show up in real DPA reviews and procurement evaluations:

  1. Per-pair quality — how does this perform on DE↔EN, EN↔AR, JA↔KO, specifically?
  2. Sample size — how many runs is your reported number based on? Ten? Ten thousand?
  3. Methodology — who is judging the translations, against what reference, with what rubric?
  4. Distribution, not average — what does the worst-case 10% look like? The best 10%? The median?
  5. Drift over time — has a given pair gotten better or worse since you last published a number?
  6. What you don't measure — what does your benchmark explicitly not capture?

None of these are unanswerable. They are just not on anyone's marketing page.


What we publish

/benchmark is our answer. The methodology is at /benchmark/methodology — written before we knew you'd be reading this.

Three things separate it from category norms.

1. Real traffic, not a curated suite

Every score in the public benchmark comes from a real /demo test run. We do not pre-select pairs that perform well. The same pipeline that serves a buyer's demo is the one being measured.

2. The judge is named

Primary: google/gemini-3.5-flash. Fallback: anthropic/claude-sonnet-5. Both via Vercel AI Gateway. The judge is part of the methodology — disclosed by name. When we change the judge (we have, as models retire), historical rows carry the judge that actually scored them; old scores never get silently re-scored.

3. The distribution is the data, not the average

Every published row shows median, p10, p90, min, max, and sample size — not a single number. A single number for a translation pair is noise. The shape of the distribution is the signal.


Practices the category hasn't adopted

  • Low-score pairs are not hidden. The public index is gated on ≥ 10 distinct IPs, ≥ 10 runs, median ≥ 60 — but anyone can deep-link to any pair directly and see the real numbers, including the pairs that are doing badly this month.
  • Known issues are documented. When the chat-test harness was broken for a few weeks earlier in 2026, that period is suppressed from the index and noted in writing on the methodology page. History does not get silently rewritten.
  • What we deliberately do NOT claim is a full section on the methodology page. We say where the LLM judge itself is imperfect. We say what we do not measure (latency, cost, user satisfaction, ASR-side errors before translation even runs). We disclose that our own automated smoke tests are part of the traffic.

A filter for the next vendor evaluation

If you are evaluating any multilingual meeting platform — ours or another — the methodology is the page worth reading. The numbers themselves are the easy part.

A practical filter for any vendor in this category:

  • Ask for per-language-pair, per-month quality data on real traffic. Not a curated benchmark. Not an aggregate.
  • Ask what their judge is, what they explicitly do not measure, and what has changed in the last six months.
  • Ask what happens when a pair's score drops — do they tell anyone, or do they fix it silently?

If the vendor has all three answers in writing, evaluate them seriously. If they don't, you are buying marketing — not translation quality.


Try it yourself

  • Try the live demo — runs the production translation pipeline on your audio, scores it against the same judge that scores the public benchmark, and shows you the output.
  • See the benchmark — every published language pair, every month, with the full distribution.
  • Read the methodology — how the numbers are computed, what they include, what they do not.

You will not need to take our word for any of it. That is the point.


Sources: the judge identifiers, gating thresholds and suppression rules above are verified against the shipped code and published at /benchmark/methodology; the judges are served via Vercel AI Gateway; the marketing claims quoted at the top are category-typical phrasings, not quotes of a named vendor; checked August 2026.

More in Live translation

All posts in Live translation
Turkish voice translator: the verb arrives last, and that decides which one you need
Live translation

Turkish voice translator: the verb arrives last, and that decides which one you need

Turkish puts the verb — and the negation, and the tense — at the end of the sentence. That single fact separates the three products sold as a 'voice translator': phone apps, translator earbuds, and live meeting translation. What each can and cannot do with a Turkish sentence, and why a vendor's language count tells you nothing about this pair.

The Mind.com Team

Simultaneous interpreter: human, RSI platform, or AI — what your multilingual meeting needs (2026)
Live translation

Simultaneous interpreter: human, RSI platform, or AI — what your multilingual meeting needs (2026)

"Simultaneous interpreter" is a profession; what most searches actually want is speech arriving in another language while it's being spoken. This guide separates the booth, the RSI platform, and AI simultaneous translation, compares the tools on their documentation — Interprefy, KUDO, Wordly, DeepL Voice, Zoom, Teams, Google Meet, InterMIND — and asks what the comparisons skip: how much of the meeting actually comes back in your language, and where the data runs.

The Mind.com Team

Simultaneous translation: booth, RSI, or AI — and which tools for your meetings (2026)
Live translation

Simultaneous translation: booth, RSI, or AI — and which tools for your meetings (2026)

"Simultaneous translation" covers three realities: the interpreter in a booth, remote simultaneous interpretation (RSI), and real-time AI translation. This guide separates the three, compares the documented tools — Interprefy, KUDO, Wordly, DeepL Voice, Zoom, Teams, Google Meet, InterMIND — and asks the question comparison posts skip: how much of the meeting actually comes back in your language?

The Mind.com Team

Get new posts and product updates by email

One email a month with new posts and product updates. Unsubscribe anytime.