Why translation-quality marketing is broken — and what we publish instead
Open any live-translation vendor's site. You will see the same kinds of numbers:
- "200+ languages"
- "6,000+ language pairs"
- "World's first" / "Highest accuracy"
- "99% accurate"
Now try to find — on any of those vendor pages — what those numbers mean for a meeting you are about to run. Per-language quality. Reproducible methodology. Sample size. Score over time. Honest disclosure of where the model is weak.
You will not find it. Not in the marketing copy, and rarely in the docs.
This is the equilibrium of the category. It exists because of three things:
- Most vendors do not own their translation engine. They route through OpenAI, Google, DeepL, Microsoft, or some combination. Publishing per-pair quality data would be benchmarking someone else's model — there is no marketing value in that.
- Honest quality data is hard to put on a billboard. A single score is noisy. A distribution is more useful but harder to compress. A
last-six-months trendis more useful still, and even harder. - Procurement has not pushed back yet. Buyers accept the marketing numbers at face value, and so the equilibrium holds.
The equilibrium will not hold. The next class of buyer — pharma, legal, financial, audit, public sector — is going to ask harder questions than "how many languages." We built /benchmark because we think they should not have to take a vendor's word for it.
What the marketing numbers don't tell you
"200+ languages" means a vendor has a model that emits text in 200 languages. Quality across those languages ranges from production-grade for major pairs (EN↔DE, EN↔ES, EN↔FR) to barely usable for low-resource pairs. Without a per-pair breakdown, you cannot tell which side of that line your meeting will land on.
"6,000+ language pairs" is N × N combinatorics on 80 source languages. Saying you support 6,000 pairs is the easy part. Saying any specific pair is good enough for a CAPA review, a contract negotiation, or an earnings call — that is the part not in the brochure.
"99% accurate", without specifying what was measured, against what reference, on what sample, by what judge — is content-free. Translation quality has no universal scalar. It has a distribution that depends on language pair, content domain, audio quality (for voice), latency budget, and what "good enough" means for the specific use case.