Translation quality, in the open

AI Translation Quality Benchmark

How well InterMIND actually translates — measured by our automated test suite over a fixed FLORES-200 benchmark set, per language pair, every month. Synthetic and reproducible, not cherry-picked averages.

Chat translation
99/100
+1 vs prev median · 21 pairs · 105 runs · Jul 2026

typed messages, translated inline

Voice translation
67/100
-1 vs prev median · 8 pairs · 8 runs · Jul 2026

one score for speech recognition + translation, end-to-end

29 language pairs·113 scored runs·22 languages·updated monthly · Jul 2026

Score = semantic fidelity to a reference translation (0–100). Higher is better.

Excellent 90+ Good 70–89 Fair 50–69 Weak <50
Quarter

Text chat translation

Typed messages translated inline as they're sent.

22 of 24 languages average Good or better (70+) this quarter · 19 improved vs Q2 2026

Average quality per language · Q3 2026 vs Q2 2026

Each bar pools every run translating to or from that language this quarter. The dashed notch marks the previous quarter.

Q3 2026 (avg) Q2 2026 median 99 · Jul 2026
100
75
50
25
0
100
100
100
100
100
100
100
100
99
99
99
99
98
97
96
94
93
92
92
91
82
70
41
0
da
▲16de
▲17es
▲16it
▲34uk
▲7ko
▲24ja
▲20pt
▲23cs
▼1sv
▲2ro
▲24zh
▲18fr
▲10nl
▲4no
▲6en
▲14pl
▼2fi
▲15hu
▲9is
▼2ru
▲70tr
▲41ar
hi

Faded bars have a thin sample (<3 runs) — read them as provisional.

Voice translation

Live speech, transcribed and translated in real time — the harder end-to-end problem.

4 of 14 languages average Good or better (70+) this quarter · 4 improved vs Q2 2026

Average quality per language · Q3 2026 vs Q2 2026

Each bar pools every run translating to or from that language this quarter. The dashed notch marks the previous quarter.

Q3 2026 (avg) Q2 2026 median 67 · Jul 2026
100
75
50
25
0
92
83
75
71
63
58
55
49
48
38
27
21
7
5
▲23es
▲12ja
▲9fr
▲20fi
▼3pl
▼12is
▼1da
▼19en
▼15uk
▼23ru
▼36sv
▼43pt
▼60ro
▼77ko

Faded bars have a thin sample (<3 runs) — read them as provisional.

Quality over time

Monthly median per modality, since we started publishing.

Chat Voice
100
75
50
25
0
62
98
73
98
68
99
67
Mar 2026
Apr 2026
May 2026
Jul 2026

These numbers come from real sessions

Every run on the live demo is scored automatically and folded into next month's benchmark. A session looks like this:

Live translated meeting
→ EN