Translation quality, in the open

AI Translation Quality Benchmark

How well InterMIND actually translates — measured by our automated test suite over a fixed FLORES-200 benchmark set, per language pair, every month. Synthetic and reproducible, not cherry-picked averages.

Chat translation
91/100
-7 vs Q2 2026median · 22 pairs · 229 runs · Q3 2026

typed messages, translated inline

Voice translation
90/100
+19 vs Q2 2026median · 22 pairs · 219 runs · Q3 2026

one score for speech recognition + translation, end-to-end

41 language pairs·82 scored runs·23 languages·updated monthly · Sep 2026

Score = semantic fidelity to a reference translation (0–100). Higher is better.

Excellent 90+ Good 70–89 Fair 50–69 Weak <50

Text chat translation

Typed messages translated inline as they're sent.

21 of 23 languages average Good or better (70+) this quarter · 15 improved vs Q2 2026

Average quality per language · Q3 2026 vs Q2 2026

Each bar pools every run translating to or from that language this quarter. The dashed notch marks the previous quarter.

Q3 2026 (avg) Q2 2026 median 91 · Q3 2026
100
75
50
25
0
99
98
96
96
96
95
93
93
93
92
92
91
90
89
88
87
86
86
85
85
79
56
53
▲19pt
▲15es
▲12it
▲12de
▲20ja
▼2ro
▲17cs
▼7sv
▼7da
▲12fr
▲5nl
▼1no
▲11pl
▲23uk
▼1en
▲10hu
▼7ko
▲11zh
▲1ru
▼9fi
▼3is
▲56tr
▲53hi

Voice translation

Live speech, transcribed and translated in real time — the harder end-to-end problem.

22 of 23 languages average Good or better (70+) this quarter · 21 improved vs Q2 2026

Average quality per language · Q3 2026 vs Q2 2026

Each bar pools every run translating to or from that language this quarter. The dashed notch marks the previous quarter.

Q3 2026 (avg) Q2 2026 median 90 · Q3 2026
100
75
50
25
0
95
95
95
94
93
92
92
91
91
91
91
90
89
88
88
88
86
84
84
84
79
78
65
▲21zh
▲26es
▲32it
▲23ja
▲26ro
▲28pt
▲26fr
▲9ko
▲16uk
▼1hi
▲19nl
▲24pl
▲26cs
▲19en
▲32da
▲31no
▲23sv
▲21de
▲23ru
▲33fi
▲14hu
▲8is
▼4tr

Quality over time

Monthly median per modality, since we started publishing.

Chat Voice
100
75
50
25
0
98
73
98
68
99
77
92
94
83
93
Mar 2026
Apr 2026
May 2026
Jul 2026
Aug 2026
Sep 2026

These numbers come from real sessions

Every run on the live demo is scored automatically and folded into next month's benchmark. A session looks like this:

Live translated meeting
→ EN