[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"blog-post-en-\u002Fspeech-to-text-on-your-own-gateway":3},{"page":4,"surround":818},{"id":5,"title":6,"authors":7,"badge":10,"body":11,"date":802,"demo-cta":803,"description":804,"extension":805,"genre":806,"heroOrder":807,"icp":808,"image":811,"meta":812,"navigation":813,"no-translate":803,"path":814,"rank-country":807,"rank-keywords":807,"rank-tag":807,"seo":815,"stem":816,"updated":807,"video":813,"__hash__":817},"blog\u002Fblog\u002Fspeech-to-text-on-your-own-gateway.md","Azure AI Speech vs Google Chirp 3 vs Amazon Transcribe on 153 voice notes: why InterMIND runs speech-to-text on your organization's own cloud gateway (2026)",[8],{"name":9},"The Mind.com Team","Architecture",{"type":12,"value":13,"toc":789},"minimark",[14,18,22,25,30,39,42,147,150,154,165,171,177,183,187,190,520,523,527,560,564,567,579,590,596,600,603,607,612,619,624,627,632,635,640,656,661,664,669,672,677,680,685,688,693,702,707,710,713,717,744,746],[15,16,6],"h1",{"id":17},"azure-ai-speech-vs-google-chirp-3-vs-amazon-transcribe-on-153-voice-notes-why-intermind-runs-speech-to-text-on-your-organizations-own-cloud-gateway-2026",[19,20,21],"p",{},"A voice note in an InterMIND channel becomes a text message that every teammate reads in their own language. The first step of that is speech-to-text, and the question we get from IT and compliance reviewers is not \"how accurate is it\" but \"whose service is it, and where does the audio go\".",[19,23,24],{},"The short answer: Azure AI Speech, the speech service of the default AI gateway, in InterMIND's own Azure tenant in Sweden Central — and the speech services of the two other gateways an organization can choose, measured on the same clips so the choice is a number, not a preference. This post shows the numbers behind that choice — the same 153 clips through Azure AI Speech, Google Cloud Speech-to-Text and Amazon Transcribe on the same day — and the two facts about each API that matter more than a percentage point of accuracy.",[26,27,29],"h2",{"id":28},"the-rule-comes-first-one-gateway-per-organization","The rule comes first: one gateway per organization",[19,31,32,33,38],{},"Every AI feature in InterMIND — meeting recaps, document summaries, the writing assistant, Ask AI and Mia in the meeting — runs on one language-model gateway the organization selects: Azure OpenAI in the EU Data Zone by default, Google Vertex AI on the EU multi-region endpoint or Amazon Bedrock in Frankfurt by choice, each in InterMIND's own tenant. We explained the reasoning in ",[34,35,37],"a",{"href":36},"\u002Fblog\u002Fmeeting-ai-on-your-own-cloud-gateway","Meeting AI on your own cloud gateway",": for most organizations that gateway is already on the approved sub-processor list, so an AI feature adds no new company to the list.",[19,40,41],{},"Speech-to-text for voice notes follows the same rule on the default gateway today, and each of the three gateways has a speech service next to its language models — which is what this post measures:",[43,44,45,64],"table",{},[46,47,48],"thead",{},[49,50,51,55,58,61],"tr",{},[52,53,54],"th",{},"Gateway",[52,56,57],{},"Speech service",[52,59,60],{},"Region used in this test",[52,62,63],{},"How the API takes a recording",[65,66,67,93,125],"tbody",{},[49,68,69,77,80,83],{},[70,71,72,76],"td",{},[73,74,75],"strong",{},"Azure OpenAI"," (Microsoft)",[70,78,79],{},"Azure AI Speech, fast transcription API",[70,81,82],{},"Sweden Central",[70,84,85,86,92],{},"One request for a file up to 5 hours and 500 MB (",[34,87,91],{"href":88,"rel":89},"https:\u002F\u002Flearn.microsoft.com\u002Fen-us\u002Fazure\u002Fai-services\u002Fspeech-service\u002Ffast-transcription-create",[90],"nofollow","fast transcription docs",", checked September 2026)",[49,94,95,101,104,111],{},[70,96,97,100],{},[73,98,99],{},"Google Vertex AI"," (Google Cloud)",[70,102,103],{},"Cloud Speech-to-Text v2, Chirp 3 model",[70,105,106,110],{},[107,108,109],"code",{},"eu"," multi-region",[70,112,113,114,119,120,92],{},"Synchronous recognition is limited to 60 seconds and 10 MB; longer audio goes through batch or streaming recognition (",[34,115,118],{"href":116,"rel":117},"https:\u002F\u002Fdocs.cloud.google.com\u002Fspeech-to-text\u002Fv2\u002Fdocs\u002Fsync-recognize",[90],"sync limits",", ",[34,121,124],{"href":122,"rel":123},"https:\u002F\u002Fdocs.cloud.google.com\u002Fspeech-to-text\u002Fv2\u002Fdocs\u002Fchirp_3-model",[90],"Chirp 3",[49,126,127,133,136,139],{},[70,128,129,132],{},[73,130,131],{},"Amazon Bedrock"," (AWS)",[70,134,135],{},"Amazon Transcribe, streaming",[70,137,138],{},"eu-central-1 (Frankfurt)",[70,140,141,142,92],{},"A streaming session over HTTP\u002F2 or WebSocket; input is PCM, FLAC or Ogg-Opus (",[34,143,146],{"href":144,"rel":145},"https:\u002F\u002Fdocs.aws.amazon.com\u002Ftranscribe\u002Flatest\u002Fdg\u002Fstreaming.html",[90],"streaming docs",[19,148,149],{},"The recogniser is told the speaker's own language in every case — the language a member set in their profile — because automatic language identification proved unreliable on short clips in our testing: a four-second Russian note came back as English. That is how the product calls Azure today, and the test calls the other two the same way.",[26,151,153],{"id":152},"what-we-measured","What we measured",[19,155,156,159,160,164],{},[73,157,158],{},"The set."," 153 clips in 17 languages: 136 dialogue turns — two business dialogues (a contract negotiation and a daily stand-up) spoken by the product's two synthetic voices in ten languages — plus 17 formal passages, the FLORES-200 business sentences of our ",[34,161,163],{"href":162},"\u002Fbenchmark\u002Fmethodology","public translation benchmark"," read by a synthetic voice, one per language and 71 to 109 seconds long. The dialogue turns are 3 to 30 seconds, the length of a real voice note.",[19,166,167,170],{},[73,168,169],{},"The metric."," Word error rate (WER) against the reference text — the fraction of words substituted, inserted or deleted — after normalizing case, punctuation and whitespace; Chinese and Japanese are compared per character. Character error rate was recorded alongside and tells the same story.",[19,172,173,176],{},[73,174,175],{},"The harness."," One script, one machine, one hour on 2026-09-16, each engine over all 153 clips. Azure and Amazon Transcribe took each clip whole. Google's synchronous recogniser takes at most 60 seconds, so the 17 formal clips were cut at silences into pieces under 55 seconds and the transcripts joined; the 136 dialogue clips went whole. Amazon Transcribe wants audio at a real-time pace, so the harness fed it four times faster than real time; its latency figure below is therefore a floor set by the feed, not the service.",[19,178,179,182],{},[73,180,181],{},"What this is not."," Synthetic voices in quiet audio, not field recordings from a phone in a car. One day, one run per clip. It is a comparison on the audio our own translation pipeline is tested on, in the languages our users write in — not a leaderboard.",[26,184,186],{"id":185},"results-word-error-rate-per-language","Results: word error rate per language",[19,188,189],{},"Mean WER on the clips of each language; every engine returned a transcript for all 153 clips.",[43,191,192,211],{},[46,193,194],{},[49,195,196,199,202,205,208],{},[52,197,198],{},"Language",[52,200,201],{},"Clips",[52,203,204],{},"Azure AI Speech",[52,206,207],{},"Google Chirp 3",[52,209,210],{},"Amazon Transcribe",[65,212,213,230,245,261,274,289,304,320,334,347,360,374,387,400,414,428,441,454,476,490,504],{},[49,214,215,218,221,224,227],{},[70,216,217],{},"Arabic",[70,219,220],{},"13",[70,222,223],{},"32%",[70,225,226],{},"46%",[70,228,229],{},"26%",[49,231,232,235,237,240,242],{},[70,233,234],{},"Chinese",[70,236,220],{},[70,238,239],{},"3%",[70,241,239],{},[70,243,244],{},"8%",[49,246,247,250,253,256,259],{},[70,248,249],{},"Czech",[70,251,252],{},"1",[70,254,255],{},"1%",[70,257,258],{},"2%",[70,260,239],{},[49,262,263,266,268,270,272],{},[70,264,265],{},"Dutch",[70,267,252],{},[70,269,258],{},[70,271,258],{},[70,273,239],{},[49,275,276,279,282,284,287],{},[70,277,278],{},"English",[70,280,281],{},"21",[70,283,258],{},[70,285,286],{},"5%",[70,288,258],{},[49,290,291,294,296,298,301],{},[70,292,293],{},"French",[70,295,220],{},[70,297,286],{},[70,299,300],{},"11%",[70,302,303],{},"12%",[49,305,306,309,311,314,317],{},[70,307,308],{},"German",[70,310,220],{},[70,312,313],{},"7%",[70,315,316],{},"9%",[70,318,319],{},"13%",[49,321,322,325,327,330,332],{},[70,323,324],{},"Hindi",[70,326,220],{},[70,328,329],{},"10%",[70,331,329],{},[70,333,303],{},[49,335,336,339,341,343,345],{},[70,337,338],{},"Hungarian",[70,340,252],{},[70,342,316],{},[70,344,313],{},[70,346,244],{},[49,348,349,352,354,356,358],{},[70,350,351],{},"Italian",[70,353,252],{},[70,355,255],{},[70,357,255],{},[70,359,239],{},[49,361,362,365,367,370,372],{},[70,363,364],{},"Japanese",[70,366,220],{},[70,368,369],{},"6%",[70,371,286],{},[70,373,286],{},[49,375,376,379,381,383,385],{},[70,377,378],{},"Korean",[70,380,252],{},[70,382,316],{},[70,384,300],{},[70,386,300],{},[49,388,389,392,394,396,398],{},[70,390,391],{},"Polish",[70,393,252],{},[70,395,255],{},[70,397,255],{},[70,399,258],{},[49,401,402,405,407,409,412],{},[70,403,404],{},"Portuguese (Brazil)",[70,406,220],{},[70,408,286],{},[70,410,411],{},"4%",[70,413,369],{},[49,415,416,419,421,424,426],{},[70,417,418],{},"Russian",[70,420,281],{},[70,422,423],{},"14%",[70,425,329],{},[70,427,423],{},[49,429,430,433,435,437,439],{},[70,431,432],{},"Spanish",[70,434,220],{},[70,436,369],{},[70,438,286],{},[70,440,369],{},[49,442,443,446,448,450,452],{},[70,444,445],{},"Turkish",[70,447,252],{},[70,449,411],{},[70,451,239],{},[70,453,411],{},[49,455,456,461,464,468,472],{},[70,457,458],{},[73,459,460],{},"All 153 clips",[70,462,463],{},"153",[70,465,466],{},[73,467,316],{},[70,469,470],{},[73,471,329],{},[70,473,474],{},[73,475,329],{},[49,477,478,481,484,486,488],{},[70,479,480],{},"Dialogue turns only",[70,482,483],{},"136",[70,485,316],{},[70,487,300],{},[70,489,329],{},[49,491,492,495,498,500,502],{},[70,493,494],{},"Formal passages only",[70,496,497],{},"17",[70,499,411],{},[70,501,286],{},[70,503,286],{},[49,505,506,509,511,514,517],{},[70,507,508],{},"Processing time ÷ audio length",[70,510],{},[70,512,513],{},"0.10",[70,515,516],{},"0.22",[70,518,519],{},"0.41 (feed-bound)",[19,521,522],{},"Languages with a single clip (the formal passage only) are shown for completeness; a one-clip figure is a data point, not a rate.",[26,524,526],{"id":525},"what-the-numbers-say","What the numbers say",[528,529,530,537,543,554],"ul",{},[531,532,533,536],"li",{},[73,534,535],{},"On the whole set the three services are within one point of each other:"," 9%, 10% and 10%. Every one of the 153 clips came back with a transcript from every engine — coverage of the 17 languages is complete on all three.",[531,538,539,542],{},[73,540,541],{},"The differences are per language, and they go both ways."," Azure had the lowest error rate on French (5% against 11% and 12%) and German (7% against 9% and 13%), and tied for lowest on English (2%, with Amazon Transcribe) and Chinese (3%, with Google). Amazon Transcribe had the lowest on Arabic (26% against 32% and 46%). Google had the lowest on Russian (10% against 14% and 14%), Portuguese, Hungarian and Turkish.",[531,544,545,548,549,553],{},[73,546,547],{},"Arabic is hard for all three."," The best result on the Arabic dialogue clips is one word in four wrong. That is consistent with what we saw on the translation side, where we withdrew Arabic from the meeting language picker in August 2026 until quality clears our bar (",[34,550,552],{"href":551},"\u002Fblog\u002Fhow-many-languages-do-you-support","how many languages we support",").",[531,555,556,559],{},[73,557,558],{},"Processing time is well under real time on all three."," A 30-second note takes about three seconds on Azure and about seven on Google in this harness; the Amazon figure is dominated by the paced feed.",[26,561,563],{"id":562},"why-azure-ai-speech-stays-the-default","Why Azure AI Speech stays the default",[19,565,566],{},"Three reasons, in the order that decided it.",[19,568,569,572,573,578],{},[73,570,571],{},"1. The default gateway is Azure."," An organization that has not chosen a gateway runs its AI features on Azure OpenAI in the EU Data Zone, so its voice notes are transcribed by Azure AI Speech in the same tenant. No additional sub-processor, no additional region. Microsoft states that Azure Speech does not store or process data outside the region of the Speech resource (",[34,574,577],{"href":575,"rel":576},"https:\u002F\u002Flearn.microsoft.com\u002Fen-us\u002Fazure\u002Fai-services\u002Fspeech-service\u002Fregions",[90],"regions page",", checked September 2026); our resource is in Sweden Central, which is listed for fast transcription.",[19,580,581,584,585,589],{},[73,582,583],{},"2. The API shape fits a voice note."," A voice note in InterMIND can be up to ten minutes long (",[34,586,588],{"href":587},"\u002Fdocs\u002Fchat\u002Fvoice-notes","voice notes","). Azure's fast transcription takes the whole recording in one request. Google's synchronous recognition stops at 60 seconds, so a ten-minute note needs batch recognition through a storage bucket or a streaming session; Amazon Transcribe streams, but takes PCM, FLAC or Ogg-Opus rather than the formats phones and browsers record in, so the audio is transcoded first. Both are solvable — the harness solves them — but they are more moving parts in the path of every note.",[19,591,592,595],{},[73,593,594],{},"3. The numbers do not argue against it."," At 9% against 10% and 10%, no engine on this set is enough better to justify moving voice notes off the default gateway. If Google or Amazon had come in at half the error rate, this post would say so.",[26,597,599],{"id":598},"what-this-means-for-organizations-on-vertex-ai-or-bedrock","What this means for organizations on Vertex AI or Bedrock",[19,601,602],{},"If your organization selected Google Vertex AI or Amazon Bedrock as its gateway, your AI features run there. Voice notes in those organizations currently arrive as a recording without a transcript: we have measured Google Cloud Speech-to-Text and Amazon Transcribe, as this post shows, but have not wired them into the product yet. The permissions and the integration code exist; this post is updated when they are in the path of every note. Until then, a voice note in a Vertex AI or Bedrock organization is a playable recording; an organization that switches its gateway to Azure gets transcripts on the notes sent from then on.",[26,604,606],{"id":605},"faq","FAQ",[19,608,609],{},[73,610,611],{},"Which speech-to-text service does InterMIND use?",[19,613,614,615,553],{},"For voice notes in chat, Azure AI Speech (fast transcription API) in InterMIND's own Azure tenant in Sweden Central — the speech service of the default AI gateway. Live speech in meetings is a different path: it runs on InterMIND's own media engine, the Mind API, hosted at OVH in France, and never touches a cloud speech service (",[34,616,618],{"href":617},"\u002Fblog\u002Fwhere-one-intermind-meeting-actually-runs","where one meeting runs",[19,620,621],{},[73,622,623],{},"Is Azure AI Speech more accurate than Google Speech-to-Text or Amazon Transcribe?",[19,625,626],{},"On our set of 153 voice-note clips in 17 languages, measured on the same day with the same harness, Azure AI Speech had a mean word error rate of 9%, Google Cloud Speech-to-Text (Chirp 3) 10% and Amazon Transcribe 10%. Per language the order changes: Azure was lowest on French, English and Chinese, Amazon Transcribe on Arabic, Google on Russian. On a different set of recordings the ranking can differ; the table above is the evidence for ours.",[19,628,629],{},[73,630,631],{},"What is word error rate, and how was it computed here?",[19,633,634],{},"Word error rate is the number of substituted, inserted and deleted words divided by the number of words in the reference text. We normalize case, punctuation and whitespace before comparing, and compare Chinese and Japanese per character because those scripts do not separate words with spaces. A rate of 9% means roughly one word in eleven differs from the reference.",[19,636,637],{},[73,638,639],{},"Does a voice note's audio leave the EU?",[19,641,642,643,645,646,650,651,655],{},"No. The recording is stored in InterMIND's object storage in the EU, and the speech service that transcribes it runs in an EU region: Sweden Central on Azure, the ",[107,644,109],{}," multi-region on Google Cloud, Frankfurt on AWS. The full list of parties and regions is on the ",[34,647,649],{"href":648},"\u002Flegal\u002Fsubprocessors","subprocessors page"," and the ",[34,652,654],{"href":653},"\u002Ftrust","trust page",".",[19,657,658],{},[73,659,660],{},"Why not recognise speech client-side, in the app, so the audio never leaves the phone?",[19,662,663],{},"We measured that too, in September 2026. Chrome's built-in recogniser with local processing recognised 0 of 17 formal clips across the product languages, and the language coverage of client-side recognisers is far narrower than the 17 languages in the table above. Client-side recognition is a direction we re-measure as the platforms improve; today the gateway path is the one that works for every member in every language.",[19,665,666],{},[73,667,668],{},"Can our organization choose which speech service transcribes its voice notes?",[19,670,671],{},"Not separately: the AI gateway is what an organization admin selects on the Integrations page. On the default gateway, voice notes are transcribed by Azure AI Speech. On Vertex AI and Amazon Bedrock, voice notes are not transcribed yet — see the section above.",[19,673,674],{},[73,675,676],{},"How long can a voice note be, and does the API limit it?",[19,678,679],{},"Up to ten minutes. Azure's fast transcription accepts files up to 5 hours and 500 MB in one request, so a note is transcribed in a single call. Google's synchronous recognition accepts 60 seconds and 10 MB, which is why the harness cut the long clips into pieces at silences.",[19,681,682],{},[73,683,684],{},"Why does the recogniser get told the speaker's language instead of detecting it?",[19,686,687],{},"Because a voice note is short. On a four-second clip, automatic language identification can return the wrong language — a Russian test note came back as English; a member's profile language is right nearly every time. The recogniser gets that language, and only when it is unknown does it get the room's languages as candidates.",[19,689,690],{},[73,691,692],{},"Is the meeting audio transcribed by the same service?",[19,694,695,696,698,699,553],{},"No. Live speech in a meeting is recognised and translated on InterMIND's own engine (the Mind API) on the media server that carries the call, hosted at OVH in France — see ",[34,697,618],{"href":617},". The cloud gateway sees text, never the audio of a call (",[34,700,701],{"href":36},"meeting AI on your own gateway",[19,703,704],{},[73,705,706],{},"Will you publish the results again?",[19,708,709],{},"The harness runs on one command, and every engine's permissions are in place, so the table can be re-measured when a vendor ships a new model. When it changes the picture, this post is updated with the date.",[711,712],"hr",{},[26,714,716],{"id":715},"see-it-for-yourself","See it for yourself",[528,718,719,727,735],{},[531,720,721,726],{},[73,722,723],{},[34,724,725],{"href":587},"Read how voice notes work"," — recording, the ten-minute limit, and where the audio goes.",[531,728,729,734],{},[73,730,731],{},[34,732,733],{"href":648},"See the subprocessors list"," — vendor, country, purpose and region for every party that processes your data.",[531,736,737,743],{},[73,738,739],{},[34,740,742],{"href":741},"\u002Fdemo","Try the live demo"," — the production translation pipeline on your own audio, no signup.",[711,745],{},[19,747,748],{},[749,750,751,752,756,757,760,761,765,766,769,770,774,775,779,780,769,784,788],"em",{},"Sources: Microsoft — Azure AI Speech fast transcription (",[34,753,755],{"href":88,"rel":754},[90],"learn.microsoft.com",") and Speech regions table (",[34,758,755],{"href":575,"rel":759},[90],"); Google Cloud — Chirp 3 model (",[34,762,764],{"href":122,"rel":763},[90],"docs.cloud.google.com","), synchronous recognition limits (",[34,767,764],{"href":116,"rel":768},[90],") and supported languages (",[34,771,764],{"href":772,"rel":773},"https:\u002F\u002Fdocs.cloud.google.com\u002Fspeech-to-text\u002Fv2\u002Fdocs\u002Fspeech-to-text-supported-languages",[90],"); AWS — Amazon Transcribe streaming (",[34,776,778],{"href":144,"rel":777},[90],"docs.aws.amazon.com","), endpoints and quotas (",[34,781,778],{"href":782,"rel":783},"https:\u002F\u002Fdocs.aws.amazon.com\u002Fgeneral\u002Flatest\u002Fgr\u002Ftranscribe.html",[90],[34,785,778],{"href":786,"rel":787},"https:\u002F\u002Fdocs.aws.amazon.com\u002Ftranscribe\u002Flatest\u002Fdg\u002Fsupported-languages.html",[90],"); all checked September 2026. Measurement: InterMIND speech bench, 2026-09-16, 153 clips, 17 languages.",{"title":790,"searchDepth":791,"depth":792,"links":793},"",2,3,[794,795,796,797,798,799,800,801],{"id":28,"depth":791,"text":29},{"id":152,"depth":791,"text":153},{"id":185,"depth":791,"text":186},{"id":525,"depth":791,"text":526},{"id":562,"depth":791,"text":563},{"id":598,"depth":791,"text":599},{"id":605,"depth":791,"text":606},{"id":715,"depth":791,"text":716},"2026-09-16",false,"We measured the speech-to-text services of the three cloud gateways an InterMIND organization can choose — Azure AI Speech, Google Cloud Speech-to-Text (Chirp 3) and Amazon Transcribe — on the same 153 voice-note clips in 17 languages, one harness, one day. Word error rates per language, the method, the limits of each API, and why the recogniser follows the gateway instead of the leaderboard.","md","editorial",null,{"role":809,"spend":810},"it-compliance","incumbent-subscription","\u002Fblog\u002Fspeech-to-text-on-your-own-gateway.svg",{},true,"\u002Fblog\u002Fspeech-to-text-on-your-own-gateway",{"title":6,"description":804},"blog\u002Fspeech-to-text-on-your-own-gateway","AnFBFj8oDSncTh4lkIQiIJwcaWMURDugJVYTcyXQE50",[819,824],{"title":820,"path":821,"stem":822,"description":823,"children":-1},"Pay only for the people who use it: InterMIND moves to per-active-member billing","\u002Fblog\u002Fpay-only-for-active-members","blog\u002Fpay-only-for-active-members","Pro and Business stop selling seats you assign by hand. Invite your whole team; each renewal counts the members who were active in the previous 28 days and bills only for them. A member who goes quiet keeps their access and costs nothing; meeting guests and outside channel participants stay free. What counts as active, how the number moves inside a billing cycle, what the letter before renewal says, and what changes on your billing page.",{"title":825,"path":826,"stem":827,"description":828,"children":-1},"Live translated captions for the whole audience: what Zoom, Teams and Google Meet require, and how each attendee reads in their own language (2026)","\u002Fblog\u002Flive-translated-captions-for-your-audience","blog\u002Flive-translated-captions-for-your-audience","One speaker, an audience in ten languages. Translated captions exist in Zoom, Microsoft Teams and Google Meet — each behind a different licence, and each answering the same question differently: who decides which languages the audience may read in. Teams lets a town-hall organizer preselect six languages (ten with Teams Premium); Zoom's pairs are set by the host's account settings; Google Meet restricts translated captions to paid Workspace editions. The documented map, next to how it works in InterMIND, where the caption language is a setting on the attendee's own side and the room is translated by voice as well."]