Choosing Languages and Voices
An agent has two speech settings that people mix up: the language it listens in (speech-to-text, STT) and the voice it speaks with (text-to-speech, TTS). They are set separately, and neither list is fixed by Wetel.
The source of truth is Google’s, not this page
Section titled “The source of truth is Google’s, not this page”Wetel does not keep its own list of supported languages or voices. The speech engines are Google’s, and Google adds and retires languages and voices on its own schedule, so any list copied here would go stale.
- Speech-to-text languages: Google’s list.
- Text-to-speech voices: Google’s voice list.
- Both links are also returned by the
agentConfigOptionsquery assttLanguageCatalogUrlandttsVoiceCatalogUrl.
If this site and Google’s lists disagree, trust Google’s, and please tell your Wetel contact so we can fix the page.
Which setting do I need?
Section titled “Which setting do I need?”| Your situation | Set this |
|---|---|
| Users speak one language | sttLanguageCode to that language, sttAlternativeLanguageCodes to [] |
| Users mix two languages (for example English and Malay) | Leave the defaults: the primary follows ttsLang, alternatives are en-US and ms-MY |
| Users speak Chinese | Make a Chinese code the primary (cmn-Hans-CN), and list up to two others as alternatives |
| You are not sure what users speak | Use the Gemini speech engine, which detects the language itself (see below) |
| You want a specific persona’s sound | Pick a specific voice, and keep ttsLang and voiceTier consistent with it |
| You only want a natural-sounding default | Pick a voice tier and a language, and leave the voice name alone |
Speech-to-text: primary and alternatives
Section titled “Speech-to-text: primary and alternatives”- Primary (
sttLanguageCode): the language the recognizer expects. It has by far the biggest effect on accuracy. If you leave it empty, the agent’sttsLangis used. - Alternatives (
sttAlternativeLanguageCodes): extra languages the recognizer may switch to. At most 3, no duplicates, and the primary must not be repeated. Google silently ignores the whole list when it holds more than three, so Wetel rejects it up front. - Default alternatives: when you set nothing, Wetel uses
en-USandms-MY. An explicit empty list ([]) means “no alternatives”, which is different from not setting it. - Never put a secondary language as the primary. A bilingual agent should keep its main language as the primary and add the other as an alternative.
- Chinese is not in the defaults on purpose. Google tended to label accented English as Mandarin, which produced empty or garbled transcripts. If you need Chinese, set it as the primary, or list
cmn-Hans-CNoryue-Hant-HKexplicitly. - Not supported at any setting: Hokkien and Hakka.
- Mixing languages in one sentence is not handled well: the recognizer picks one dominant language and usually loses the other’s words.
The two speech engines behave differently. Cloud Speech-to-Text uses the primary and alternatives above. The Gemini engine works out the language itself and ignores both settings. See the /stt reference for which engine your environment uses by default and how to choose one per request.
Text-to-speech: tier, language and voice
Section titled “Text-to-speech: tier, language and voice”A Google voice name has the form <language>-<tier>-<name>, for example en-US-Chirp3-HD-Kore. The name already contains the language and the tier, so three settings have to agree:
ttsLang: the language, such asen-USorms-MY.voiceTier:STANDARD,WAVENET,NEURAL2,CHIRP3_HDorSTUDIO.ttsVoice: the full voice name.
If they disagree (for example a Neural2 voice name with voiceTier set to STANDARD), the request may be rejected or the voice may be ignored, so don’t rely on either outcome. Set all three from the same row of Google’s voice list.
| Tier | Use it for | Cost |
|---|---|---|
STANDARD | Cheapest, clearly synthetic | Lowest |
WAVENET | Better quality, still budget | Low |
NEURAL2 | Natural, a safe default | Medium |
CHIRP3_HD | The most natural for live conversation | About 2x NEURAL2 |
STUDIO | Narration and recorded content. Not for real-time chat | About 10x NEURAL2 |
Use agentConfigOptions for a short, curated voice list, but remember it is a starting point and not the full catalog.
Worked examples
Section titled “Worked examples”One language, no guessing:
mutation { updateAgent( id: 12 input: { ttsLang: "ms-MY" sttLanguageCode: "ms-MY" sttAlternativeLanguageCodes: [] } ) { id }}Mandarin primary with English as a backup:
mutation { updateAgent( id: 12 input: { ttsLang: "cmn-CN" sttLanguageCode: "cmn-Hans-CN" sttAlternativeLanguageCodes: ["en-US"] } ) { id }}Troubleshooting
Section titled “Troubleshooting”- Empty transcript, or the wrong script: the primary language is probably not what the speaker used. Fix the primary first; adding alternatives helps less.
confidenceof0: it means “not reported”, not “low”. See the/sttreference.- A voice sounds wrong or is rejected: check the voice name against Google’s list, and check that
ttsLangandvoiceTiermatch it.