Sending Voice Messages to an Agent
If your integration already receives voice notes — a WhatsApp/Telegram bridge, a call-center recording, a voice-memo upload in your own app — you can transcribe one and feed it into a Wetel agent using the same avatarToken the <vai-avatar> embed widget uses internally. Nothing restricts that token to the widget; this page documents calling it directly from your own backend.
This is a practical, working walkthrough. For the concepts behind each call, see Core API Flow and Authentication.
Prerequisite
Section titled “Prerequisite”You need a working sdkStart call already. If you don’t have one yet, set that up first — see Sessions API: sdkStart. Everything below assumes you can already call sdkStart server-side with your API key and get back a sessionId and avatarToken.
Step 1: start a session, keep the avatarToken
Section titled “Step 1: start a session, keep the avatarToken”const startRes = await fetch(WETEL_GRAPHQL_URL, { method: "POST", headers: { "Content-Type": "application/json", "x-huat-platform": "customer", "x-api-key": WETEL_API_KEY, // server-side only, never send this to a client }, body: JSON.stringify({ query: `mutation Start($input: SdkStartInput!) { sdkStart(input: $input) { sessionId avatarToken } }`, variables: { input: { agentId: YOUR_AGENT_ID } }, }),});const { data } = await startRes.json();const { sessionId, avatarToken } = data.sdkStart;avatarToken is the credential every call below authenticates with — it’s session-scoped (bound to this sessionId, since 2026-09-08), so keep the pair together.
Step 2: send the audio to /stt
Section titled “Step 2: send the audio to /stt”/stt lives on the avatar-service host, not the main GraphQL API — a separate REST service with its own hostname. Base64-encode your audio and POST it with the avatarToken as a bearer token.
curl -X POST https://avatar.wetel.dev/stt \ -H "Authorization: Bearer <AVATAR_TOKEN>" \ -H "Content-Type: application/json" \ -d '{ "audioBase64": "T2dnUwACAAAAAAAAAAB...", "audioEncoding": "OGG_OPUS", "languageCode": "en-US" }'Or from the same backend that just called sdkStart:
const audioBase64 = rawAudioBuffer.toString("base64"); // your voice-note bytes
const sttRes = await fetch("https://avatar.wetel.dev/stt", { method: "POST", headers: { "Content-Type": "application/json", Authorization: `Bearer ${avatarToken}`, }, body: JSON.stringify({ audioBase64, audioEncoding: "OGG_OPUS", // native WhatsApp/Telegram voice-note format languageCode: "en-US", }),});const { transcript, confidence, languageCode } = await sttRes.json();OGG_OPUS is exactly what WhatsApp and Telegram voice notes already are — an OGG container with Opus-encoded audio. You do not need to transcode the file before sending it; forward the raw bytes as-is.
Step 3: read the response
Section titled “Step 3: read the response”{ "transcript": "What time does the library close today?", "confidence": 0.94, "languageCode": "en-US"}| Field | Type | Notes |
|---|---|---|
transcript | string | The transcribed text |
confidence | number | 0–1. 0 means “not reported”, not “low” |
languageCode | string | The language the audio was actually recognised in. Since 2026-09-28 this can differ from the one you sent |
There’s no audio or intermediate state to clean up — this call is stateless; the transcript is everything you need going forward.
Step 4: feed the transcript into the agent
Section titled “Step 4: feed the transcript into the agent”Send the transcript as a normal turn via sdkSendMessage — same mutation, same credential, as if the user had typed it.
await fetch(WETEL_GRAPHQL_URL, { method: "POST", headers: { "Content-Type": "application/json", "x-huat-platform": "customer", Authorization: `Bearer ${avatarToken}`, // avatarToken here, not x-api-key }, body: JSON.stringify({ query: `mutation Send($input: SdkSendMessageInput!) { sdkSendMessage(input: $input) }`, variables: { input: { sessionId, text: transcript } }, }),});sdkSendMessage returns immediately — the agent’s reply arrives asynchronously over the sessionEvents subscription, exactly like any other turn. See Events & Subscriptions if you haven’t wired that up yet.
Constraints to know before you build on this
Section titled “Constraints to know before you build on this”- You don’t have to guess the user’s language. Since 2026-09-28,
/sttrecognises against English, Mandarin, Malay and Cantonese as fallbacks, capped at 3 alongside your primarylanguageCode. The response’slanguageCodetells you which language it actually heard. Before this, sending an English code for a Mandarin or Malay voice note didn’t error; it returned a short garbage transcript. Two things to handle on your side. (1) If you drop low-confidence transcripts, apply the threshold only whenconfidence > 0. Some correct fallback-language transcripts, Malay in particular, come back with no confidence reported, which shows up as0. (2) A clip that switches language mid-sentence is still transcribed in one dominant language. Hokkien and Hakka are not supported. Override the fallback list withalternativeLanguageCodes(max 3), or send[]to switch fallback off. See Language fallback and the detected language. - Accepted
audioEncodingvalues:MP3,LINEAR16,WEBM_OPUS,OGG_OPUS. Don’t send anything else — it’s rejected as a validation error before it reaches the speech provider. - No chunking needed for realistic voice-note lengths.
/sttautomatically switches from synchronous transcription to an asynchronous long-running path for audio over ~50 seconds — you send the whole file in one call either way, and the endpoint transparently handles both. (Before 2026-09-25 this endpoint had a hard, silent ~60-second ceiling and the workaround was client-side chunking; that’s no longer necessary and there’s no server-side way to stitch chunks back together, so don’t build a chunking path against this endpoint.) A long call (multiple minutes of audio) takes proportionally longer to respond — the request stays open until transcription finishes, there’s no separate polling step on your side. - Pauses don’t cut the transcript short. The speech provider often splits a clip into several segments at a pause, and
/sttjoins every segment into onetranscript(withconfidenceaveraged across them). Before 2026-09-26 this was broken for clips under ~50 seconds, which covers most real voice notes. On that path only the first segment was kept, so a voice note like “I want to check on my package… [pause] …the tracking number is 1234” came back as just the first half, with no error and a normal-looking confidence. The agent then answered half a question. Clips over ~50 seconds were never affected. If you added client-side workarounds for “the bot only heard the start of my message,” such as asking users not to pause or re-sending audio, you can remove them. - Confirmed working up to ~5 minutes; a much longer clip (10+ minutes) may still fail. The underlying speech provider has its own upper bound on audio embedded directly in the request, which we haven’t fully mapped for every audio encoding — a typical voice note or conversational turn is nowhere near it, but a channel with no recording-length cap (a user can record a very long WhatsApp voice message, for instance) could theoretically hit it. If this happens, the call fails with a clear
500and an explicit reason in the response body — not silently, and not with corrupted output. If you’re seeing this in practice, reach out — support for much longer audio (tens of minutes) is a straightforward addition on our side if it’s actually needed. - 10 requests/minute per avatar token, enforced consistently across the fleet. STT is billed per 15-second audio block upstream, which is why this limit is tighter than
/tts’s 30/min — see Avatar Service REST API:POST /sttfor the full reference. - Max request size: roughly 4MB of raw audio (~5,600,000 base64 characters). Comfortably more than any single voice note; if you’re hitting this, you’re likely sending something other than a short recorded message.
- Sample rate: leave
sampleRateHertzunset forOGG_OPUS— it’s auto-detected from the file itself, not guessed. Different apps record voice notes at different real sample rates (we confirmed this the hard way: Telegram’s own voice notes are 48kHz, and a partner separately reported WhatsApp uses a different rate than Telegram) — so/sttreads the real rate directly out of the audio’s own OGG container header instead of assuming one value for every source. This means you don’t need to know or track which rate your particular voice-note source uses; just send the file as-is. If you DO pass an explicitsampleRateHertz, it’s used verbatim and skips auto-detection — only do this if you’re certain of the real rate (e.g. you re-encoded the audio yourself), since a wrong explicit value produces a clean but empty transcript with no error, not a helpful failure. This only applies toOGG_OPUS;WEBM_OPUS(the browser’s MediaRecorder API) still defaults to a fixed 48kHz, which has always been correct for that source.
Matching the agent’s languages
Section titled “Matching the agent’s languages”The languageCode and alternativeLanguageCodes you send are yours to choose on a direct /stt call. To use the same languages the agent is configured for, read sttLanguageCode and sttAlternativeLanguageCodes from the sdkStart result (live in production since 2026-10-01) and forward them. See Speech-to-text languages.
See also
Section titled “See also”- Avatar Service REST API — the full
/stt//ttsreference this page builds on. - Sessions API —
sdkStart,sdkSendMessage, and the rest of the SDK flow. - Voice Interface — how
<vai-avatar>uses this same/sttendpoint internally for a live, in-browser voice conversation. - Headless Agents (No UI Required) — if you’re bridging a messaging channel end to end, not just transcribing one message.
- Image Understanding for an Agent — the equivalent walkthrough for image attachments feeding a vision-capable model call.