Skip to content

Sending Voice Messages to an Agent

If your integration already receives voice notes — a WhatsApp/Telegram bridge, a call-center recording, a voice-memo upload in your own app — you can transcribe one and feed it into a Wetel agent using the same avatarToken the <vai-avatar> embed widget uses internally. Nothing restricts that token to the widget; this page documents calling it directly from your own backend.

This is a practical, working walkthrough. For the concepts behind each call, see Core API Flow and Authentication.

You need a working sdkStart call already. If you don’t have one yet, set that up first — see Sessions API: sdkStart. Everything below assumes you can already call sdkStart server-side with your API key and get back a sessionId and avatarToken.

Step 1: start a session, keep the avatarToken

Section titled “Step 1: start a session, keep the avatarToken”
const startRes = await fetch(WETEL_GRAPHQL_URL, {
method: "POST",
headers: {
"Content-Type": "application/json",
"x-huat-platform": "customer",
"x-api-key": WETEL_API_KEY, // server-side only, never send this to a client
},
body: JSON.stringify({
query: `mutation Start($input: SdkStartInput!) {
sdkStart(input: $input) { sessionId avatarToken }
}`,
variables: { input: { agentId: YOUR_AGENT_ID } },
}),
});
const { data } = await startRes.json();
const { sessionId, avatarToken } = data.sdkStart;

avatarToken is the credential every call below authenticates with — it’s session-scoped (bound to this sessionId, since 2026-09-08), so keep the pair together.

/stt lives on the avatar-service host, not the main GraphQL API — a separate REST service with its own hostname. Base64-encode your audio and POST it with the avatarToken as a bearer token.

Terminal window
curl -X POST https://avatar.wetel.dev/stt \
-H "Authorization: Bearer <AVATAR_TOKEN>" \
-H "Content-Type: application/json" \
-d '{
"audioBase64": "T2dnUwACAAAAAAAAAAB...",
"audioEncoding": "OGG_OPUS",
"languageCode": "en-US"
}'

Or from the same backend that just called sdkStart:

const audioBase64 = rawAudioBuffer.toString("base64"); // your voice-note bytes
const sttRes = await fetch("https://avatar.wetel.dev/stt", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${avatarToken}`,
},
body: JSON.stringify({
audioBase64,
audioEncoding: "OGG_OPUS", // native WhatsApp/Telegram voice-note format
languageCode: "en-US",
}),
});
const { transcript, confidence, languageCode } = await sttRes.json();

OGG_OPUS is exactly what WhatsApp and Telegram voice notes already are — an OGG container with Opus-encoded audio. You do not need to transcode the file before sending it; forward the raw bytes as-is.

{
"transcript": "What time does the library close today?",
"confidence": 0.94,
"languageCode": "en-US"
}
FieldTypeNotes
transcriptstringThe transcribed text
confidencenumber0–1. 0 means “not reported”, not “low”
languageCodestringThe language the audio was actually recognised in. Since 2026-09-28 this can differ from the one you sent

There’s no audio or intermediate state to clean up — this call is stateless; the transcript is everything you need going forward.

Step 4: feed the transcript into the agent

Section titled “Step 4: feed the transcript into the agent”

Send the transcript as a normal turn via sdkSendMessage — same mutation, same credential, as if the user had typed it.

await fetch(WETEL_GRAPHQL_URL, {
method: "POST",
headers: {
"Content-Type": "application/json",
"x-huat-platform": "customer",
Authorization: `Bearer ${avatarToken}`, // avatarToken here, not x-api-key
},
body: JSON.stringify({
query: `mutation Send($input: SdkSendMessageInput!) { sdkSendMessage(input: $input) }`,
variables: { input: { sessionId, text: transcript } },
}),
});

sdkSendMessage returns immediately — the agent’s reply arrives asynchronously over the sessionEvents subscription, exactly like any other turn. See Events & Subscriptions if you haven’t wired that up yet.

Constraints to know before you build on this

Section titled “Constraints to know before you build on this”
  • You don’t have to guess the user’s language. Since 2026-09-28, /stt recognises against English, Mandarin, Malay and Cantonese as fallbacks, capped at 3 alongside your primary languageCode. The response’s languageCode tells you which language it actually heard. Before this, sending an English code for a Mandarin or Malay voice note didn’t error; it returned a short garbage transcript. Two things to handle on your side. (1) If you drop low-confidence transcripts, apply the threshold only when confidence > 0. Some correct fallback-language transcripts, Malay in particular, come back with no confidence reported, which shows up as 0. (2) A clip that switches language mid-sentence is still transcribed in one dominant language. Hokkien and Hakka are not supported. Override the fallback list with alternativeLanguageCodes (max 3), or send [] to switch fallback off. See Language fallback and the detected language.
  • Accepted audioEncoding values: MP3, LINEAR16, WEBM_OPUS, OGG_OPUS. Don’t send anything else — it’s rejected as a validation error before it reaches the speech provider.
  • No chunking needed for realistic voice-note lengths. /stt automatically switches from synchronous transcription to an asynchronous long-running path for audio over ~50 seconds — you send the whole file in one call either way, and the endpoint transparently handles both. (Before 2026-09-25 this endpoint had a hard, silent ~60-second ceiling and the workaround was client-side chunking; that’s no longer necessary and there’s no server-side way to stitch chunks back together, so don’t build a chunking path against this endpoint.) A long call (multiple minutes of audio) takes proportionally longer to respond — the request stays open until transcription finishes, there’s no separate polling step on your side.
  • Pauses don’t cut the transcript short. The speech provider often splits a clip into several segments at a pause, and /stt joins every segment into one transcript (with confidence averaged across them). Before 2026-09-26 this was broken for clips under ~50 seconds, which covers most real voice notes. On that path only the first segment was kept, so a voice note like “I want to check on my package… [pause] …the tracking number is 1234” came back as just the first half, with no error and a normal-looking confidence. The agent then answered half a question. Clips over ~50 seconds were never affected. If you added client-side workarounds for “the bot only heard the start of my message,” such as asking users not to pause or re-sending audio, you can remove them.
  • Confirmed working up to ~5 minutes; a much longer clip (10+ minutes) may still fail. The underlying speech provider has its own upper bound on audio embedded directly in the request, which we haven’t fully mapped for every audio encoding — a typical voice note or conversational turn is nowhere near it, but a channel with no recording-length cap (a user can record a very long WhatsApp voice message, for instance) could theoretically hit it. If this happens, the call fails with a clear 500 and an explicit reason in the response body — not silently, and not with corrupted output. If you’re seeing this in practice, reach out — support for much longer audio (tens of minutes) is a straightforward addition on our side if it’s actually needed.
  • 10 requests/minute per avatar token, enforced consistently across the fleet. STT is billed per 15-second audio block upstream, which is why this limit is tighter than /tts’s 30/min — see Avatar Service REST API: POST /stt for the full reference.
  • Max request size: roughly 4MB of raw audio (~5,600,000 base64 characters). Comfortably more than any single voice note; if you’re hitting this, you’re likely sending something other than a short recorded message.
  • Sample rate: leave sampleRateHertz unset for OGG_OPUS — it’s auto-detected from the file itself, not guessed. Different apps record voice notes at different real sample rates (we confirmed this the hard way: Telegram’s own voice notes are 48kHz, and a partner separately reported WhatsApp uses a different rate than Telegram) — so /stt reads the real rate directly out of the audio’s own OGG container header instead of assuming one value for every source. This means you don’t need to know or track which rate your particular voice-note source uses; just send the file as-is. If you DO pass an explicit sampleRateHertz, it’s used verbatim and skips auto-detection — only do this if you’re certain of the real rate (e.g. you re-encoded the audio yourself), since a wrong explicit value produces a clean but empty transcript with no error, not a helpful failure. This only applies to OGG_OPUS; WEBM_OPUS (the browser’s MediaRecorder API) still defaults to a fixed 48kHz, which has always been correct for that source.

The languageCode and alternativeLanguageCodes you send are yours to choose on a direct /stt call. To use the same languages the agent is configured for, read sttLanguageCode and sttAlternativeLanguageCodes from the sdkStart result (live in production since 2026-10-01) and forward them. See Speech-to-text languages.