Avatar Service REST API
Unlike the rest of the platform, avatar/voice operations are exposed over a separate REST service, not the GraphQL API at api.wetel.dev/graphql. This service handles text-to-speech (TTS) synthesis, speech-to-text (STT) transcription, and serving the catalog of available 3D avatar models. Requests go to your assigned avatar-service endpoint (referred to below as https://avatar.wetel.dev — confirm the exact hostname for your environment with the platform team).
This page assumes you’re already familiar with Avatar Config and Sessions from the main GraphQL API. See also <vai-avatar> and Voice Interface for the higher-level client components that call these endpoints for you — most integrations should reach for those rather than calling this REST API directly.
Authentication
Section titled “Authentication”Every endpoint below (except the public model catalog and health check) requires an avatar token — a short-lived, HMAC-signed token that is distinct from your regular JWT or API key. Send it as a standard bearer token:
Authorization: Bearer <AVATAR_TOKEN>You obtain an avatar token from the main GraphQL API, not from this service directly:
- the
avatarTokenquery (see Avatar Config), or - automatically, as part of
sdkStart’s response, when using the embed SDK entry point.
Rate limiting: two independent ceilings per route
Section titled “Rate limiting: two independent ceilings per route”Added 2026-09-17. Every authenticated route below (/tts, /tts/google, /stt, /avatar-session) is governed by two separate, named throttles that must both allow the request:
- A per-client-IP ceiling (the number quoted in each route’s own “Rate limit” line below) — this existed before 2026-09-17 and is unchanged.
- A per-tenant ceiling, new as of 2026-09-17, shared across every caller currently authenticated with that tenant’s avatar tokens (regardless of which IP each caller connects from) — set to 20x the per-IP figure on every route.
A request without a valid, current avatarToken (missing header, malformed token, expired token, bad signature) is governed by the per-IP ceiling only — the per-tenant throttle is skipped entirely for it, not silently pooled into a shared “no tenant” bucket. Exceeding either ceiling returns the same 429-equivalent throttling response; there’s no way from the response alone to tell which of the two was hit.
POST /tts
Section titled “POST /tts”Synthesizes speech from text or SSML.
Auth required: Avatar token (Bearer).
Rate limit: 30 requests/minute per client IP, and 600 requests/minute per tenant (added 2026-09-17 — see Rate limiting: two independent ceilings per route above).
Request body:
| Field | Type | Required | Notes |
|---|---|---|---|
text | string | One of text/ssml | Plain text to synthesize |
ssml | string | One of text/ssml | SSML markup; takes precedence over text if both are set |
voiceName | string | Yes | |
languageCode | string | Yes | e.g. en-US |
speakingRate | number | No | 0.25–4.0 |
audioEncoding | string | No | MP3 or LINEAR16 |
Response: 200 OK
| Field | Type | Notes |
|---|---|---|
audioContent | string | Base64-encoded audio |
timepoints | array | [{ markName: string, timeSeconds: number }] — SSML mark timing, for lip-sync |
curl -X POST https://avatar.wetel.dev/tts \ -H "Authorization: Bearer <AVATAR_TOKEN>" \ -H "Content-Type: application/json" \ -d '{ "text": "Hello, how can I help you today?", "voiceName": "en-US-Neural2-F", "languageCode": "en-US" }'{ "audioContent": "//uQxAAAAAAAAAAAAAAAAAAAAAAAAAAA...", "timepoints": []}POST /tts/google
Section titled “POST /tts/google”A request/response-shape-compatible variant of /tts, for third-party avatar rendering libraries (such as TalkingHead) that expect the Google Cloud Text-to-Speech request format instead of this service’s native shape.
Auth required: Avatar token (Bearer).
Rate limit: 30 requests/minute per client IP, and 600 requests/minute per tenant (added 2026-09-17 — see Rate limiting: two independent ceilings per route above).
Request body (Google Cloud TTS shape):
{ "input": { "text": "Hello there", "ssml": null }, "voice": { "languageCode": "en-US", "name": "en-US-Neural2-F" }, "audioConfig": { "speakingRate": 1.0 }}Only input.text/input.ssml, voice.name, voice.languageCode, and audioConfig.speakingRate are read — any other fields a third-party library sends alongside these (e.g. enableTimePointing) are accepted and ignored, not rejected.
Response: Same shape as /tts — { audioContent: string, timepoints: [...] }.
POST /tts/gemini
Section titled “POST /tts/gemini”Added 2026-10-01, live in production. An optional text-to-speech engine for audio that is not played back in real time, such as voice notes. It uses a Gemini speech model; /tts and /tts/google are unchanged and remain the right choice for the lip-synced avatar.
Same authentication and rate limits as the other /tts routes. Request body:
| Field | Type | Notes |
|---|---|---|
text | string | Required, up to 5,000 characters. Plain text only. |
voiceName | string | Optional Gemini voice such as Kore (the default) or Puck. Any Cloud-style id (it contains -) selects the default voice. |
Response 200:
{ "audioContent": "<base64 WAV>", "timepoints": [], "mimeType": "audio/wav" }What differs from /tts/google:
- Audio is a 24 kHz, 16-bit mono WAV. MP3 and OGG are not produced.
timepointsis always empty, so it cannot drive lip-sync.- Slower: about 3 to 4 seconds for a two-sentence reply in our checks, against under a second for Cloud voices.
- The language is detected from the text.
languageCode,speakingRateand SSML are not accepted. - No silent fallback. If the engine is not configured on the environment, the route answers
503; it never returns audio from another engine. - Metered like
/tts, per character.
POST /stt
Section titled “POST /stt”Transcribes an audio clip to text.
Auth required: Avatar token (Bearer).
Rate limit: 10 requests/minute per client IP — lower than TTS because STT is billed per 15-second audio block upstream, so this limit is intentionally more conservative — and 200 requests/minute per tenant (added 2026-09-17 — see Rate limiting: two independent ceilings per route above).
Request body:
| Field | Type | Required | Notes |
|---|---|---|---|
audioBase64 | string | Yes | Base64-encoded audio, max ~4MB of raw audio (~5,600,000 base64 characters) |
audioEncoding | string | Yes | MP3, LINEAR16, WEBM_OPUS, or OGG_OPUS |
languageCode | string | Yes | The primary language, e.g. en-US. See Language fallback below. |
alternativeLanguageCodes | string[] | No | Up to 3 fallback languages (BCP-47). Omit to use the default set; send [] to recognise in languageCode only. More than 3 is rejected with a 400. See below. |
sampleRateHertz | number | No | 8000–48000. For OGG_OPUS, leave this unset — see below. If omitted for WEBM_OPUS: 48000. If omitted for MP3/LINEAR16: 16000. |
OGG_OPUS (OGG container, Opus codec) is the native format WhatsApp and Telegram voice notes are recorded in — send one straight through without transcoding it first. It’s a distinct value from WEBM_OPUS because the two use different container formats despite both carrying Opus audio.
For OGG_OPUS, sampleRateHertz is auto-detected from the file itself when omitted, not guessed. Different source apps record at different real rates — Telegram’s voice notes are 48kHz; a different real device/app is not guaranteed to match. Rather than assume one value, this endpoint reads the real rate directly out of the audio’s own OGG container header (the same field standard tools like ffprobe read). If you do pass an explicit sampleRateHertz for OGG_OPUS, it’s used as given and skips auto-detection entirely — only do this if you know the real rate for certain, since a wrong value produces a clean 200 with an empty transcript, not an error.
Response: 200 OK
| Field | Type | Notes |
|---|---|---|
transcript | string | |
confidence | number | 0–1. 0 means “not reported”, not “low” — see below. |
languageCode | string | The language the audio was actually recognised in, which can differ from the one you requested. See below. |
attempts | number | How many recognition calls were made: 1 normally, or 2 when a fallback language won but returned an empty or mixed-script (Han plus Latin) transcript and was automatically rechecked. |
curl -X POST https://avatar.wetel.dev/stt \ -H "Authorization: Bearer <AVATAR_TOKEN>" \ -H "Content-Type: application/json" \ -d '{ "audioBase64": "UklGRi...", "audioEncoding": "WEBM_OPUS", "languageCode": "en-US" }'{ "transcript": "What time does the library close today?", "confidence": 0.94, "languageCode": "en-US"}transcript covers the whole clip. When the speech provider splits audio into several segments (typically at a pause), they are joined with spaces into one string, and confidence is their average. Before 2026-09-26, audio under ~50 seconds returned only the first segment, which silently dropped everything after the first pause. See Sending Voice Messages to an Agent.
Language fallback and the detected language
Section titled “Language fallback and the detected language”Added 2026-09-28. Every /stt call is now sent to the speech provider with fallback languages as well as your primary languageCode, and the provider picks whichever one the audio is actually in. Before this, each call recognised in exactly one language. A voice note in a different language didn’t fail: it came back as a short, confident-looking garbage transcript (a Mandarin clip sent as en-US came back as "for").
The default fallback set, in priority order, is en-US (English) and ms-MY (Malay). Your primary is removed from that list, leaving at most two fallbacks. (Shipping 2026-09-29: Mandarin and Cantonese are being removed from the defaults because accented English clips are being misclassified as Chinese — see below.)
- Override it with
alternativeLanguageCodes: up to 3 BCP-47 codes of your choice. The Wetel-owned voice clients fill this in for you from the agent’s stored settings (live in production since 2026-10-01): see Speech-to-text languages for the per-agentsttLanguageCode/sttAlternativeLanguageCodesfields, the Telegram connector, and the<vai-avatar>stt-lang/stt-alternative-langsattributes. - Turn it off with
"alternativeLanguageCodes": [], which recognises inlanguageCodeonly (the old behaviour). - Why the hard cap of 3: that is the speech provider’s real limit. Past it, the provider doesn’t return an error. It silently ignores the whole list, which quietly turns fallback off. We verified this live, so
/sttrejects a 4th entry with a400rather than let that happen.
languageCode in the response is now the detected language. It used to echo whatever you sent. It now reports the language the provider actually recognised, spelled the way you spelled it in the request (so cmn-Hans-CN comes back as cmn-Hans-CN, Cantonese comes back as yue-Hant-HK if you sent it that way). When a clip has several segments in different languages, the language covering the most transcript text wins. If nothing was recognised, it falls back to your requested languageCode. Use this field to decide, for example, which language the agent should reply in.
Fallback re-check: automatic quality gate (live in production since 2026-10-01). When a fallback language wins the recognition but returns an empty transcript, or returns mixed-script garbage (a small Chinese fragment plus an English clause, like "蓝色in Chinese"), a short clip (under ~50 seconds) is automatically re-checked once in the primary languageCode only, with no fallbacks. If the primary-only result is at least as good as the first, it replaces the fallback’s result. This re-check is metered as an extra 15-second audio block, so attempts: 2 in the response means two billed recognition calls. Known limits: (1) Cantonese heard on a Mandarin primary (zh/zh-*/cmn-*) counts as the primary — every Han-script winner on a Han-script primary is treated as the primary, so it is never retried. (2) Japanese fallbacks are re-checked only when the transcript is empty, never for mixed-script. (3) A genuinely garbled transcript (Mandarin with short English phrases woven in naturally) can trigger the re-check but the original is kept unless the primary-only result is at least as confident.
confidence: 0 means “not reported”. The provider sometimes returns no confidence at all when it recognised the audio through a fallback language. We have seen this with correct Malay transcripts. /stt reports that as 0, which is the provider’s own “not set” value. If your integration throws away low-confidence transcripts, apply your threshold only when confidence > 0. Otherwise you will discard correct transcripts. Wetel’s own Telegram connector was changed to work this way.
Known limits:
- Mandarin and Cantonese are not in the default fallback set (the default is
en-USandms-MY, live since 2026-10-01). The provider classifies accented English clips as Chinese (cmn-Hans-CNoryue-Hant-HK), even when the transcript is 100% English — a misclassifiedlanguageCodecan be corrected when the script doesn’t match the content, but preventing the mislabel in the first place is more reliable than correcting it after the fact. See Choosing Languages and Voices. If your application genuinely needs Chinese detection, passalternativeLanguageCodes: ["cmn-Hans-CN", "yue-Hant-HK"]or similar explicitly. - Code-switching within one clip is not handled. For mixed speech (English, then Mandarin, then English again), the provider picks one dominant language and the other language’s words are usually lost.
- Hokkien and Hakka are not supported by the speech provider at any setting.
- Tested on synthetic voices only so far. Real, accented, conversational voice notes have not been measured yet. Check against your own traffic before you depend on this.
Update, 2026-09-29 — the default fallback set already changed (live in production and staging):
The speech provider classifies accented English clips (especially short clips like “Hi testing one two three”) as Mandarin or Cantonese, even though the transcript is pure English. Mandarin and Cantonese have been removed from the defaults as of 2026-09-29 (now live on staging and production) — the default set is now [en-US, ms-MY] only. This false-positive misclassification is costlier than the fix-after-the-fact language-tag correction. If a segment’s languageCode is cmn-Hans-CN or yue-Hant-HK but the transcript contains zero Han characters, the tag is still corrected to the first non-Han language you requested — but this correction only works if the fallback was included in the first place. If your integration needs Chinese detection, pass it explicitly: "alternativeLanguageCodes": ["cmn-Hans-CN", "yue-Hant-HK"] (or whichever Chinese variants you support) in the request. There is also a separate issue, unfixed: a genuinely garbled transcript traces to which language you send as languageCode (primary), not to the fallback list. Testing against real voice notes found the primary carries much more weight in the provider’s recognition than the alternatives do: the same audio produced a clean transcript with en-US as primary, and a corrupted transcript with cmn-Hans-CN as primary — reproduced exactly, byte for byte. Send your best actual guess as the primary languageCode, not just a placeholder — it isn’t “one candidate among equals” the way the alternatives are. We tested narrowing alternativeLanguageCodes as a potential fix and found it does not help (it made results worse in some cases) — the fix is a better primary guess, not a shorter fallback list.
The same detected-language behaviour applies to the LLM Playground’s transcribeSpeech query.
If the underlying speech provider is rate-limited or unavailable, this endpoint returns 503 Service Unavailable rather than a generic 500 — worth handling as a distinct, retryable case in your client.
As of 2026-09-25, this endpoint automatically routes audio over ~50 seconds through an asynchronous long-running transcription path instead of synchronous recognition — you still send the whole file in a single call and get the transcript back in the same response, just with proportionally longer latency for longer audio. (Previously this endpoint had a hard, silently-enforced ~60-second ceiling; audio longer than that failed outright. That limit no longer applies — no client-side chunking workaround is needed, and none is supported.) Confirmed working up to ~5 minutes; a much longer clip could still fail with a 500 (a real, explicit error — not silent) — see Sending Voice Messages to an Agent: constraints if this matters for your integration.
Speech engines: /stt, /stt/google, /stt/gemini
Section titled “Speech engines: /stt, /stt/google, /stt/gemini”Added 2026-10-01, live in production. There are two speech-to-text engines, and you can pick one per call:
| Route | Engine |
|---|---|
POST /stt | The deployment’s default engine. Since 2026-10-02 this is Gemini, with Cloud Speech-to-Text answering automatically if Gemini errors. |
POST /stt/google | Always Cloud Speech-to-Text (the behaviour described above). |
POST /stt/gemini | Always the Gemini speech model (gemini-3.5-transcribe). |
All three take the same request body, the same avatar token and the same rate limits, and return the same response shape. Calling an explicit route on an environment where that engine isn’t configured returns 503, never a transcript from the other engine. Plain /stt uses the default engine (see above); confidence is 0 when Gemini answered. Call /stt/google to keep the previous behaviour exactly.
What differs on /stt/gemini, measured on the same audio:
- Faster and cleaner text. About 1.4–3.0 s against 3.2–3.4 s for Cloud STT in our checks, and the text comes back capitalised and punctuated.
- No language hint needed. The model detects the language itself, so
languageCodedoes not steer it andalternativeLanguageCodesis ignored. A Malay clip sent asen-USstill came back correct. ThelanguageCodein the response simply echoes the one you sent: this engine does not report a detected language. - No confidence score.
confidenceis0, which means “not reported”, the same as it does for the other engine when it has no score. - No fallback re-check.
attemptsis not returned. - Formats.
MP3,WEBM_OPUS,OGG_OPUSandLINEAR16all work;LINEAR16can be a WAV file or raw 16-bit mono PCM (setsampleRateHertzfor raw PCM, default 16000).
Billing and credit metering are unchanged: both engines are metered per started 15-second block.
POST /avatar-session
Section titled “POST /avatar-session”Negotiates a HOSTED_API (photorealistic video avatar) rendering session. Only relevant if your agent’s avatarBackend is configured to HOSTED_API — see Avatar & 3D Rendering: Avatar rendering options. Not used for the default 3D (CLIENT_3D) rendering path, which is driven entirely by sdkStart’s response instead.
Auth required: Avatar token (Bearer).
Rate limit: 10 requests/minute per client IP, and 200 requests/minute per tenant (added 2026-09-17 — see Rate limiting: two independent ceilings per route above). Lower than TTS/STT because this negotiates a real, per-minute-billed hosted video session with the underlying vendor, not a per-utterance call.
Request body:
| Field | Type | Required | Notes |
|---|---|---|---|
anamAvatarId | string | No | Must match the value already configured on the agent’s avatarBackend config — sent only so a mismatch can be rejected loudly. |
anamAvatarModel | string | No | Same as above. |
Your own client code never chooses or overrides the vendor persona/model here — the real value comes from the agent’s server-side avatarBackendConfig, embedded in the avatarToken at sdkStart time. These fields exist only so the server can detect and reject a mismatch, not so the browser can request an arbitrary persona.
Response: 200 OK
| Field | Type | Notes |
|---|---|---|
sessionToken | string | A vendor session credential. The browser never sees the vendor’s own API key. |
curl -X POST https://avatar.wetel.dev/avatar-session \ -H "Authorization: Bearer <AVATAR_TOKEN>" \ -H "Content-Type: application/json" \ -d '{}'{ "sessionToken": "eyJhbGciOi..."}A negotiation failure (missing vendor credentials, vendor outage) returns a generic 500 with no vendor-specific detail — the underlying error is logged server-side, never surfaced to the browser, since it may reference internal config state.
GET /avatars/models
Section titled “GET /avatars/models”Returns the catalog of available 3D avatar models.
Auth required: None — this is a public catalog listing.
Response: 200 OK
{ "models": [ { "id": "avatar-01", "label": "Aria", "gender": "female", "path": "/avatars/models/aria.glb" } ]}| Field | Type | Notes |
|---|---|---|
id | string | |
label | string | Display name |
gender | string | |
path | string | Relative path to the .glb model asset |
curl https://avatar.wetel.dev/avatars/modelsGET /health
Section titled “GET /health”A basic health check endpoint — returns a 200 OK when the service is up. No auth required, no meaningful response body to parse; use it only for uptime/liveness checks, not as a functional API.
For the config that determines which voice/avatar a given agent uses by default, see Avatar Config. For how sessions and avatar tokens fit together, see Sessions. For the client-side components that wrap this API, see <vai-avatar> and Voice Interface.