Skip to content

LLM API

Wetel already sits in front of multiple LLM providers (Bedrock Mantle, WebbyxAI/Claude Haiku, Gemini) for its own agent/workflow traffic. This page documents the two operations that let you call an LLM directly, as a simple middleman — you send a prompt, Wetel routes it to a provider and bills the call, you get a completion back. No session, no agent, no workflow required, and nothing to set up on your own account with AWS, Google, or any other model vendor.

This is deliberately the simplest possible version of “bring your model calls to Wetel” — a flat, metered pass-through billed the same way any other generation already is on your account. It is not a prepaid-balance or model-marketplace product (no wallet, no per-model markup, no arbitrary third-party model selection) — this page covers what’s actually built today.

generateText accepts either a dashboard JWT or an X-Api-Key header (as of 2026-09-24 — see the changelog) — a partner backend holding only an API key works exactly as before, and a dashboard-authenticated caller can now use it too. startTextGeneration and the textGenerationEvents subscription remain X-Api-Key only — generate one via Tenant: API keys. Set it as an X-Api-Key header on HTTP requests, and in connectionParams (not a header) for the WebSocket subscription — see Streaming below.

Every request must also include the x-huat-platform: customer header, same as every other operation in this API.

Single-response — sends prompt to a chosen provider chain and returns the full completion. Real cost is incurred and metered the same as any other generation call on your account.

generateText(input: GenerateTextInput!): GenerateTextResultDto!

GenerateTextInput fields

FieldTypeRequiredNotes
promptString!yesMax 4000 characters.
providerLlmProviderTypenoBEDROCK_MANTLE | WEBBYXAI_HAIKU | GEMINI_DIRECT. Omit to use the app-wide default.
modelIdStringnoOverrides which specific model the selected provider chain calls. For WEBBYXAI_HAIKU: claude-haiku-4-5-20251001 (default), claude-sonnet-4-6, claude-opus-4-6, claude-opus-4-6-low, -medium, -high, -max. Ignored by the app-wide default.
images[String!]noFull data: URLs. Consumed by BEDROCK_MANTLE with a vision-capable modelId, and by GEMINI_DIRECT (native inline-image support). WEBBYXAI_HAIKU also reads images (every Claude model on the gateway; an image over 5 MB as base64 is answered by Gemini instead). Max 1 image.

Request size: a request body can be up to 10 MB (about 7 MB of raw image data once base64-encoded). Larger bodies are rejected with HTTP 413 before they reach the API. Because that response comes from the gateway in front of the API and carries no CORS headers, a browser reports it as a CORS error, not as a size error. Resize or compress large photos before sending, and send one image per call where you can.

GenerateTextResultDto fields

FieldTypeNotes
textString!The completion.
providerIdString!Which provider actually served this call — can differ from input.provider on a silent fallback.
modelIdString!Which model actually served the call.
costUsdFloat!Real, incurred cost for this call.
creditsUsedApproxInt!Added 2026-09-24. Approximate credit-equivalent of costUsd, for previewing what this call cost under the credits pricing model — a blended estimate, not an exact per-model breakdown.

Request

mutation GenerateText($input: GenerateTextInput!) {
generateText(input: $input) {
text
providerId
modelId
costUsd
creditsUsedApprox
}
}
{ "input": { "prompt": "Write a haiku about databases." } }
POST /graphql
Content-Type: application/json
X-Api-Key: <your-api-key>
x-huat-platform: customer

Or, as of 2026-09-24, a dashboard JWT works in place of X-Api-Key:

POST /graphql
Content-Type: application/json
Authorization: Bearer <jwt>
x-huat-platform: customer

Response

{
"data": {
"generateText": {
"text": "Rows in silence wait\nA query breaks the stillness\nJoins reveal the truth",
"providerId": "bedrock-mantle",
"modelId": "qwen.qwen3-235b-a22b-2507",
"costUsd": 0.00003,
"creditsUsedApprox": 1
}
}
}

To analyze an image, set provider: GEMINI_DIRECT and pass one image as a full data: URL in images. Gemini is the recommended starting point for this — it needs no modelId (it always uses its own vision-capable default model), unlike BEDROCK_MANTLE, which only sees the image at all if you also pick a vision-capable modelId yourself (and silently ignores images if you don’t — see the field table above).

mutation GenerateText($input: GenerateTextInput!) {
generateText(input: $input) {
text
providerId
modelId
costUsd
}
}
{
"input": {
"provider": "GEMINI_DIRECT",
"prompt": "What is shown in this image? Be specific about any text you can read.",
"images": ["data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAA..."]
}
}
{
"data": {
"generateText": {
"text": "The image shows a shipping label with the tracking number 4471002200 printed near a barcode.",
"providerId": "gemini",
"modelId": "gemini-3.1-flash-lite",
"costUsd": 0.00019
}
}
}

A few things worth knowing before your first real call:

  • Max 1 image per call, and it must be a real data:image/(png|jpeg|webp|gif);base64,... URL — a plain HTTP(S) image URL is rejected by input validation, not fetched server-side. If you only have a URL, download it and base64-encode it yourself first.
  • WEBBYXAI_HAIKU (the app-wide default when provider is omitted) has no vision support at all — it silently ignores images and answers based on prompt alone, with no error telling you the image was dropped. Always set provider explicitly when sending an image; don’t rely on the default.
  • Check providerId/modelId in the response, not just what you requested — BEDROCK_MANTLE falls back through its own chain (ultimately landing on Gemini) on a provider-side failure, so a BEDROCK_MANTLE request can come back served by gemini. GEMINI_DIRECT has no fallback of its own — if it fails, the call fails; there’s nothing further to fall back to.

Beyond Gemini, BEDROCK_MANTLE offers several vision-capable models you can select via modelId:

Model IDAccuracy on clean textNotes
mistral.magistral-small-25093/3 exact matchRecommended for text/code extraction; ~1.1s latency. Occasionally renders en-dash instead of hyphen.
mistral.mistral-large-3-675b-instruct3/3 exact matchStrongest overall accuracy. ~1.0s latency.
nvidia.nemotron-nano-12b-v23/3 exact matchLow-cost alternative. ~1.0s latency.
mistral.ministral-3-14b-instruct2/3 exact; 3/3 normalizedCheapest option. Occasionally uses en-dash or drops minor formatting. ~1.1s latency.
google.gemma-3-27b-it2/3 exact; 3/3 normalizedGenerally capable. May occasionally drop currency prefixes. ~1.1s latency.
moonshotai.kimi-k2.52/3 exact; 3/3 normalizedMultilingual. May drop currency prefixes or add leading whitespace. ~1.2s latency.

Important caveat: These models have been tested on clean, synthetic images only (high-contrast printed text, no real photographs). Real-world accuracy on photographed documents, receipts, or labels may differ. If you use image analysis in production, validate against your own image sources before committing to a model. See Workflows: Image Understanding for how to add image input to an agent workflow.

Streaming: startTextGeneration + textGenerationEvents

Section titled “Streaming: startTextGeneration + textGenerationEvents”

For a streamed reply, start generation with a mutation, then subscribe to its events.

startTextGeneration(input: GenerateTextInput!): StartTextGenerationResultDto!

input is the same GenerateTextInput shape as generateText above. Returns immediately:

{ "data": { "startTextGeneration": { "requestId": "b3f1..." } } }

Then subscribe:

subscription TextGenerationEvents($requestId: String!) {
textGenerationEvents(requestId: $requestId) {
requestId
text
done
providerId
modelId
error
}
}

TextGenerationEventDto fields

FieldTypeNotes
requestIdString!Echoes the id from startTextGeneration.
textString!This chunk’s incremental text — append it to what you already have. Empty on the terminal chunk unless error is set.
doneBoolean!true on the final event — no further events arrive after this one.
providerIdStringOnly populated on the terminal (done: true) event, same convention as generateText.
modelIdStringOnly populated on the terminal event.
errorStringSet (with done: true) if generation failed — the provider’s own error message, safe to show directly.

WebSocket auth for this subscription specifically

Section titled “WebSocket auth for this subscription specifically”

Unlike every other subscription in this API (which authenticate via a dashboard JWT), textGenerationEvents authenticates via the same API key as the mutations above — pass it as x-api-key in connectionParams when you open the graphql-ws connection (not as an HTTP header — WebSocket connections don’t carry HTTP headers the same way):

import { createClient } from "graphql-ws";
const client = createClient({
url: "wss://api.wetel.dev/graphql",
// See "Troubleshooting: connect before you start generation" below for
// why lazy: false matters here.
lazy: false,
connectionParams: {
"x-api-key": "<your-api-key>",
"x-huat-platform": "customer",
},
});

A requestId you didn’t start yourself (or that already expired — see below) is rejected with a Forbidden error at subscribe time, not silently ignored.

A requestId is valid to subscribe against for 10 minutes after startTextGeneration returns it — after that it’s treated as never having existed. In practice, subscribe immediately after starting generation; there’s no legitimate reason to hold onto a requestId for later.

Every tenant has a monthly spend ceiling on this API, based on plan tier. Both generateText and startTextGeneration check it up front — if you’re already at or over the cap, the call fails immediately with a BadRequestException (before any provider is called, so it never incurs cost) and a message naming your limit and current spend:

This tenant has reached its monthly LLM spend limit of $5.00 (spent $5.12 so far this month). Upgrade your plan or wait until next month.

This is a hard stop, not a soft warning — plan ahead if you’re running a burst of calls near the end of a billing month. The cap resets at the start of each calendar month.

generateText is rate-limited by two independent ceilings that must both pass (as of 2026-09-24): a per-client-IP limit of 30 requests/minute, and a new per-tenant limit of 60 requests/minute shared across every caller authenticated as that tenant — so a partner backend proxying many end users through one server IP isn’t bottlenecked by the IP tier alone. startTextGeneration keeps its original single per-IP limit of 10 requests/minute. See Rate Limiting: LLM API for the full table.

Chunks never arrive — connect (and subscribe) before calling startTextGeneration

Section titled “Chunks never arrive — connect (and subscribe) before calling startTextGeneration”

Events are delivered over plain Redis pub/sub, not a queue — there is no replay buffer. If generation finishes before your subscription is actually listening, those chunks are gone forever; your subscription just sits there until it times out on your end, with no error from the server (the requestId is still valid, ownership still checks out, there’s simply nothing left to deliver).

For a short prompt against a fast model, generation can complete in well under a second — faster than a fresh WebSocket handshake, in practice. Get the connection established first, then start generation:

const client = createClient({
url: "wss://api.wetel.dev/graphql",
lazy: false, // see below
connectionParams: {
"x-api-key": "<your-api-key>",
"x-huat-platform": "customer",
},
});
// Wait for the socket to actually be open before doing anything else.
await new Promise(resolve => {
client.on("connected", resolve);
});
// NOW start generation and subscribe — not before.
const { requestId } = await startTextGeneration(prompt);
client.subscribe(
{ query: TEXT_GENERATION_EVENTS, variables: { requestId } },
sink
);

graphql-ws’s client defaults to a lazy connection — by design, it doesn’t open the socket at all until the first subscribe() call, so a naive “wait for connected, then do stuff” pattern deadlocks with the default options (connected never fires because nothing has triggered a connection yet). Pass lazy: false at client creation so the socket opens immediately and connected actually fires — this is why the code sample above sets it explicitly.