LLM API
Wetel already sits in front of multiple LLM providers (Bedrock Mantle, WebbyxAI/Claude Haiku, Gemini) for its own agent/workflow traffic. This page documents the two operations that let you call an LLM directly, as a simple middleman — you send a prompt, Wetel routes it to a provider and bills the call, you get a completion back. No session, no agent, no workflow required, and nothing to set up on your own account with AWS, Google, or any other model vendor.
This is deliberately the simplest possible version of “bring your model calls to Wetel” — a flat, metered pass-through billed the same way any other generation already is on your account. It is not a prepaid-balance or model-marketplace product (no wallet, no per-model markup, no arbitrary third-party model selection) — this page covers what’s actually built today.
generateText accepts either a dashboard JWT or an X-Api-Key header (as of 2026-09-24 — see the changelog) — a partner backend holding only an API key works exactly as before, and a dashboard-authenticated caller can now use it too. startTextGeneration and the textGenerationEvents subscription remain X-Api-Key only — generate one via Tenant: API keys. Set it as an X-Api-Key header on HTTP requests, and in connectionParams (not a header) for the WebSocket subscription — see Streaming below.
Every request must also include the x-huat-platform: customer header, same as every other operation in this API.
generateText
Section titled “generateText”Single-response — sends prompt to a chosen provider chain and returns the full completion. Real cost is incurred and metered the same as any other generation call on your account.
generateText(input: GenerateTextInput!): GenerateTextResultDto!GenerateTextInput fields
| Field | Type | Required | Notes |
|---|---|---|---|
prompt | String! | yes | Max 4000 characters. |
provider | LlmProviderType | no | BEDROCK_MANTLE | WEBBYXAI_HAIKU | GEMINI_DIRECT. Omit to use the app-wide default. |
modelId | String | no | Overrides which specific model the selected provider chain calls. For WEBBYXAI_HAIKU: claude-haiku-4-5-20251001 (default), claude-sonnet-4-6, claude-opus-4-6, claude-opus-4-6-low, -medium, -high, -max. Ignored by the app-wide default. |
images | [String!] | no | Full data: URLs. Consumed by BEDROCK_MANTLE with a vision-capable modelId, and by GEMINI_DIRECT (native inline-image support). WEBBYXAI_HAIKU also reads images (every Claude model on the gateway; an image over 5 MB as base64 is answered by Gemini instead). Max 1 image. |
Request size: a request body can be up to 10 MB (about 7 MB of raw image data once base64-encoded). Larger bodies are rejected with HTTP 413 before they reach the API. Because that response comes from the gateway in front of the API and carries no CORS headers, a browser reports it as a CORS error, not as a size error. Resize or compress large photos before sending, and send one image per call where you can.
GenerateTextResultDto fields
| Field | Type | Notes |
|---|---|---|
text | String! | The completion. |
providerId | String! | Which provider actually served this call — can differ from input.provider on a silent fallback. |
modelId | String! | Which model actually served the call. |
costUsd | Float! | Real, incurred cost for this call. |
creditsUsedApprox | Int! | Added 2026-09-24. Approximate credit-equivalent of costUsd, for previewing what this call cost under the credits pricing model — a blended estimate, not an exact per-model breakdown. |
Request
mutation GenerateText($input: GenerateTextInput!) { generateText(input: $input) { text providerId modelId costUsd creditsUsedApprox }}{ "input": { "prompt": "Write a haiku about databases." } }POST /graphqlContent-Type: application/jsonX-Api-Key: <your-api-key>x-huat-platform: customerOr, as of 2026-09-24, a dashboard JWT works in place of X-Api-Key:
POST /graphqlContent-Type: application/jsonAuthorization: Bearer <jwt>x-huat-platform: customerResponse
{ "data": { "generateText": { "text": "Rows in silence wait\nA query breaks the stillness\nJoins reveal the truth", "providerId": "bedrock-mantle", "modelId": "qwen.qwen3-235b-a22b-2507", "costUsd": 0.00003, "creditsUsedApprox": 1 } }}Image analysis
Section titled “Image analysis”To analyze an image, set provider: GEMINI_DIRECT and pass one image as a full data: URL in images. Gemini is the recommended starting point for this — it needs no modelId (it always uses its own vision-capable default model), unlike BEDROCK_MANTLE, which only sees the image at all if you also pick a vision-capable modelId yourself (and silently ignores images if you don’t — see the field table above).
mutation GenerateText($input: GenerateTextInput!) { generateText(input: $input) { text providerId modelId costUsd }}{ "input": { "provider": "GEMINI_DIRECT", "prompt": "What is shown in this image? Be specific about any text you can read.", "images": ["data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAA..."] }}{ "data": { "generateText": { "text": "The image shows a shipping label with the tracking number 4471002200 printed near a barcode.", "providerId": "gemini", "modelId": "gemini-3.1-flash-lite", "costUsd": 0.00019 } }}A few things worth knowing before your first real call:
- Max 1 image per call, and it must be a real
data:image/(png|jpeg|webp|gif);base64,...URL — a plain HTTP(S) image URL is rejected by input validation, not fetched server-side. If you only have a URL, download it and base64-encode it yourself first. WEBBYXAI_HAIKU(the app-wide default whenprovideris omitted) has no vision support at all — it silently ignoresimagesand answers based onpromptalone, with no error telling you the image was dropped. Always setproviderexplicitly when sending an image; don’t rely on the default.- Check
providerId/modelIdin the response, not just what you requested —BEDROCK_MANTLEfalls back through its own chain (ultimately landing on Gemini) on a provider-side failure, so aBEDROCK_MANTLErequest can come back served bygemini.GEMINI_DIRECThas no fallback of its own — if it fails, the call fails; there’s nothing further to fall back to.
Vision-capable models on Bedrock Mantle
Section titled “Vision-capable models on Bedrock Mantle”Beyond Gemini, BEDROCK_MANTLE offers several vision-capable models you can select via modelId:
| Model ID | Accuracy on clean text | Notes |
|---|---|---|
mistral.magistral-small-2509 | 3/3 exact match | Recommended for text/code extraction; ~1.1s latency. Occasionally renders en-dash instead of hyphen. |
mistral.mistral-large-3-675b-instruct | 3/3 exact match | Strongest overall accuracy. ~1.0s latency. |
nvidia.nemotron-nano-12b-v2 | 3/3 exact match | Low-cost alternative. ~1.0s latency. |
mistral.ministral-3-14b-instruct | 2/3 exact; 3/3 normalized | Cheapest option. Occasionally uses en-dash or drops minor formatting. ~1.1s latency. |
google.gemma-3-27b-it | 2/3 exact; 3/3 normalized | Generally capable. May occasionally drop currency prefixes. ~1.1s latency. |
moonshotai.kimi-k2.5 | 2/3 exact; 3/3 normalized | Multilingual. May drop currency prefixes or add leading whitespace. ~1.2s latency. |
Important caveat: These models have been tested on clean, synthetic images only (high-contrast printed text, no real photographs). Real-world accuracy on photographed documents, receipts, or labels may differ. If you use image analysis in production, validate against your own image sources before committing to a model. See Workflows: Image Understanding for how to add image input to an agent workflow.
Streaming: startTextGeneration + textGenerationEvents
Section titled “Streaming: startTextGeneration + textGenerationEvents”For a streamed reply, start generation with a mutation, then subscribe to its events.
startTextGeneration(input: GenerateTextInput!): StartTextGenerationResultDto!input is the same GenerateTextInput shape as generateText above. Returns immediately:
{ "data": { "startTextGeneration": { "requestId": "b3f1..." } } }Then subscribe:
subscription TextGenerationEvents($requestId: String!) { textGenerationEvents(requestId: $requestId) { requestId text done providerId modelId error }}TextGenerationEventDto fields
| Field | Type | Notes |
|---|---|---|
requestId | String! | Echoes the id from startTextGeneration. |
text | String! | This chunk’s incremental text — append it to what you already have. Empty on the terminal chunk unless error is set. |
done | Boolean! | true on the final event — no further events arrive after this one. |
providerId | String | Only populated on the terminal (done: true) event, same convention as generateText. |
modelId | String | Only populated on the terminal event. |
error | String | Set (with done: true) if generation failed — the provider’s own error message, safe to show directly. |
WebSocket auth for this subscription specifically
Section titled “WebSocket auth for this subscription specifically”Unlike every other subscription in this API (which authenticate via a dashboard JWT), textGenerationEvents authenticates via the same API key as the mutations above — pass it as x-api-key in connectionParams when you open the graphql-ws connection (not as an HTTP header — WebSocket connections don’t carry HTTP headers the same way):
import { createClient } from "graphql-ws";
const client = createClient({ url: "wss://api.wetel.dev/graphql", // See "Troubleshooting: connect before you start generation" below for // why lazy: false matters here. lazy: false, connectionParams: { "x-api-key": "<your-api-key>", "x-huat-platform": "customer", },});A requestId you didn’t start yourself (or that already expired — see below) is rejected with a Forbidden error at subscribe time, not silently ignored.
requestId lifetime
Section titled “requestId lifetime”A requestId is valid to subscribe against for 10 minutes after startTextGeneration returns it — after that it’s treated as never having existed. In practice, subscribe immediately after starting generation; there’s no legitimate reason to hold onto a requestId for later.
Monthly spend cap
Section titled “Monthly spend cap”Every tenant has a monthly spend ceiling on this API, based on plan tier. Both generateText and startTextGeneration check it up front — if you’re already at or over the cap, the call fails immediately with a BadRequestException (before any provider is called, so it never incurs cost) and a message naming your limit and current spend:
This tenant has reached its monthly LLM spend limit of $5.00 (spent $5.12 so far this month). Upgrade your plan or wait until next month.This is a hard stop, not a soft warning — plan ahead if you’re running a burst of calls near the end of a billing month. The cap resets at the start of each calendar month.
Rate limits
Section titled “Rate limits”generateText is rate-limited by two independent ceilings that must both pass (as of 2026-09-24): a per-client-IP limit of 30 requests/minute, and a new per-tenant limit of 60 requests/minute shared across every caller authenticated as that tenant — so a partner backend proxying many end users through one server IP isn’t bottlenecked by the IP tier alone. startTextGeneration keeps its original single per-IP limit of 10 requests/minute. See Rate Limiting: LLM API for the full table.
Troubleshooting
Section titled “Troubleshooting”Chunks never arrive — connect (and subscribe) before calling startTextGeneration
Section titled “Chunks never arrive — connect (and subscribe) before calling startTextGeneration”Events are delivered over plain Redis pub/sub, not a queue — there is no replay buffer. If generation finishes before your subscription is actually listening, those chunks are gone forever; your subscription just sits there until it times out on your end, with no error from the server (the requestId is still valid, ownership still checks out, there’s simply nothing left to deliver).
For a short prompt against a fast model, generation can complete in well under a second — faster than a fresh WebSocket handshake, in practice. Get the connection established first, then start generation:
const client = createClient({ url: "wss://api.wetel.dev/graphql", lazy: false, // see below connectionParams: { "x-api-key": "<your-api-key>", "x-huat-platform": "customer", },});
// Wait for the socket to actually be open before doing anything else.await new Promise(resolve => { client.on("connected", resolve);});
// NOW start generation and subscribe — not before.const { requestId } = await startTextGeneration(prompt);client.subscribe( { query: TEXT_GENERATION_EVENTS, variables: { requestId } }, sink);graphql-ws’s client defaults to a lazy connection — by design, it doesn’t open the socket at all until the first subscribe() call, so a naive “wait for connected, then do stuff” pattern deadlocks with the default options (connected never fires because nothing has triggered a connection yet). Pass lazy: false at client creation so the socket opens immediately and connected actually fires — this is why the code sample above sets it explicitly.
See also
Section titled “See also”- Knowledge Base & Embeddings:
embed— the other API-key-tier utility operation, for building your own vector index - Tenant: API keys — generating the
X-Api-Keythis page’s operations use - Events & Subscriptions — general
graphql-wsconnection setup (written for the dashboard-JWT case; this page’s WebSocket auth note above is the one thing that differs fortextGenerationEvents) - API Reference: Overview