Endpoints

Text to speech

Generate spoken audio from text. OpenAI-compatible speech endpoint, WAV output.

POST /api/v1/audio/speech
New to text to speech? Start with the Text to speech guide. This page is the field-by-field reference.

POST /api/v1/audio/speech is mirrored at the OpenAI-style alias /v1/audio/speech. By default the response body is the audio itself — raw WAV bytes with Content-Type: audio/wav — so save it straight to a file or stream it to a player, with no JSON envelope to unwrap. If you'd rather receive JSON, send Accept: application/json and the audio comes back base64-encoded as { "audio": "<base64 wav>", "format": "wav" }; this keeps the connection active on very long generations that an intermediary might otherwise drop for being idle.

Body application/json
model string required
Speech model ID. List the speech models available to you — with their voices — at /models?type=speech.
input string required
The text to speak, up to 4,096 characters.
voice string default: model's default voice
One of the model's preset voices. Each model's voices and its default_voice are listed in its /models entry; omit this field to use the default.
response_format string default: wav
Only wav is supported — responses always return WAV audio (24 kHz, 16-bit mono).

Response

A 200 with the generated audio as the response body (Content-Type: audio/wav). Generation is typically faster than the audio is long, but allow a few seconds for longer passages — render a spinner in interactive UIs.

Which worker served the request

Every speech response names the worker that generated it in the X-Worker-Id and X-Worker-Name headers, next to X-Request-Id. A JSON response (Accept: application/json) also carries it in the body, which is the reliable place to read it: on a very long generation the response starts before a worker has finished, so the headers can't include one.

{
  "audio": "<base64 wav>",
  "format": "wav",
  "duration_ms": 4210,
  "request_id": "3f1c8a2e-5b7d-4e19-9c0a-6d2f8b1e4a73",
  "pendra": {"worker": {"id": "wrk-1a2b3c4d", "name": "london-gpu-01", "version": "3.115.1"},
             "request_id": "3f1c8a2e-5b7d-4e19-9c0a-6d2f8b1e4a73"}
}

The Python and Node SDKs use the JSON form and give you the worker as speech.pendra.worker. See Which worker served the request for what each field means.

Errors

Errors follow the standard error shape:

  • 400 — unsupported response_format, an unknown voice, or input over 4,096 characters.
  • 404 — the model doesn't exist or isn't available on any of your workers.
{
  "detail": "Model 'your-speech-model' is not available on any connected worker"
}

Usage tracking

Speech requests appear under Speech in the console usage view, with the seconds of audio generated per request. This is visibility only — it doesn't change what you pay.