Endpoints
Usage & workers
The usage and worker endpoints give your own code the same view the console's Usage and Workers pages show: which requests are running right now, what finished and how, and whether your workers are healthy. Typical uses are a monitoring job that alerts when a worker goes offline, or a service that cancels a request and wants to confirm it has stopped before retrying it.
Permissions
These endpoints need a key with the Usage & workers permission. It's off by default, because most keys only send requests and never need to see your organisation's history. Tick it in the console when you create the key. It's read-only. To change worker settings from your code, see Managing workers.
A key without the permission gets a 403 naming what it's
missing: usage:read for the /usage endpoints,
workers:read for the /workers endpoints. Keys
created before permissions existed are unrestricted and can already call
them. See the authentication page
for how permissions work.
A key only ever sees its own organisation's requests and workers, so there
is no X-Org-Id header to send.
Requests in progress
GET /api/v1/usage/inflight lists the requests running on your
workers right now, oldest first. Add ?worker_id=… to narrow it
to one worker.
curl https://api.pendra.ai/api/v1/usage/inflight \
-H "Authorization: Bearer pdr_sk_..."
[
{
"task_id": "8f14e45f-ceea-467a-9575-2f9d6b1c2a10",
"model": "qwen3.6:27b",
"worker_id": "wrk-3f9b8a2d9e104c1a",
"worker_name": "gpu-box-1",
"request_type": "chat",
"started_at": "2026-09-25T10:14:03Z",
"elapsed_ms": 4210,
"tokens_generated": 118
}
]
tokens_generated is a running estimate for streaming
/api/v1/chat/completions requests. It is null
until the first tokens arrive, and for other kinds of request. For a
privately encrypted request
Pendra can't read the output, so the estimate counts the encrypted chunks
streamed so far instead.
A request stays on this list for as long as it's running, however long that is. It drops off when it finishes, is cancelled or fails.
Confirming a cancelled request has stopped
You cancel a request by closing
the connection. Pendra stops the generation and frees the worker's slot
straight away, so the request drops off the in-progress list. To confirm
that before you retry, record the time you cancelled, then poll
/usage/inflight until no request for that model that started
before then is still listed:
import time
from datetime import datetime, timezone
import requests
API = "https://api.pendra.ai/api/v1"
HEADERS = {"Authorization": "Bearer pdr_sk_..."} # key with Usage & workers
def wait_until_stopped(model: str, cancelled_at: datetime, timeout_s: float = 30) -> bool:
deadline = time.monotonic() + timeout_s
while time.monotonic() < deadline:
running = requests.get(f"{API}/usage/inflight", headers=HEADERS, timeout=10).json()
still_there = [
r for r in running
if r["model"] == model
and datetime.fromisoformat(r["started_at"]) <= cancelled_at
]
if not still_there:
return True
time.sleep(1)
return False
cancelled_at = datetime.now(timezone.utc)
# ... close the connection for the request you're abandoning ...
if wait_until_stopped("qwen3.6:27b", cancelled_at):
... # safe to retry
Once the request finishes stopping, it appears in the
usage log with status 499 and the error
message "Client disconnected".
Usage log
GET /api/v1/usage/logs returns completed requests, newest
first: successes, failures and cancellations. Page through it with
limit (1–1000, default 50) and offset, and add
worker_id to see one worker's requests.
curl "https://api.pendra.ai/api/v1/usage/logs?limit=20" \
-H "Authorization: Bearer pdr_sk_..."
[
{
"id": "0b8e3c3e-5d0a-4a8f-9d2f-6c1f0e2b7a41",
"created_at": "2026-09-25T10:14:07Z",
"model": "qwen3.6:27b",
"request_type": "chat",
"status_code": 499,
"error_message": "Client disconnected",
"prompt_tokens": 512,
"completion_tokens": 118,
"total_tokens": 630,
"request_duration_ms": 4380,
"time_to_first_token_ms": 310,
"tokens_per_second": 31.4,
"worker_id": "wrk-3f9b8a2d9e104c1a",
"worker_name": "gpu-box-1"
}
]
Each row carries more fields than shown here, such as reasoning and cached
token counts, queue wait, and cold-start timings. Every field a request
doesn't apply to is null or 0.
A request that used
speculative decoding also
carries draft_acceptance_rate: the share of draft tokens the
model kept, from 0 to 1. It is null on every other row.
Token counts come from the worker when the request finishes. If a streaming
request to /api/v1/chat/completions ends before the worker can
report them (you cancel it, or the worker fails partway through), the row
carries an estimate instead, worked out from the text that was sent and
streamed back (about 4 characters per token). The estimate usually reads
lower than the true count, because images and tool definitions aren't
counted. An encrypted request's content can't be read, so no estimate is
made: its counts stay 0 unless the worker reported them.
A transcription row (request_type of audio) has no
token counts. It carries audio_seconds, the length of the audio,
and audio_format, which is wav, mp3 or
flac. Your worker reads the format from the audio itself, not
from the file name, and reports it the same way for encrypted requests. It is
null when the request failed. A live transcription row has
audio/pcm;stream instead.
Totals
Three endpoints add the log up for you. Each takes days
(1–365, default 30):
GET /api/v1/usage/summarygives total requests, tokens and images.GET /api/v1/usage/dailygives the same totals for each day.GET /api/v1/usage/breakdowngives totals by model and by worker (top 10 of each). Passsort=requeststo rank by request count instead of tokens.
Worker status
GET /api/v1/workers lists your organisation's workers,
connected and recently disconnected.
curl https://api.pendra.ai/api/v1/workers \
-H "Authorization: Bearer pdr_sk_..."
{
"workers": [
{
"worker_id": "wrk-3f9b8a2d9e104c1a",
"worker_name": "gpu-box-1",
"status": "healthy",
"version": "3.104.0",
"models": [{ "id": "qwen3.6:27b" }],
"active_tasks": 1,
"queue_depth": 0,
"max_concurrent": 4,
"last_seen_seconds_ago": 2
}
]
}
status is one of:
healthy: connected and serving.stale: connected, but it hasn't checked in recently.unhealthy: connected, but its last requests failed, so Pendra is holding traffic back from it for a short while.offline: disconnected.offline_sincesays when.
active_tasks is how many requests the worker is running now,
and queue_depth is how many are waiting for it. Each entry
also includes the rest of what the console's Workers page shows, such as
GPU details and whether the worker is waiting for
approval.
Worker activity
GET /api/v1/workers/events returns the activity history shown
on a worker's Activity tab, newest first: connects and
disconnects, health changes, model loads and installs, and setting changes.
Filter with worker_id and event_type (for example
disconnected), and page with limit and
offset.
curl "https://api.pendra.ai/api/v1/workers/events?worker_id=wrk-3f9b8a2d9e104c1a&limit=20" \
-H "Authorization: Bearer pdr_sk_..."
[
{
"id": "5a1c2d3e-4f50-4617-8293-a4b5c6d7e8f9",
"worker_id": "wrk-3f9b8a2d9e104c1a",
"worker_name": "gpu-box-1",
"event_type": "disconnected",
"detail": null,
"metadata": {},
"created_at": "2026-09-25T09:02:44Z",
"transient": false,
"actor": null
}
]
transient is true when a disconnect was a brief
drop that the worker recovered from on its own. Ignore those if you only
want to alert on real outages.
Power usage
GET /api/v1/workers/{worker_id}/power returns how much power
a worker's machine drew over a period: the history behind the
Power usage chart on the worker's page.
Pass range as 1h, 24h, 7d
or 30d. Each point averages one minute on 1h, five
minutes on 24h, and one hour on 7d and
30d (resolution_seconds says which).
curl "https://api.pendra.ai/api/v1/workers/wrk-3f9b8a2d9e104c1a/power?range=24h" \
-H "Authorization: Bearer pdr_sk_..."
{
"worker_id": "wrk-3f9b8a2d9e104c1a",
"range": "24h",
"resolution_seconds": 300,
"machine": "estimated",
"channels": [
{ "key": "gpu:GPU-3f1c2a4b", "scope": "gpu", "index": 0, "name": "NVIDIA GeForce RTX 3090",
"source": "nvml", "method": "measured", "power_limit_w": 350 },
{ "key": "cpu", "scope": "cpu", "source": "rapl", "method": "measured" },
{ "key": "system", "scope": "system", "source": "estimate", "method": "estimated" }
],
"points": [
{ "t": "2026-09-25T10:40:00Z",
"values": { "gpu:GPU-3f1c2a4b": 197.0, "cpu": 61.0, "system": 312.3 } }
],
"totals": { "energy_kwh": 4.65, "avg_w": 194.0, "peak_w": 588.0, "coverage": 0.998 }
}
channelslists what the machine reports: one per GPU, the CPU, andsystem, the whole machine. A machine whose sensor covers its whole processor chip reports asocchannel in place ofcpu. A channel read from one of the machine's own power sensors carries that sensor'sname.methodismeasuredwhen a part was read from a sensor andestimatedwhen it was modelled.machineis the method of the whole-machine figure, ornullwhen there are no readings. See Measured or Estimated.- Each point's
valuesare average watts for that window, keyed by channel. A channel missing from a point had no reading then; treat it as a gap, not as zero. Windows with no readings at all are left out. totalscover the whole machine over the range: energy in kWh, average and peak watts, andcoverage, the share of the range that has readings (0 to 1).
A worker with no power readings, such as one too old to report them, returns
empty channels and points. A worker that isn't in
your organisation returns 404. Power history is part of the Pro
and Enterprise plans: on the Free plan this endpoint returns 403
with the error type plan_required. Each worker in the
worker list also carries power, its live
reading (null when it isn't reporting), on every plan.
GET /api/v1/workers/latest-version returns the newest worker
release, as {"version": "3.104.0"}. Compare it with each
worker's version to find the ones that need an upgrade.