Endpoints

Usage & workers

The usage and worker endpoints give your own code the same view the console's Usage and Workers pages show: which requests are running right now, what finished and how, and whether your workers are healthy. Typical uses are a monitoring job that alerts when a worker goes offline, or a service that cancels a request and wants to confirm it has stopped before retrying it.

Permissions

These endpoints need a key with the Usage & workers permission. It's off by default, because most keys only send requests and never need to see your organisation's history. Tick it in the console when you create the key. It's read-only. To change worker settings from your code, see Managing workers.

A key without the permission gets a 403 naming what it's missing: usage:read for the /usage endpoints, workers:read for the /workers endpoints. Keys created before permissions existed are unrestricted and can already call them. See the authentication page for how permissions work.

A key only ever sees its own organisation's requests and workers, so there is no X-Org-Id header to send.

Requests in progress

GET /api/v1/usage/inflight lists the requests running on your workers right now, oldest first. Add ?worker_id=… to narrow it to one worker.

curl
curl https://api.pendra.ai/api/v1/usage/inflight \
  -H "Authorization: Bearer pdr_sk_..."
[
  {
    "task_id": "8f14e45f-ceea-467a-9575-2f9d6b1c2a10",
    "model": "qwen3.6:27b",
    "worker_id": "wrk-3f9b8a2d9e104c1a",
    "worker_name": "gpu-box-1",
    "request_type": "chat",
    "started_at": "2026-09-25T10:14:03Z",
    "elapsed_ms": 4210,
    "tokens_generated": 118
  }
]

tokens_generated is a running estimate for streaming /api/v1/chat/completions requests. It is null until the first tokens arrive, and for other kinds of request. For a privately encrypted request Pendra can't read the output, so the estimate counts the encrypted chunks streamed so far instead.

A request stays on this list for as long as it's running, however long that is. It drops off when it finishes, is cancelled or fails.

Confirming a cancelled request has stopped

You cancel a request by closing the connection. Pendra stops the generation and frees the worker's slot straight away, so the request drops off the in-progress list. To confirm that before you retry, record the time you cancelled, then poll /usage/inflight until no request for that model that started before then is still listed:

import time
from datetime import datetime, timezone

import requests

API = "https://api.pendra.ai/api/v1"
HEADERS = {"Authorization": "Bearer pdr_sk_..."}  # key with Usage & workers


def wait_until_stopped(model: str, cancelled_at: datetime, timeout_s: float = 30) -> bool:
    deadline = time.monotonic() + timeout_s
    while time.monotonic() < deadline:
        running = requests.get(f"{API}/usage/inflight", headers=HEADERS, timeout=10).json()
        still_there = [
            r for r in running
            if r["model"] == model
            and datetime.fromisoformat(r["started_at"]) <= cancelled_at
        ]
        if not still_there:
            return True
        time.sleep(1)
    return False


cancelled_at = datetime.now(timezone.utc)
# ... close the connection for the request you're abandoning ...
if wait_until_stopped("qwen3.6:27b", cancelled_at):
    ...  # safe to retry

Once the request finishes stopping, it appears in the usage log with status 499 and the error message "Client disconnected".

Usage log

GET /api/v1/usage/logs returns completed requests, newest first: successes, failures and cancellations. Page through it with limit (1–1000, default 50) and offset, and add worker_id to see one worker's requests.

curl
curl "https://api.pendra.ai/api/v1/usage/logs?limit=20" \
  -H "Authorization: Bearer pdr_sk_..."
[
  {
    "id": "0b8e3c3e-5d0a-4a8f-9d2f-6c1f0e2b7a41",
    "created_at": "2026-09-25T10:14:07Z",
    "model": "qwen3.6:27b",
    "request_type": "chat",
    "status_code": 499,
    "error_message": "Client disconnected",
    "prompt_tokens": 512,
    "completion_tokens": 118,
    "total_tokens": 630,
    "request_duration_ms": 4380,
    "time_to_first_token_ms": 310,
    "tokens_per_second": 31.4,
    "worker_id": "wrk-3f9b8a2d9e104c1a",
    "worker_name": "gpu-box-1"
  }
]

Each row carries more fields than shown here, such as reasoning and cached token counts, queue wait, and cold-start timings. Every field a request doesn't apply to is null or 0.

A request that used speculative decoding also carries draft_acceptance_rate: the share of draft tokens the model kept, from 0 to 1. It is null on every other row.

Token counts come from the worker when the request finishes. If a streaming request to /api/v1/chat/completions ends before the worker can report them (you cancel it, or the worker fails partway through), the row carries an estimate instead, worked out from the text that was sent and streamed back (about 4 characters per token). The estimate usually reads lower than the true count, because images and tool definitions aren't counted. An encrypted request's content can't be read, so no estimate is made: its counts stay 0 unless the worker reported them.

A transcription row (request_type of audio) has no token counts. It carries audio_seconds, the length of the audio, and audio_format, which is wav, mp3 or flac. Your worker reads the format from the audio itself, not from the file name, and reports it the same way for encrypted requests. It is null when the request failed. A live transcription row has audio/pcm;stream instead.

Totals

Three endpoints add the log up for you. Each takes days (1–365, default 30):

  • GET /api/v1/usage/summary gives total requests, tokens and images.
  • GET /api/v1/usage/daily gives the same totals for each day.
  • GET /api/v1/usage/breakdown gives totals by model and by worker (top 10 of each). Pass sort=requests to rank by request count instead of tokens.

Worker status

GET /api/v1/workers lists your organisation's workers, connected and recently disconnected.

curl
curl https://api.pendra.ai/api/v1/workers \
  -H "Authorization: Bearer pdr_sk_..."
{
  "workers": [
    {
      "worker_id": "wrk-3f9b8a2d9e104c1a",
      "worker_name": "gpu-box-1",
      "status": "healthy",
      "version": "3.104.0",
      "models": [{ "id": "qwen3.6:27b" }],
      "active_tasks": 1,
      "queue_depth": 0,
      "max_concurrent": 4,
      "last_seen_seconds_ago": 2
    }
  ]
}

status is one of:

  • healthy: connected and serving.
  • stale: connected, but it hasn't checked in recently.
  • unhealthy: connected, but its last requests failed, so Pendra is holding traffic back from it for a short while.
  • offline: disconnected. offline_since says when.

active_tasks is how many requests the worker is running now, and queue_depth is how many are waiting for it. Each entry also includes the rest of what the console's Workers page shows, such as GPU details and whether the worker is waiting for approval.

Worker activity

GET /api/v1/workers/events returns the activity history shown on a worker's Activity tab, newest first: connects and disconnects, health changes, model loads and installs, and setting changes. Filter with worker_id and event_type (for example disconnected), and page with limit and offset.

curl
curl "https://api.pendra.ai/api/v1/workers/events?worker_id=wrk-3f9b8a2d9e104c1a&limit=20" \
  -H "Authorization: Bearer pdr_sk_..."
[
  {
    "id": "5a1c2d3e-4f50-4617-8293-a4b5c6d7e8f9",
    "worker_id": "wrk-3f9b8a2d9e104c1a",
    "worker_name": "gpu-box-1",
    "event_type": "disconnected",
    "detail": null,
    "metadata": {},
    "created_at": "2026-09-25T09:02:44Z",
    "transient": false,
    "actor": null
  }
]

transient is true when a disconnect was a brief drop that the worker recovered from on its own. Ignore those if you only want to alert on real outages.

Power usage

GET /api/v1/workers/{worker_id}/power returns how much power a worker's machine drew over a period: the history behind the Power usage chart on the worker's page. Pass range as 1h, 24h, 7d or 30d. Each point averages one minute on 1h, five minutes on 24h, and one hour on 7d and 30d (resolution_seconds says which).

curl
curl "https://api.pendra.ai/api/v1/workers/wrk-3f9b8a2d9e104c1a/power?range=24h" \
  -H "Authorization: Bearer pdr_sk_..."
{
  "worker_id": "wrk-3f9b8a2d9e104c1a",
  "range": "24h",
  "resolution_seconds": 300,
  "machine": "estimated",
  "channels": [
    { "key": "gpu:GPU-3f1c2a4b", "scope": "gpu", "index": 0, "name": "NVIDIA GeForce RTX 3090",
      "source": "nvml", "method": "measured", "power_limit_w": 350 },
    { "key": "cpu", "scope": "cpu", "source": "rapl", "method": "measured" },
    { "key": "system", "scope": "system", "source": "estimate", "method": "estimated" }
  ],
  "points": [
    { "t": "2026-09-25T10:40:00Z",
      "values": { "gpu:GPU-3f1c2a4b": 197.0, "cpu": 61.0, "system": 312.3 } }
  ],
  "totals": { "energy_kwh": 4.65, "avg_w": 194.0, "peak_w": 588.0, "coverage": 0.998 }
}
  • channels lists what the machine reports: one per GPU, the CPU, and system, the whole machine. A machine whose sensor covers its whole processor chip reports a soc channel in place of cpu. A channel read from one of the machine's own power sensors carries that sensor's name.
  • method is measured when a part was read from a sensor and estimated when it was modelled. machine is the method of the whole-machine figure, or null when there are no readings. See Measured or Estimated.
  • Each point's values are average watts for that window, keyed by channel. A channel missing from a point had no reading then; treat it as a gap, not as zero. Windows with no readings at all are left out.
  • totals cover the whole machine over the range: energy in kWh, average and peak watts, and coverage, the share of the range that has readings (0 to 1).

A worker with no power readings, such as one too old to report them, returns empty channels and points. A worker that isn't in your organisation returns 404. Power history is part of the Pro and Enterprise plans: on the Free plan this endpoint returns 403 with the error type plan_required. Each worker in the worker list also carries power, its live reading (null when it isn't reporting), on every plan.

GET /api/v1/workers/latest-version returns the newest worker release, as {"version": "3.104.0"}. Compare it with each worker's version to find the ones that need an upgrade.