Endpoints
Managing workers
The worker management endpoints let your own code change the settings you'd otherwise change on a worker's page in the console. For example, a deployment script can pin a model's context size, install the models a worker should serve, or raise how many requests a worker accepts once it's on bigger hardware. The change reaches the running worker within a few seconds, and the worker doesn't need a restart.
Permissions
These endpoints need a key with the Manage workers permission. It's off by default. Tick it in the console when you create the key. A key that holds it can also read worker status and the worker endpoints that need Usage & workers.
A key acts with the role of the person who created it. Only an
organisation owner or operator can change
worker settings, so a key created by a member gets a 403 even
with the permission, and a key stops working here if its creator's role
changes to member.
Some actions stay in the console, whatever permissions a key has:
- Approving a worker, because an approved worker sees your organisation's requests.
- Removing a worker.
- Turning on Allow remote image URLs and Web tools, which let the worker fetch pages from the internet. The console shows what each one changes before you switch it on.
A key without the permission gets a 403 naming it:
{"detail": "API key is missing the required scope: workers:write"}
Keys created before permissions existed are unrestricted and can already
call these endpoints. See the
authentication page
for how permissions work. A key only reaches its own organisation's
workers, so there's no X-Org-Id header to send.
Example: pin a model's context size
Send the worker id (from GET /api/v1/workers)
and the model id in the path, and the context size in tokens in the body.
Send 0 to go back to Auto.
curl -X PATCH https://api.pendra.ai/api/v1/workers/wrk-3f9b8a2d9e104c1a/models/qwen3.6:27b/config \
-H "Authorization: Bearer pdr_sk_..." \
-H "Content-Type: application/json" \
-d '{"context_size": 32768}'
HTTP/1.1 202 Accepted
{"status": "queued", "model_id": "qwen3.6:27b", "context_size": 32768}
The same limits apply as in the console. The largest value is
1048576, and the worker serves the largest window that fits the
model's trained context and the worker's memory. The model picks up the new
size on its next request. See
Choosing a context size for when to
raise or lower it.
Confirming a change has applied
A 202 means Pendra has passed the change to the worker, not
that the worker has finished applying it. The worker confirms each change a
moment later, and the new value then shows up in
GET /api/v1/workers. For a
context size, look for the model under context_size_override:
import os, time, requests
API = "https://api.pendra.ai/api/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['PENDRA_API_KEY']}"}
WORKER, MODEL = "wrk-3f9b8a2d9e104c1a", "qwen3.6:27b"
# Ask the worker to serve this model with a 32k context window.
requests.patch(
f"{API}/workers/{WORKER}/models/{MODEL}/config",
headers=HEADERS,
json={"context_size": 32768},
).raise_for_status()
# Wait for the worker to confirm it.
for _ in range(10):
workers = requests.get(f"{API}/workers", headers=HEADERS).json()["workers"]
worker = next(w for w in workers if w["worker_id"] == WORKER)
if (worker.get("context_size_override") or {}).get(MODEL) == 32768:
break
time.sleep(1)
Changes are only sent to a connected worker. If the worker is offline the
request gets a 503 and nothing changes, so retry once it's back.
Endpoints
All paths are under https://api.pendra.ai/api/v1. Every
change returns 202 once it has been sent to the worker.
Per model
| Setting | Request | Body |
|---|---|---|
| Context size | PATCH /workers/{worker_id}/models/{model_id}/config | {"context_size": 32768}. 0 is Auto. |
| Serving | POST /workers/{worker_id}/models/{model_id}/serving | {"enabled": false} stops this worker serving the model without uninstalling it. |
| Always-on | POST /workers/{worker_id}/models/{model_id}/loadPOST /workers/{worker_id}/models/{model_id}/unload | None. load loads the model and makes it Always-on. unload releases it from memory. It stays installed and loads again on its next request. |
| Idle unload | PATCH /workers/{worker_id}/models/{model_id}/idle-timeout | {"idle_timeout_minutes": 60}, from 1 to 1440. 0 keeps the model loaded until the memory is needed for another model. |
Installing and removing models
| Action | Request | Body |
|---|---|---|
| Install | POST /workers/{worker_id}/models | {"model_id": "qwen3.6:27b"}, a model id from the catalogue. Returns a job_id. |
| Uninstall | DELETE /workers/{worker_id}/models/{model_id} | None. Returns a job_id. |
| Check progress | GET /workers/{worker_id}/model-jobs | None. Lists install and uninstall jobs, newest first, with their status and any error. Filter with ?status=active. Needs only Usage & workers. |
| Cancel | POST /workers/{worker_id}/model-jobs/{job_id}/cancel | None. Stops an install that's still running. |
| Resume | POST /workers/{worker_id}/model-jobs/{job_id}/resume | None. Continues an interrupted or failed install from where the download stopped. |
A download can take several minutes for a large model. Poll
model-jobs until the job's status is
success or error.
Worker-wide
These match the settings on a worker's Settings tab in the console. See worker configuration for what each one does and its default.
| Setting | Request | Body |
|---|---|---|
| Name | PATCH /workers/{worker_id}/name | {"worker_name": "gpu-box-1"} |
| Requests accepted at once | PATCH /workers/{worker_id}/max-concurrent | {"max_concurrent": 8}, from 1 to 1024 |
| Run requests together | PATCH /workers/{worker_id}/batching | {"enabled": true, "max_running": 4}. Send either field on its own to leave the other unchanged. max_running: 0 runs every accepted request at once. |
| Task timeout | PATCH /workers/{worker_id}/task-timeout | {"timeout_seconds": 1800}, from 60 to 7200 |
| Auto-unload idle models | PATCH /workers/{worker_id}/idle-default | {"idle_timeout_minutes": 15}, from 1 to 1440. 0 turns it off. |
| KV cache (Memory) | PATCH /workers/{worker_id}/kv-prefix-cache | {"enabled": true} |
| KV cache (NVMe) | PATCH /workers/{worker_id}/kv-disk-cache | {"enabled": true} |
| NVMe cache size cap | PATCH /workers/{worker_id}/kv-disk-cache-max-gb | {"max_gb": 200}. 0 lets the worker size it. |
| RAM cache | PATCH /workers/{worker_id}/kv-ram-cache | {"percent": 50}: 0 (off) or 10 to 90. null is Automatic. percent is required. |
| Conversation memory density | PATCH /workers/{worker_id}/kv-cache-precision | {"precision": "q8_0"} for Compact, "f16" for Standard |
Errors
400: a value is out of range. Thedetailgives the allowed range.403: the key is missing the Manage workers permission, its creator isn't an owner or operator, or the worker is one Pendra operates for you, whose settings only Pendra can change.404: no worker with that id in your organisation.409: the model isn't installed on that worker.503: the worker isn't connected right now. Nothing was changed.