Endpoints
GPU rental
On-demand GPU rental lets you provision a GPU without going near the console: rent one, it joins your organisation as a worker, you run inference on it exactly like any other Pendra worker, then you stop it. You're billed per GPU-hour for as long as it's running. Everything here is also available as a click-through flow in the console — see Rent a GPU for the equivalent UI walkthrough, jurisdictions, and how billing and spend limits work.
Permissions
Renting needs a key with the GPU rental permission ticked — it's off by default because it spends money. Tick it in the console when you create the key, or replace an existing key with one that has it. A key created before permissions existed is unrestricted and can already rent. Full details, including what a missing permission looks like, are on the authentication page.
Renting and stopping also require the key's owner to be an owner
or operator in the organisation — a key belonging to a
plain member gets a 403 on both, even with the
right permission ticked. Reading the catalogue, listing instances, and
checking spend have no such restriction.
What you can rent
GET /api/v1/compute/catalogue lists the GPUs available to your
organisation, one entry per GPU model, each with a price per jurisdiction
tier. Pendra chooses which underlying provider and region actually serve
the instance — you choose the GPU and how strict a jurisdiction guarantee
you need.
curl https://api.pendra.ai/api/v1/compute/catalogue \
-H "Authorization: Bearer pdr_sk_..."
[
{
"gpu_model_id": "gpu_l4_24",
"label": "NVIDIA L4 24GB",
"vendor": "NVIDIA",
"model": "L4",
"vram_gb": 24,
"generation": "Ada Lovelace",
"gpu_count": 1,
"tiers": [
{
"tier_id": "uk_region",
"label": "UK Hosted",
"rank": 3,
"description": "UK data centre, but the operator may be foreign-owned.",
"price_per_hour_gbp": 1.15,
"allowed": true,
"available": true
},
{
"tier_id": "global",
"label": "Global",
"rank": 1,
"description": "Runs wherever there is capacity — the cheapest option, with no residency guarantee.",
"price_per_hour_gbp": 1.05,
"allowed": true,
"available": true
}
]
}
]
A tier only appears on a GPU if Pendra can actually place you there right
now. allowed is false when your organisation's
jurisdiction floor rules it out; available is a soft,
best-effort signal that can flip to false when that tier is
briefly out of capacity — still worth showing, not worth hiding.
A tier may also carry max_ram_gb: the most system memory
(GB) any machine serving it has. It is null when Pendra
can't say, which never rules the GPU out.
Picking a model to pre-install
GET /api/v1/compute/models lists the catalogue models your
organisation can rent a GPU for, one entry per size. This is optional —
you can rent a bare GPU and install models yourself afterwards — but
passing a variant's variant_id in the models
array on POST /instances (below) pre-installs it during boot,
so the instance is ready to serve the moment it connects.
curl https://api.pendra.ai/api/v1/compute/models \
-H "Authorization: Bearer pdr_sk_..."
[
{
"model_id": "qwen3.5:4B",
"name": "Qwen 3.5 · 4B",
"publisher": "Alibaba",
"parameter_size": "4B",
"default_variant_id": "qwen3.5:4b",
"variants": [
{
"variant_id": "qwen3.5:4b",
"label": "Q4_K_M",
"parameter_size": "4B",
"quantization": "Q4_K_M",
"min_vram_gb": 6
},
{
"variant_id": "qwen3.5:4b-q6_k",
"label": "Q6_K",
"parameter_size": "4B",
"quantization": "Q6_K",
"min_vram_gb": 7
},
{
"variant_id": "qwen3.5:4b-q8_0",
"label": "Q8_0",
"parameter_size": "4B",
"quantization": "Q8_0",
"min_vram_gb": 8
}
]
}
]
min_vram_gb is the minimum single-GPU memory to run that
variant fully on-GPU — use it to sanity-check a variant against the
vram_gb on the GPU you're about to rent. Some large models
keep part of themselves in the machine's system memory instead of on the
GPU; for those a variant also carries min_ram_gb, the system
memory (GB) the machine needs. Compare it with the tier's
max_ram_gb. It's absent or null when the model
has no such requirement.
Renting an instance
POST /api/v1/compute/instances provisions the GPU.
gpu_model_id and tier_id are required;
models is an optional array of catalogue variant ids to
pre-install (each is set to Always-on
once it's installed, so it stays loaded), and name is an optional friendly label for the
worker (defaults to "Rented <GPU>").
curl https://api.pendra.ai/api/v1/compute/instances \
-H "Authorization: Bearer pdr_sk_..." \
-H "Content-Type: application/json" \
-d '{
"gpu_model_id": "gpu_l4_24",
"tier_id": "uk_region",
"models": ["qwen3.5:4b"],
"name": "batch-job-worker"
}'
Returns 202 Accepted with the new rental:
{
"id": "5c9b6b0a-6e2b-4b8a-9c1a-3f9b8a2d9e10",
"gpu_model_id": "gpu_l4_24",
"label": "NVIDIA L4 24GB",
"requested_tier_id": "uk_region",
"tier_label": "UK Hosted",
"status": "provisioning",
"worker_id": "wrk-3f9b8a2d9e104c1a",
"models": ["qwen3.5:4b"],
"price_per_hour_gbp": 1.15,
"accrued_cost_gbp": 0,
"hours_billed": 0,
"connected": false,
"started_at": null,
"stopped_at": null,
"created_at": "2026-09-09T14:02:11Z",
"error": null,
"boot_phase": null,
"model_download": null
}
status always starts at provisioning on a
successful call — the GPU is real hardware booting up, not something
handed to you instantly. Nothing is billed yet; the clock only starts once
the worker actually connects (see polling below).
A rental request can also fail before anything is provisioned:
400 if no machine with that GPU in that tier has enough
system memory for the models you asked to pre-install (choose a larger
GPU); 402 if your organisation has no payment card on file, not
enough credit to cover the new GPU's first hour plus the next hour of its
running instances, or has hit its monthly budget cap; 409 if
there's no capacity for that GPU in the tier you asked for right now (try
again shortly, or choose another GPU or tier). The card, credit balance,
and budget cap are all managed from the console — see
Credits & payment.
Waiting for it to come online
Poll GET /api/v1/compute/instances/{id} until
status reaches running and connected
is true — that's usually a few minutes. While it's still
booting, boot_phase reports the last milestone reached
(e.g. docker_ready, pulling_image,
worker_started) so you can show real progress instead of a
bare "provisioning" spinner. Polling is only for your benefit: the
instance comes online, and starts billing, as soon as its worker connects
whether or not anything is watching.
curl https://api.pendra.ai/api/v1/compute/instances/5c9b6b0a-6e2b-4b8a-9c1a-3f9b8a2d9e10 \
-H "Authorization: Bearer pdr_sk_..."
If you passed models when renting, a freshly-connected
instance may still be downloading them — check for that with the
model_download field, which is non-null (with
total, completed, and a progress_pct)
only while at least one install is still in flight, and drops back to
null once every model is ready to serve. Once
model_download is null and status is
running, send inference to it exactly like any other worker —
it's just another entry behind
chat completions, picked
automatically once it's serving the model you ask for.
Stopping an instance
DELETE /api/v1/compute/instances/{id} destroys the
underlying GPU and disconnects the worker.
Billing stops the moment this call succeeds — an instance
you leave running keeps costing you money until you stop it, whether or
not you're actively sending it requests. A rental can't be restarted once
stopped; renting is always start-fresh.
curl -X DELETE https://api.pendra.ai/api/v1/compute/instances/5c9b6b0a-6e2b-4b8a-9c1a-3f9b8a2d9e10 \
-H "Authorization: Bearer pdr_sk_..."
Returns the rental with status: "stopped" and stopped_at set.
Evidence of where it ran
GET /api/v1/compute/instances/{id}/audit returns the
placement record for a rental: which jurisdiction tier was actually
delivered, the operator's legal entity and domicile, and the data-centre
region — everything you need to evidence a residency claim to an auditor.
It 404s until the instance has actually been placed.
curl https://api.pendra.ai/api/v1/compute/instances/5c9b6b0a-6e2b-4b8a-9c1a-3f9b8a2d9e10/audit \
-H "Authorization: Bearer pdr_sk_..."
Checking your spend
GET /api/v1/compute/spend returns this calendar month's GPU
bill, broken down per instance.
curl https://api.pendra.ai/api/v1/compute/spend \
-H "Authorization: Bearer pdr_sk_..."
{
"period_start": "2026-09-01T00:00:00Z",
"month_to_date_gbp": 18.40,
"monthly_budget_gbp": 200.0,
"active_count": 1,
"lines": [
{
"rental_id": "5c9b6b0a-6e2b-4b8a-9c1a-3f9b8a2d9e10",
"gpu_model_id": "gpu_l4_24",
"label": "NVIDIA L4 24GB",
"tier_label": "UK Hosted",
"status": "running",
"price_per_hour_gbp": 1.15,
"hours": 16,
"cost_gbp": 18.40,
"started_at": "2026-09-08T22:04:02Z",
"stopped_at": null
}
]
}
Rate limits
Renting is limited to one call to POST /instances per
API key per minute. Over that, you get a 429 with a
Retry-After header (seconds to wait).
This limit counts the call itself, not just successful rentals — so a
request that's rejected for another reason (an unrecognised model, a
declined card, or no matching capacity) still counts against your minute.
If you get a 400, 402, or 409, fix
the underlying issue before retrying rather than immediately resubmitting,
or you'll also collect a 429 and have to wait out the minute.
429 here does not mean your rental failed.
It means Pendra refused to even attempt a second provision within
the same minute — the first request may well have succeeded and already be
provisioning a GPU. Before you retry, call
GET /api/v1/compute/instances and check whether the earlier
request went through. Retrying blindly on a 429 is how you
end up renting — and paying for — two GPUs instead of one.
curl https://api.pendra.ai/api/v1/compute/instances \
-H "Authorization: Bearer pdr_sk_..."
Budgets and auto-stop
Monthly spending caps and the per-instance max-runtime auto-stop policy are configured in the console only, under Settings — there is no API to read or change them. See Spending limits for how they work.