On-demand GPUs: rent a card by the hour, in the jurisdiction you choose
Until today, putting a model on Pendra meant bringing your own hardware. You owned the GPU, installed the worker, and Pendra turned it into an API. That is still the heart of the platform, but plenty of teams told us the same thing. They want the sovereignty without the shopping. So now you can rent the GPU too. Pick a card and a jurisdiction in the console, and a few minutes later a fully configured worker joins your organisation, billed by the hour, stoppable at any time.
From click to connected
Open Workers, click Add a Worker and choose On-demand GPU. If you already know which model you want to serve, pick it and a precision, and the console narrows the list to GPUs that can actually run it, suggests one, and installs the model for you automatically once the instance is up. Leave it blank to browse every card and install models yourself later.
The instance appears in your Workers list straight away as a loading row and walks through its startup stages (finding a GPU, booting, installing the worker, downloading your model). Once it connects, it looks and behaves like any worker you installed by hand.
Billing is by the hour, at the rate shown when you rent, and the clock starts when the worker connects rather than when you click Rent, so the minutes an instance spends booting cost nothing. Part-hours round up, so a 90-minute run bills as two. When you are done, press Stop on the worker's page and billing ends.
Choose where it may run
The rent screen never asks you to pick a cloud provider. What you choose is the jurisdiction the GPU must stay within, and Pendra finds capacity inside it. The plan is four tiers, so every workload can run where its owner is comfortable. UK Sovereign means a UK data centre operated by a UK-domiciled company, so neither the hardware nor the operator answers to a foreign jurisdiction. UK Hosted keeps your data in a UK data centre but allows a foreign-owned operator. Europe Hosted widens that to the EEA and Switzerland, and Global runs wherever there is capacity.
We are starting at the strict end. Everything you can rent today is UK Sovereign, the hardest tier to offer and the reason most of our customers are here. The wider tiers open up as we bring capacity online.
The promise underneath is blunt. If we cannot place your workload inside the jurisdiction you chose, the rental fails and says so. It never quietly runs somewhere else. Every rental also keeps a record of where it actually ran, covering the operator, its legal entity and country of domicile, and the data-centre region, against the exact version of the jurisdiction rules in force at the time. If an auditor asks where your inference happened, you have an answer you can show them.
Indicative pricing
Five A100 configurations to start, every one of them UK Sovereign. Prices are in pounds per hour and correct on 14 August 2026.
| GPU | Memory | Per hour |
|---|---|---|
| A100 | 40 GB | £1.29 |
| 2× A100 | 2× 40 GB | £2.58 |
| A100 | 80 GB | £2.12 |
| 2× A100 | 2× 80 GB | £4.24 |
| 4× A100 | 4× 80 GB | £8.48 |
Each card shows its exact price before you rent, and that rate holds for as long as the instance runs.
Rentals run on prepaid credit. An organisation owner adds a card under Settings and tops up the wallet, £20 minimum, and hours are settled from the balance as they are used. When you upgrade to Pro you get a one-off £99 of GPU credit to experiment with, roughly a month of Pro back, which lasts 90 days. Credit you buy yourself never expires.
The balance is also the hard stop. An instance runs until you stop it or the credit runs out, and owners can add two softer guards, a monthly budget and a maximum runtime per instance, for the afternoon someone forgets to press Stop. Auto top-up keeps a long-running instance fed, and however a rental ends, your owners get an email saying why.
A worker like any other
A rented GPU arrives as a normal worker. It has a worker page, you install and remove models from the catalogue, its requests show up in Usage, and your existing API keys route to it with no code changes. Rented workers also sit outside your plan's worker limit, so a full fleet of self-hosted machines does not stop you renting a burst of extra capacity for an evaluation or a launch week.
If you are wondering what fits on what, our guide to GPU memory and model sizes applies to rented cards exactly as it does to owned ones.
On-demand GPUs are live in the console today. Open Workers, choose Add a Worker, and pick On-demand GPU. The full detail, covering jurisdictions, credits and spending limits, is in the docs. If you need a GPU or a jurisdiction you don't see, tell us.