Worker

Provisioning with Ansible

If you run more than one GPU box, you don't want to click through setup on each one. This page gives you a single Ansible playbook that provisions a Linux worker end to end: it installs the Pendra worker, writes its worker key, starts the service, and installs the models you list — the same steps the manual install walks through, but hands-off and repeatable across every host in your inventory.

It targets native package installs on Linux (the .deb / .rpm served from the signed Pendra repository). Re-running it is safe: each step only changes what's out of date, so the same playbook both provisions a new host and reconciles an existing one.

Worker keys. Generate a key in the console (Workers → Worker Keys → Generate key) — the base64 string it shows you is an Ed25519 private key that authenticates your workers to your organisation. You can give every host its own key or share one across the inventory: each host becomes its own worker either way (see Worker IDs). If you build hosts from a golden image, don't bake a worker's .pendra-instance file into it — hosts cloned with that file look like the same worker. It is not a pdr_sk_ API key: worker keys authenticate machines, API keys authenticate clients.

Before you start

  • A control machine with Ansible installed and SSH access to each worker host.
  • One worker key generated per host in the console.
  • Linux hosts on glibc 2.38 or newer (Ubuntu 24.04+, Debian 13+, or equivalent) — the inference runtime needs it.
NVIDIA hosts (pendra-cuda) need the CUDA 13 runtime first. The pendra-cuda package requires NVIDIA driver R580 or newer plus the CUDA 13 runtime (libcudart.so.13 / libcublas.so.13), which the package doesn't include. Install it before this playbook runs — a fresh install of pendra-cuda refuses rather than come up CPU-only. See Install the worker for the runtime commands. Hosts using the CPU (pendra) or Vulkan (pendra-vulkan) packages have no such prerequisite.

1. Inventory and variables

List your workers and set a few group-wide variables. The GPU variant and the model list live in group_vars; the worker key is a secret, so keep it per host in an encrypted vault file (below).

inventory.ini
# inventory.ini
[pendra_workers]
gpu-01.example.com
gpu-02.example.com

[pendra_workers:vars]
ansible_user=ubuntu
group_vars/pendra_workers.yml
# group_vars/pendra_workers.yml
pendra_gpu_variant: cuda            # cpu | cuda | vulkan
pendra_ws_url: wss://api.pendra.ai

# The models every worker in this group should serve, by catalogue id.
# These are the ids you see on console.pendra.ai → Models.
pendra_models:
  - qwen3.6:27b
  - nomic-embed-text

pendra_gpu_variant selects the package: cpu installs pendra, cuda installs pendra-cuda, and vulkan installs pendra-vulkan (NVIDIA via Vulkan, or AMD / Intel GPUs). Set it per host in host_vars if your fleet is mixed.

pendra_models is your list — every entry is a model id from the catalogue on console.pendra.ai → Models. Add or remove ids here and the next run reconciles each worker to match.

Keep the key in a vault. Put each host's pendra_worker_key in an encrypted host_vars file so the secret never sits in plaintext in your repo:
host_vars/gpu-01.example.com/vault.yml
# host_vars/gpu-01.example.com/vault.yml
# Encrypt this file:  ansible-vault encrypt host_vars/gpu-01.example.com/vault.yml
pendra_worker_key: "MC4CAQAwBQYDK2VwBCIEIExampleBase64Ed25519PrivateKey=="

2. The playbook

Register the signed repository, install the worker package for the host's GPU, template the config file, start the service, then install any models the worker is missing. It handles both Debian/Ubuntu (apt) and RHEL/Fedora/openSUSE (yum/dnf) hosts.

pendra.yml
# pendra.yml — install the worker, give it a key, install its models
- name: Provision Pendra workers
  hosts: pendra_workers
  become: true

  handlers:
    - name: restart pendra
      ansible.builtin.systemd:
        name: pendra
        state: restarted

  tasks:
    # 1. Register the signed Pendra package repository.
    - name: Add Pendra apt repository (Debian / Ubuntu)
      when: ansible_os_family == "Debian"
      block:
        - name: Install the Pendra signing key
          ansible.builtin.get_url:
            url: https://get.pendra.ai/pendra-archive-keyring.gpg
            dest: /usr/share/keyrings/pendra-archive-keyring.gpg
            mode: "0644"
        - name: Add the Pendra apt source
          ansible.builtin.apt_repository:
            repo: "deb [signed-by=/usr/share/keyrings/pendra-archive-keyring.gpg] https://get.pendra.ai/apt stable main"
            filename: pendra

    - name: Add Pendra yum repository (RHEL / Fedora / openSUSE)
      when: ansible_os_family in ["RedHat", "Suse"]
      ansible.builtin.yum_repository:
        name: pendra
        description: Pendra Worker
        baseurl: https://get.pendra.ai/yum/$basearch
        gpgcheck: true
        repo_gpgcheck: true
        gpgkey: https://get.pendra.ai/RPM-GPG-KEY-pendra

    # 2. Install the package that matches this host's GPU.
    - name: Install the Pendra worker
      ansible.builtin.package:
        name: "{{ 'pendra' if pendra_gpu_variant == 'cpu' else 'pendra-' + pendra_gpu_variant }}"
        state: present
        update_cache: true

    # 3. Write the worker key + config, owned by the pendra service user.
    - name: Configure the worker
      ansible.builtin.template:
        src: config.yaml.j2
        dest: /var/lib/pendra/config.yaml
        owner: pendra
        group: pendra
        mode: "0600"
      notify: restart pendra

    # 4. Enable the service so it starts now and on every boot.
    - name: Enable and start pendra
      ansible.builtin.systemd:
        name: pendra
        enabled: true
        state: started

    # 5. Install any models the worker doesn't already have.
    #    Running a task as the pendra user needs acl on the target, so
    #    make sure it's installed first.
    - name: Ensure acl is installed (needed to run tasks as the pendra user)
      ansible.builtin.package:
        name: acl
        state: present

    - name: Find the models already installed
      ansible.builtin.command: pendra models
      become_user: pendra
      changed_when: false
      register: pendra_models_present

    #    --wait returns once the worker is serving the model, and the
    #    exit code says what happened (0 serving, 1 failed, 2 unknown model),
    #    so a failed download fails the play instead of passing silently.
    - name: Install any missing models
      ansible.builtin.command: "pendra models install {{ item }} --wait"
      become_user: pendra
      loop: "{{ pendra_models }}"
      when: item not in pendra_models_present.stdout

And the tiny config template it renders:

templates/config.yaml.j2
# templates/config.yaml.j2 — rendered to /var/lib/pendra/config.yaml
app_server_ws_url: "{{ pendra_ws_url }}"
gpu_worker_private_key: "{{ pendra_worker_key }}"
worker_name: "{{ pendra_worker_name | default(inventory_hostname, true) }}"

A few things worth knowing about why the playbook is shaped this way:

  • The config file is owned by the pendra service user. On a package install the worker runs as the pendra user and reads /var/lib/pendra/config.yaml at mode 0600, so the template task sets owner: pendra, group: pendra, and mode: "0600". Templating the file directly (rather than shelling out to pendra config set) keeps the task declarative and idempotent.
  • There's no worker ID to template. Pendra gives each worker its ID from the host's worker key, so re-running the playbook keeps every host the same worker in the console. If an older version of this template set worker_id, you can drop the line — Pendra ignores it.
  • Models install as the pendra user. The install step runs become_user: pendra so the downloaded model files land in the service's own model directory, readable by the running worker. Running a task as the pendra user needs the acl package on the target, so the playbook installs it first. It also lists what's already installed and only pulls the models that are missing, so re-running skips anything the worker already has.
  • A key change restarts the worker. The template task notifies a handler that restarts pendra, so rotating a worker key (or changing the WS URL) takes effect at the end of the run.

3. Run it

bash
# Encrypt the per-host key file, then run the playbook
ansible-vault encrypt host_vars/gpu-01.example.com/vault.yml
ansible-playbook -i inventory.ini pendra.yml --ask-vault-pass

Ansible connects to each host, installs and configures the worker, and pulls the model list. The first run downloads model files, so give it a few minutes per large model; later runs skip anything already installed.

4. Verify

Ask every worker for its status over SSH — each should report connected and a model count:

bash
# Ask every worker for its status over SSH
ansible pendra_workers -i inventory.ini --become -a "pendra status"
gpu-01.example.com | CHANGED | rc=0 >>

  Status          ✔ connected
  Worker ID       wrk-3f9b8a2d9e104c1a
  Name            gpu-01.example.com
  Server          wss://api.pendra.ai
  Version         3.60.5
  Variant         cuda
  GPU             NVIDIA L4

  Models (2)
  ·  qwen3.6:27b
  ·  nomic-embed-text

Your machines should also appear under console.pendra.ai → Workers within a few seconds, with the models listed once each download finishes.

Keeping the fleet current

Because the playbook registered the Pendra package repository, worker upgrades are a normal package upgrade — a small second playbook pulls the latest worker onto every host, and the package restarts the running worker onto the new version for you:

upgrade.yml
# upgrade.yml — pull the latest worker onto the whole fleet
- name: Update Pendra workers
  hosts: pendra_workers
  become: true
  tasks:
    - name: Upgrade the worker package
      ansible.builtin.package:
        name: "{{ 'pendra' if pendra_gpu_variant == 'cpu' else 'pendra-' + pendra_gpu_variant }}"
        state: latest
        update_cache: true

To change what a fleet serves, edit pendra_models and re-run pendra.yml — new ids get installed, and everything already present is left untouched.

Next steps