Forming the Cluster — Participant Handout
Ansible for Proxmox VE · Module 03
What a Proxmox VE cluster is
Three things happen when nodes form a cluster:
- corosync provides cluster communication and decides who is in the membership.
- pmxcfs, the Proxmox cluster file system, mounts
/etc/pveand replicates every configuration file to all members. Editing a guest configuration on one node makes it visible on all of them. - Quorum, a majority of votes, is required before a node may change anything. With three nodes the quorum is two.
None of this is optional knowledge for the automation: quorum is the reason we join nodes sequentially, and pmxcfs is the reason a single API call changes the state of the whole cluster.
The collection
The Proxmox modules live in community.proxmox. They previously lived in community.general and were removed there in version 11, which is why many examples found online still use the old names.
Requirements worth checking before the first run:
| Requirement | Where | Check |
|---|---|---|
| ansible-core ≥ 2.17 | control node | ansible --version |
proxmoxer ≥ 2.3, requests |
control node | python3 -c 'import proxmoxer' |
| API reachable on port 8006 | from control node | curl -k https://192.168.0.1:8006/ |
Where API modules run
An SSH module runs on the managed node. An API module does not need to: it opens an HTTPS connection to port 8006 and talks to the Proxmox VE API. It therefore runs wherever proxmoxer is installed, which here is the control node.
The play still iterates over the Proxmox hosts, so ansible_host, hostvars and per-host variables all behave as usual. Only the execution location changes to the control node.
validate_certs: false accepts the self-signed certificate that every fresh Proxmox VE node ships with. In production you install a certificate from your own CA (or use the ACME integration) and leave validation switched on.
Handling credentials: lab vs. production
Credentials never belong directly in task definitions; they belong in variables. In inventory/group_vars/pve.yml:
In production, credentials belong in Ansible Vault or your CI/CD secret store. In this lab we keep them plain to avoid typing overhead.
The playbook
Play 1: create the cluster
playbooks/20-cluster.yml
---
- name: Create the cluster on the primary node
hosts: pve_primary
gather_facts: false
tasks:
- name: Ensure the cluster exists
community.proxmox.proxmox_cluster:
state: present
api_host: "{{ ansible_host }}"
api_user: root@pam
api_password: "{{ pve_root_password }}"
validate_certs: false
cluster_name: "{{ pve_cluster_name }}"
link0: "{{ ansible_host }}"
delegate_to: localhostlink0 is the address corosync uses for its first ring. Setting it explicitly means the cluster communicates over the network you chose, not over whichever address the node happened to resolve first. A second ring for redundancy would be link1.
The module is idempotent: it reads the cluster status first, reports ok if a cluster with that name exists, and fails clearly if the node belongs to a different cluster.
Play 2: join the remaining nodes
- name: Join the remaining nodes
hosts: pve_secondary
gather_facts: false
serial: 1
vars:
primary: "{{ groups['pve_primary'][0] }}"
tasks:
- name: Read the join information from the primary
community.proxmox.proxmox_cluster_join_info:
api_host: "{{ hostvars[primary].ansible_host }}"
api_user: root@pam
api_password: "{{ pve_root_password }}"
validate_certs: false
delegate_to: localhost
register: join_info
- name: Join this node to the cluster
community.proxmox.proxmox_cluster:
state: present
api_host: "{{ ansible_host }}"
api_user: root@pam
api_password: "{{ pve_root_password }}"
validate_certs: false
master_ip: "{{ hostvars[primary].ansible_host }}"
fingerprint: "{{ (join_info.cluster_join.nodelist
| selectattr('name', 'equalto', primary)
| first).pve_fp }}"
link0: "{{ ansible_host }}"
delegate_to: localhostThree details deserve attention:
serial: 1. Without it, Ansible would run the play for pve02 and pve03 in parallel and both would try to join at the same moment. corosync rewrites its configuration for every new member, and the joining node restarts its own services mid-way. Sequential joins are not a style preference; they are the supported path.
The fingerprint. proxmox_cluster_join_info returns the current cluster configuration, including a nodelist with one entry per member. We filter that list for the primary node by name and take its pve_fp value. This is why your inventory host names must match the real Proxmox VE node names.
master_ip. The address of a node that is already in the cluster. The module also uses it to recognise that a node is already a member, in which case it reports ok and does nothing.
Play 2, continued: wait for the node to actually arrive
The API accepts the join request and returns immediately; the work happens in the background, and the joining node’s web service restarts in the process.
- name: Wait until this node is a quorate cluster member
community.proxmox.proxmox_cluster_status_info:
api_host: "{{ ansible_host }}"
api_user: root@pam
api_password: "{{ pve_root_password }}"
validate_certs: false
delegate_to: localhost
register: cl
until: cl.cluster_status | default([])
| selectattr('type', 'equalto', 'cluster')
| selectattr('quorate', 'equalto', true) | list | count > 0
retries: 30
delay: 5Two details in that condition are not decoration:
quorate, not just “a cluster exists”. POST /cluster/config returns before corosync has reached a quorum. A node that joins in that window is rejected with cluster not ready - no quorum?, and the join task reports changed anyway because the API accepted the request. Waiting for port 8006 does not help either, because pveproxy never went away.
default([]). The joining node restarts its own web service during the join, so the status call fails for a few seconds. Without the default, the until expression raises on an undefined key and the task aborts instead of retrying.
Skipping this is the classic way to build a playbook that works when you run it by hand and fails in a pipeline: the next play starts before the previous change has landed.
The same thing with curl
Every module in this collection is a wrapper around one or two REST calls. Seeing the call underneath tells you what the module can possibly do, makes error messages readable, and gives you a fallback when a module lacks an option.
Set up the credentials once. An API token goes in as an HTTP header, with no login round-trip:
-k accepts the self-signed lab certificate, jq formats the JSON envelope, which always wraps the payload in a data key.
Nothing in this workshop creates a token for root@pam, so make one on a node before you try these calls:
pveum user token add root@pam ansible --privsep 0The secret is printed once. If you use the automation@pve token from module 05 instead, the cluster-wide endpoints come back empty rather than failing: that account holds PVEVMAdmin, which does not include Sys.Audit. An empty result from the Proxmox VE API usually means missing privileges, not missing data.
| Ansible module | HTTP call | API viewer |
|---|---|---|
proxmox_cluster (create) |
POST /cluster/config |
/cluster/config |
proxmox_cluster (join) |
POST /cluster/config/join |
/cluster/config/join |
proxmox_cluster_join_info |
GET /cluster/config/join |
same page |
proxmox_cluster_status_info |
GET /cluster/status |
/cluster/status |
Creating the cluster
Reading the join information
That pve_fp is exactly the value the playbook extracts with selectattr.
Joining a node
Note the host: this call goes to the joining node, not to the primary.
The response is a UPID, a task id rather than a result. The join runs in the background, which is precisely why the playbook waits afterwards. You can follow the task:
curl -sk -H "Authorization: $TOKEN2" "$PVE2/nodes/pve02/tasks/$UPID/status" | jqChecking the status
The complete API of your cluster is browsable on every node at https://<node>:8006/pve-docs/api-viewer/, including the exact parameters of your version. The public copy is at https://pve.proxmox.com/pve-docs/api-viewer/.
The CLI fallback
If the collection is unavailable (an air-gapped environment, an older control node), the same result is reachable over SSH. You then own the idempotency:
Joining over the CLI additionally needs passwordless root SSH from the joining node to the primary, because pvecm add --use_ssh authenticates that way. That is extra moving parts for the same outcome, so prefer the module when you can.
Verifying the result
pvecm status should show Quorum information with Quorate: Yes, three votes, and the expected cluster name. ls /etc/pve/nodes returns the same three directories on every node, which is pmxcfs replication at work.
Then run the playbook again. Everything reports ok. The cluster has stopped being something that happened and has become something you describe.
/root/.ssh/authorized_keys on a Proxmox VE node is a symlink into pmxcfs:
/root/.ssh/authorized_keys -> /etc/pve/priv/authorized_keysSo it is not a per-node file, it is cluster-wide. A node that joins gets the primary’s /etc/pve, and with it the primary’s copy of that file. Keys that existed only on the joining node are gone afterwards.
deploy.sh installs your key on all three nodes before you form anything, so the primary’s copy already contains it and the join changes nothing for you. Worth knowing before you distribute keys to a cluster you are about to build: put them on the primary, or put them everywhere.
Quorum, in one table
| Nodes online (of 3) | Quorate | What still works |
|---|---|---|
| 3 | yes | everything |
| 2 | yes | everything, without further redundancy |
| 1 | no | /etc/pve read-only, no guest starts, no configuration changes |
For two-node clusters a QDevice provides the third vote. That is a topic of the Clustering & Shared Storage course.
The Proxmox VE API can create and join clusters, but it cannot make a node leave one. Removing a node is a deliberate, partly manual procedure (pvecm delnode, followed by reinstalling the removed node before any reuse). Plan cluster membership before you automate it.
Further reading
The collection
community.proxmoxcollection index (all modules, the inventory plugin, and the authentication guide)proxmox_clusterproxmox_cluster_join_infoproxmox_cluster_status_info- Source and issue tracker on GitHub (when a module does not do what you need, the source is short and readable, and the issue tracker usually already has your question)
- Installing collections
Ansible concepts from this module
- Controlling where tasks run: delegation (
delegate_to,run_once) - Controlling playbook execution: strategies and more (
serial, and why it exists) - Retrying a task until a condition is met (
until,retries,delay) ansible.builtin.wait_for- Jinja filters:
selectattr,map,first
Proxmox VE
- Cluster Manager (
pvecm): read the sections on requirements, adding nodes, and separating the cluster network - Cluster Manager wiki page
- API viewer: look up
/cluster/configand/cluster/config/join, which is exactly what the module calls man pvecmon any node
Offline
ansible-doc community.proxmox.proxmox_cluster
ansible-doc -l community.proxmox # everything the collection offers