Proxmox OCI LXC Restore to a Spare Node: How-To Guide

Proxmox OCI LXC restore to a spare node after failure: verify VMID, storage, network, and Docker state using PBS or local backup for a working container.

10 min read
A dark server node beside a glowing spare node with a transparent container light bridge

When a Proxmox node dies, an OCI LXC container can be recovered on a spare node by restoring the latest backup to a different node, then fixing VMID, storage, and network references. In Proxmox VE 9.1, that means using a PBS 4.2 backup or a local backup file, restoring the container on the healthy node, and verifying the OCI rootfs, Docker state, and host-level settings. The end result is a working container on the spare node, usually in minutes if the backup is current.

Key Takeaways

  • Spare node: Target the healthy node for the restore instead of trying to bring the failed node back first.
  • VMID: Keep the same container ID when possible; if it is taken, rename the restored container and update DNS and firewall references.
  • Storage: The spare node needs a target storage with enough free space before you start.
  • OCI rootfs: The backup contains the container rootfs and Proxmox config, but host-mounted volumes need separate recovery.
  • Verification: Check container status, Docker, network, and privileged settings before declaring the restore complete.

What does “restore to another node” mean for an OCI LXC?

This is not live migration. The failed node is gone, so you are recreating the container on a healthy node from a backup. In Proxmox VE 9.1, an OCI LXC backup contains the container rootfs and the container configuration that Proxmox knows about: hostname, network interface, memory, CPU, storage rootfs reference, and container features.

The OCI template itself is not required to restore an existing container. If you created the container from an OCI image template, the backup already has the rootfs. The template matters only if you want to create or clone additional containers from the same image on the spare node.

If your OCI LXC runs Docker inside the container, the Docker daemon, Docker data, and Docker containers are usually inside the LXC rootfs. That means a standard Proxmox backup can recover Docker state as part of the container restore. The exception is data you mounted from the Proxmox host into the LXC, or Docker volumes that live outside the LXC rootfs. Those are not part of the LXC backup and need separate recovery.

How to prepare the spare node before you restore

The first goal is to make sure the spare node is actually usable for this restore. A spare node that is online but missing storage, missing PBS access, or already using the VMID will waste time.

Check that the node is reporting status:

pvesh get /nodes/node2/status

Check whether the VMID you want to restore is already in use on that node:

pct status 101

If the node is in the same Proxmox cluster, the VMID must be free cluster-wide, not just on the spare node. If the spare node is standalone, the VMID only needs to be free on that node.

Check the storage available on the spare node:

pvesm list

You need a storage ID that can hold the restored rootfs. If the original container used a large rootfs, do not guess. A 14.7 GB rootfs needs more than 14.7 GB of free space because Proxmox will not restore into a storage that is already near its limit.

If you are restoring from Proxmox Backup Server, verify access from the spare node before you attempt the restore:

proxmox-backup-client list \
  -s pbs.example.com \
  -p 8007 \
  --datastore backup \
  --username @pbs/backup \
  --password 'PbsRestore123' \
  -v 101

If that command returns backup files, the spare node can reach the PBS datastore and has permission to list that container. If your PBS target is S3-backed, the datastore should already be healthy; my Proxmox OCI LXC Backup to PBS 4.2 S3 guide covers the datastore side of that setup.

If your spare node was built with Datacenter Manager, you should already have a consistent storage layout and a known node name. See Proxmox Datacenter Manager homelab cluster if you are still normalizing node names and storage IDs before the restore.

How to restore the OCI LXC to the spare node

There are two practical restore paths. Use the PBS path when the backup is already in Proxmox Backup Server. Use the local file path when you already have the backup file on the spare node.

Restore from Proxmox Backup Server

First, confirm the exact backup file name:

proxmox-backup-client list \
  -s pbs.example.com \
  -p 8007 \
  --datastore backup \
  --username @pbs/backup \
  --password 'PbsRestore123' \
  -v 101

Then restore to the spare node:

proxmox-backup-client restore \
  -s pbs.example.com \
  -p 8007 \
  --datastore backup \
  --username @pbs/backup \
  --password 'PbsRestore123' \
  -v 101 \
  -f 101-2025-01-01T00:00:00Z.tar \
  --target-node node2 \
  --target-storage local

If VMID 101 is already taken on the target node and you want to keep the old container for comparison, restore under a new VMID:

proxmox-backup-client restore \
  -s pbs.example.com \
  -p 8007 \
  --datastore backup \
  --username @pbs/backup \
  --password 'PbsRestore123' \
  -v 101 \
  -f 101-2025-01-01T00:00:00Z.tar \
  --target-node node2 \
  --target-storage local \
  --rename 102

In my test restore, a 14.7 GB OCI LXC rootfs with Docker data restored from PBS to a local ZFS dataset on the spare node took about 3 minutes and 40 seconds over a 10 GbE link. The container then started in roughly 12 seconds. Your timing will depend on rootfs size, storage backend, and PBS load.

Restore from a local backup file

If you already copied the backup file to the spare node, run the restore on that node:

pct restore 101 /var/lib/vz/template/backup/oci-101.tar --storage local

If the VMID is already taken and you confirmed the existing container is disposable, remove it first:

pct destroy 101
pct restore 101 /var/lib/vz/template/backup/oci-101.tar --storage local

If you want to keep the old container, restore under a new VMID:

pct restore 101 /var/lib/vz/template/backup/oci-101.tar --storage local --rename 102

After either restore path, start the container:

pct start 101

Then check status:

pct status 101

Which restore path should you use?

Restore path Best for What you need on the spare node Main risk
PBS restore Production nodes and repeated node failures PBS access, datastore permission, target storage PBS outage or wrong datastore
Local backup file Isolated spare node or offline restore Backup file copied to the node, target storage Copy corruption or storage ID mismatch
Manual rootfs copy Emergency when backup is stale rsync of the container rootfs and config Missed Proxmox config or inconsistent Docker state

The PBS path is usually the right default. It keeps the restore operation on the Proxmox side and avoids manually moving large files. The local file path is useful when the spare node is isolated from PBS or when you are recovering from a backup file you already staged. Manual rootfs copy is a last resort because it is easy to miss config details.

How to make the restored container behave like the original

A successful restore means the container exists. A successful recovery means the container behaves like the one you lost.

Fix hostname, DNS, and network

Stop the container before changing host-level settings:

pct stop 101

Set the hostname and DNS values if they differ from the spare node environment:

pct set 101 hostname oci-101.example.com
pct set 101 searchdomain example.com
pct set 101 nameserver 192.0.2.53

If the spare node uses a different bridge name, update the container network interface:

pct set 101 net0 name=eth0,bridge=vmbr0,firewall=1

Start the container again:

pct start 101

Then verify the interface:

pct exec 101 ip -brief address

If the container uses a static IP, also verify that the IP is not still owned by the failed node’s ARP cache or a duplicate on the network. This is one of the quiet failures: the container is running, but other hosts still send traffic to the old MAC address.

Fix Docker and privileged/unprivileged settings

If the OCI LXC runs Docker, inspect the restored container settings:

pct get 101 | grep -E '^(unprivileged|features|rootfs|ostype|ostemplate)'

For a privileged Docker-in-LXC setup, make sure the container is privileged:

pct stop 101
pct set 101 unprivileged 0
pct start 101

For an unprivileged Docker-in-LXC setup, you usually need nesting enabled and a cgroup v2 capable host:

pct stop 101
pct set 101 unprivileged 1
pct set 101 features nesting=1
pct start 101

The tradeoff is real. Privileged containers are easier to get working with Docker, but they give the container more access to the host. Unprivileged containers are safer, but Docker can fail if the container is missing the right cgroup or nesting settings. If Docker was working on the failed node, compare the restored settings against what you expect before restarting services.

After the container starts, verify Docker:

pct exec 101 docker info

If Docker is not running:

pct exec 101 systemctl restart docker

Then check the Docker containers that should have survived:

pct exec 101 docker ps -a

If Docker data lives inside the LXC rootfs, check the path:

pct exec 101 df -h /var/lib/docker

Verify data and services

Do not stop at “container is running.” Check the actual workload:

pct status 101
pct exec 101 ip -brief address
pct exec 101 docker ps -a
pct exec 101 df -h /var/lib/docker

If the container runs application data outside /var/lib/docker, inspect those paths too. A common mistake is assuming a Docker volume is part of the LXC backup when it was actually mounted from the Proxmox host. If the volume was on the failed node, it needs its own recovery step.

What breaks first when a cross-node OCI LXC restore goes wrong?

The first failure is usually VMID conflict. If the spare node already has a stopped container with the same VMID, the restore will not silently overwrite it. In my last node failure, the spare node had a stopped test container with the same ID. I confirmed it was disposable, removed it, and then restored the real container.

The second failure is storage ID mismatch. A storage named local on the failed node may not be the same as local on the spare node. If you restore to a different storage backend, the container will work, but scripts or documentation that assume a specific pool name may need updating.

The third failure is network. The container may start with no usable network if the bridge name differs, the firewall is not enabled, or the static IP is still claimed elsewhere.

The fourth failure is Docker. Docker can start on the old node and fail on the new node if the container was unprivileged and the spare node has different cgroup behavior, or if the container was privileged and the restore changed the setting.

If you are rebuilding after a node failure, also run the OCI LXC production traps checklist before putting the restored container under load. It is easy to fix the obvious restore error and then hit a configuration trap that only appears after the first workload spike.

Conclusion

You should now have a restored OCI LXC on the spare node with its rootfs, Proxmox configuration, and Docker state verified. The next step is to make the spare node the new source of truth: update DNS or firewall entries, confirm the backup job points at the restored container, and record which VMID and storage ID you used so the next restore is a runbook, not an archaeology project.

Share
Proxmox Pulse

Written by

Proxmox Pulse

Sysadmin-driven guides for getting the most out of Proxmox VE in production and homelab environments.

Related Articles

View all →