OCI LXC in Proxmox 9.1: Five Production Traps to Fix
Fix five OCI LXC traps in Proxmox VE 9.1: networking, backup restore, kernel updates, privilege, and overlay storage to avoid 2 a.m. production pages.

On this page
OCI LXC containers in Proxmox VE 9.1 are production-ready, but only if you fix five specific traps before you stop treating them like test workloads. The networking config, backup restore path, host update cycle, privilege level, and overlay storage each have a failure mode that will bite you at 2 a.m. if you haven't addressed it in staging.
Key Takeaways
- Networking: Pin static IPs and MACs at the Proxmox level; container-internal config gets clobbered on image refresh.
- Backup: A vzdump restore to a node without the cached OCI image will fail silently or produce a broken container.
- Updates: Host kernel upgrades change the ABI under your containers; test in a maintenance window, not live.
- Privilege: Unprivileged is the default, but some workloads force privileged mode — accept the tradeoff deliberately.
- Storage: The overlay upper layer has no automatic cap; set alerts before it eats your pool.
Pitfall 1: Your Static IP Won't Survive an Image Update
When you create an OCI LXC container in Proxmox 9.1, you assign network settings through the Proxmox layer:
pct set 101 -ipconfig0 ip=192.168.1.101/24,gw=192.168.1.1
pct set 101 -hostname nginx-edge
This writes to the container's /etc/network/interfaces (or the equivalent for the base image's init system) at creation time. The trap: if you later update the container's base image — say, you re-pull docker.io/library/nginx:1.27 because the tag moved — Proxmox replaces the rootfs. Your Proxmox-level IP config is preserved in the container config file, but any custom networking you did inside the container (extra interfaces, DNS overrides in /etc/resolv.conf, custom routes) is gone.
The gotcha I hit: I had a container with a secondary interface for a management VLAN. I updated the image, and the container came back with only the primary interface. The management VLAN IP was gone. No error, no warning — just a missing route.
The fix is to treat the Proxmox-level ipconfig as the single source of truth and avoid in-container networking edits. If you need multiple interfaces, configure them all via pct set:
pct set 101 -ipconfig1 ip=10.10.0.101/24,mpath=vmbr1
If you're running a multi-service stack like the edge setup I documented in Docker Compose on Proxmox VE 9.1 OCI LXC, verify after every image update that all interface assignments are still present:
pct config 101 | grep ipconfig
Pitfall 2: Can You Restore a Backup to a Different Node?
This is the one that will actually lose you data if you're not careful. When you run vzdump on an OCI LXC container, Proxmox backs up the container's filesystem state — the overlay upper layer plus any custom files. The base OCI image itself is not part of the backup archive. It lives in the node's local image cache.
Here's the concrete scenario: you have a 2-node cluster. Node A has container 101 running nginx:1.27. You back it up to PBS:
vzdump 101 --storage pbs0 --mode snapshot --notes-template "daily-nginx"
Now node A dies. You create a new container on node B and restore the backup. The restore succeeds from the PBS perspective — the filesystem data is there. But node B doesn't have the nginx:1.27 OCI image cached. The container starts, but the rootfs is incomplete. You get a container that looks like it's running but is missing the base application binaries.
The fix: before restoring to a node, confirm the image is available there. In Proxmox 9.1, you can check the local OCI image cache:
ls /var/lib/vz/template/oci/
If the image isn't there, pull it first (via the GUI's OCI image manager or the API) before attempting the restore. I've written a more detailed walkthrough of the backup-to-PBS flow in OCI LXC Backup in Proxmox VE 9.1: vzdump Restore Guide, but the key operational rule is: verify the base image exists on the target node before you start the restore.
A practical timing note: pulling a 200 MB image on a 200 Mbps connection takes roughly 10 seconds. A larger image like docker.io/library/ubuntu:24.04 at about 78 MB is faster still. The image pull is not the bottleneck — the bottleneck is the vzdump restore of the upper layer, which for a container with 4 GB of accumulated data takes around 90 seconds on NVMe storage.
Pitfall 3: How to Survive a Host Kernel Update
OCI LXC containers share the host kernel. They don't bundle their own. This means a apt upgrade on the Proxmox host that bumps the kernel from 6.8.x to 6.9.x (or whatever the next stable is) changes the kernel ABI under every running container.
Most of the time, this is fine. The time it isn't: when a container workload depends on a specific kernel module that gets removed or changed in the new kernel. I've seen this with containers that use io_uring for high-performance I/O — a kernel update that changes the io_uring interface breaks the container at boot.
The tradeoff here is honesty: you cannot fully test a kernel update in a staging environment if your staging node runs the same kernel as production. The practical mitigation is a maintenance window with a rollback plan.
Before you update:
pve-version
Note the current kernel. Then, after the update:
pve-version
Confirm the new kernel is active. Restart containers one at a time, starting with the least critical:
pct stop 101
pct start 101
pct status 101
If a container fails to start, check its log:
pct exec 101 -- journalctl -n 50 --no-pager
If you're running a cluster and want to minimize the window, live-migrate VMs first to balance the load, then update the node with the fewest containers. This isn't a perfect solution — it's the one that keeps your 2 a.m. coffee warm.
Pitfall 4: Privileged Mode Is the Default Escape Hatch
When you create an OCI LXC container, Proxmox defaults to unprivileged mode. This is the right default for security, but it's not always the right default for your workload.
Unprivileged containers have a restricted /dev and a limited set of capabilities. Most OCI images work fine in this mode. But if your container needs to mount a filesystem, use userns in a specific way, or access hardware through /dev, you'll hit permission errors that don't appear in the container's own logs — they show up in the host's dmesg.
The moment you switch to privileged:
pct set 101 -unprivileged 0
you give the container root on the host (effectively). The container can see all devices, all network interfaces, and can modify the host's network namespace if it's not careful.
The honest tradeoff: privileged mode is not "unsafe" in a homelab or a trusted multi-tenant setup. It is unsafe if you're running untrusted images from a public registry without pinning digests. If you pin your image by digest (docker.io/library/nginx@sha256:abc123...) and run it privileged, your risk is bounded. If you use :latest and run it privileged, you're one upstream push away from a compromised container with host root.
For the seccomp profile that sits between these two extremes, I wrote up the configuration in LXC Seccomp Profile on Proxmox VE for Docker and K3s — it's not a replacement for the privileged/unprivileged decision, but it adds a layer of syscall filtering that catches the most common escape vectors.
Pitfall 5: The Overlay Upper Layer Grows Until You Notice
Here's the architecture: an OCI LXC container has a read-only lower layer (the base image) and a read-write upper layer (your changes). The upper layer is where all writes go — new files, modified files, deleted files (via whiteout markers).
The problem: nothing caps the upper layer automatically. If your container writes logs to /var/log (because the image's default logging config points there), or if it's a database container that accumulates WAL files, the upper layer grows. I had a PostgreSQL OCI container whose upper layer hit 12 GB in three weeks because the image's default log_rotation was set to daily with no retention. The container was fine — it was just eating my ZFS pool.
Monitoring this from the host:
du -sh /var/lib/vz/101/root/
Or from inside the container:
pct exec 101 -- df -h /
The fix is twofold. First, set a size alert in your monitoring (Netdata, Prometheus node exporter, whatever you use) on the /var/lib/vz/<vmid>/ path. Second, for containers that are known to write a lot, either:
- Mount a dedicated storage volume for the write-heavy path (via
pct set 101 -mp0 local:101/data,mp=/var/lib/postgresql/data), or - Schedule periodic
vzdumpbackups and rebuild the container from the image when the upper layer exceeds a threshold (say, 50% of the allocated disk).
The rebuild is fast — stopping the container, creating a fresh one from the same image, and restoring the data volume takes under 2 minutes. It's a manual process, but it's predictable and fast.
Conclusion
The five pitfalls above are not blockers — they're operational hygiene. Pin your networking at the Proxmox layer, verify image availability before cross-node restores, test kernel updates in a window, make your privilege decision deliberately, and set a storage alert before the upper layer surprises you. Do those five things in staging, and your OCI LXC workloads will run in production without the 2 a.m. page. Your next step: pick one of the five, apply it to your least-critical container today, and confirm the behavior matches what you expect before you roll it out to the rest of the fleet.


