PBS 4.2 S3 Backup Pre-Flight Checklist for Production

PBS 4.2 S3 backup setup can saturate uplinks and inflate costs. Use this checklist to validate endpoints, throttle syncs, and set retention that saves budget.

9 min read
proxmox-backup-servers3-object-storagebandwidth-throttlingbackup-retentioniam-policy
A dark server room with storage racks, glowing cables, and a luminous arc of data particles.

Pointing production PBS backups at S3 in PBS 4.2 is straightforward to configure but easy to misconfigure in ways that hit your wallet or saturate your uplink within the first night. This is the pre-flight checklist I run before any S3 target goes live: endpoint validation, bandwidth throttling, and retention math that actually matches your restore requirements.

Key Takeaways

  • Throttle first: Always set upstream bandwidth limits before your first sync job runs — an unthrottled 500 GB sync will eat a 1 Gbps uplink for hours.
  • Retention compounds cost: A 1 TB backup with aggressive retention policies can cost 3–5× more per month than a conservative one on the same S3 provider.
  • Endpoint style matters: Path-style vs. virtual-hosted-style URLs break silently on some providers, and PBS won't tell you which format it resolved.
  • Bucket permissions: PBS needs specific S3 operations (PUT, GET, LIST, DELETE); a missing policy fails mid-sync and leaves partial data.

Why S3 Targets Are Tempting (and Where They Bite)

PBS 4.2 lets you configure S3-compatible object storage as a backup target directly from the PBS web interface. You don't need a second PBS node, a WireGuard tunnel, or a dedicated replication link. You point a backup job at a bucket and PBS syncs the data. It's the simplest off-site backup path PBS has ever offered.

The temptation is real: you skip the hardware, skip the network configuration, and your backups land in the cloud. But I've seen three failure modes show up consistently in the first week of a new S3 target:

  1. Ulink saturation — the first full sync runs at line rate and your remote site can't reach the internet for four hours.
  2. Retention bloat — you set keep-daily: 30 without doing the math, and your monthly bill triples by month three.
  3. Silent endpoint misconfig — the target shows as "connected" in the UI, but the first sync fails with a cryptic 403 or hangs on a LIST operation.

I'll walk through each one and the specific checks that prevent them.

Pre-Flight: S3 Target Configuration

Validate the Endpoint Before You Save

The S3 endpoint field in PBS accepts a base URL. How PBS constructs the final request URL depends on the provider's expected addressing style. AWS supports both path-style and virtual-hosted-style. Backblaze B2 and Wasabi expect a specific format. MinIO depends on your deployment.

Before you add the datastore in PBS, test the endpoint from the PBS host:

# Test endpoint reachability and TLS
curl -sI https://s3.us-east-1.amazonaws.com

# For path-style (AWS):
curl -sI https://s3.us-east-1.amazonaws.com/my-backup-bucket

# For virtual-hosted-style (AWS):
curl -sI https://my-backup-bucket.s3.us-east-1.amazonaws.com

# Backblaze B2:
curl -sI https://s3.us-west-004.backblazeb2.com

# Wasabi:
curl -sI https://s3.us-east-1.wasabisys.com

A 200 or 403 response means the endpoint is reachable. A 404 or connection timeout means you have the wrong region or the bucket doesn't exist yet. The bucket must exist before PBS can use it — PBS will not create it for you.

Bucket and IAM Policy

PBS needs the following S3 operations on your bucket:

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": {
        "AWS": "arn:aws:iam::YOUR_ACCOUNT_ID:user/pbs-service"
      },
      "Action": [
        "s3:PutObject",
        "s3:GetObject",
        "s3:ListBucket",
        "s3:DeleteObject",
        "s3:ListMultipartUploadParts",
        "s3:AbortMultipartUpload"
      ],
      "Resource": [
        "arn:aws:s3:::my-backup-bucket",
        "arn:aws:s3:::my-backup-bucket/*"
      ]
    }
  ]
}

The ListBucket permission on the bucket ARN and the object-level permissions on the /* ARN are both required. I've seen people grant only the object-level permissions and wonder why the initial sync hangs.

Adding the Datastore in PBS

In the PBS web interface, go to Datacenter → Datastores → Add. Select type s3 and fill in:

{
  "type": "s3",
  "endpoint": "https://s3.us-east-1.amazonaws.com",
  "bucket": "my-backup-bucket",
  "access_key": "AKIAIOSFODNN7EXAMPLE",
  "secret_key": "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
}

After saving, PBS will show the datastore status. If it shows "connected," the credentials and endpoint are at least parseable. That's not the same as "fully functional" — the first sync is the real test.

This is where most people get burned. The first full sync of a 500 GB backup set will push data at whatever rate your uplink allows. On a 1 Gbps residential connection, that's roughly 110 MB/s sustained. If you have other users on that link, they'll notice.

Set Bandwidth Limits on the Backup Job

In PBS, bandwidth throttling is configured per backup job (or per sync/replication job). In the job configuration, set the upstream bandwidth limit:

# Via API (example for a job named "prod-vm-backup")
# Bandwidth is in bytes per second
# 50 MB/s = 52428800 bytes/sec
curl -k -X POST "https://pbs.example.com:8007/api/json/backup/job/prod-vm-backup" \
  -H "Authorization: PVEAPIToken root@pbs!pbs" \
  -d '{
    "bandwidth": 52428800,
    "bandwidth-downstream": 104857600
  }'

Or in the web UI: Datacenter → Backup → [job name] → Edit → General tab → Bandwidth.

My rule of thumb for homelab and small office uplinks:

Uplink Speed Recommended Upstream Limit Rationale
1 Gbps 50 MB/s (400 Mbps) Leaves 60% headroom for other traffic
500 Mbps 25 MB/s (200 Mbps) Leaves 60% headroom
100 Mbps 10 MB/s (80 Mbps) Leaves 20 Mbps for everything else
50 Mbps 5 MB/s (40 Mbps) Leaves 10 Mbps for interactive use

Schedule the Sync During Off-Peak Hours

Throttling helps, but a full sync of 1 TB at 50 MB/s still takes about 5.5 hours. Schedule it for 01:00–06:00 when the link is idle. In PBS, the schedule field uses standard cron syntax:

# Run at 01:00 daily
"schedule": "0 1 * * *"

I cover the full scheduling and pause workflow in my PBS replication throttling and scheduling guide, which applies the same principles to PBS-to-PBS replication. The S3 target uses the same job-level bandwidth fields.

The First Sync Is Different

The first sync is a full transfer. Every subsequent sync is incremental — PBS only sends new or changed data blocks. If your backup set is 1 TB and you change 50 GB per day, your nightly sync after the first will be roughly 50 GB, not 1 TB. Plan your throttling for the first sync specifically, then relax it for steady-state operation.

Retention Rules and Their Cost Implications

Retention is where S3 costs quietly grow. PBS retention policies work by keeping snapshots that match a set of rules. Every snapshot that survives pruning stays in the bucket and incurs storage charges.

The Math

Let's say you have a 1 TB backup set with daily changes of 50 GB. Here's what different retention policies cost in stored data (approximate, assuming PBS deduplication reduces stored size by ~40%):

Policy Retained Snapshots Approx. Stored Data Monthly Cost (S3 Standard)
keep-last: 7, keep-weekly: 4 ~11 ~660 GB ~$15
keep-last: 30, keep-monthly: 12 ~42 ~1.8 TB ~$41
keep-last: 90, keep-yearly: 5 ~95 ~4.5 TB ~$103

These are rough figures assuming 60% of the data is unique after dedup. Your numbers will differ based on workload. The point is: keep-last: 90 is not a free safety net. It's a 6× cost increase over keep-last: 7.

Configure Retention Deliberately

In the PBS backup job configuration:

{
  "prune-backup": 1,
  "prune-keep-last": 7,
  "prune-keep-daily": 0,
  "prune-keep-weekly": 4,
  "prune-keep-monthly": 6,
  "prune-keep-yearly": 1
}

This gives you:

  • Last 7 daily snapshots (about a week of recovery points)
  • 4 weekly snapshots (about a month)
  • 6 monthly snapshots (about half a year)
  • 1 yearly snapshot

For most production workloads, this is the right balance. You can restore to any point in the last week at daily granularity, any point in the last month at weekly granularity, and any point in the last six months at monthly granularity.

The Tradeoff

If you need point-in-time recovery to a specific minute within a 30-day window, you need keep-last: 30 or keep-daily: 30. That's a real requirement for some databases and mail servers. But if your RPO is "last night's backup is fine," then keep-last: 7 saves you real money. I've seen homelab operators set keep-last: 365 on a 200 GB backup set and not notice the $50/month bill until the annual statement.

Common Mistakes I'd Catch in a Pre-Flight Review

Here are the specific things I check before approving an S3 target for production:

  1. No bandwidth limit set. The job will run at line rate. If your uplink is 1 Gbps and the first sync is 800 GB, that's 12+ hours of saturated uplink.

  2. Bucket in a different region than the PBS host. Cross-region S3 transfers are slower and cost more for data transfer. If your PBS node is in Frankfurt, put the bucket in eu-central-1.

  3. Using S3 Standard for long-term retention. If your retention policy keeps data for a year, S3 Standard is the wrong tier. S3 Standard-IA or S3 Glacier Instant Retrieval cuts storage cost by 40–60% with no access penalty for your use case. I compare the tiers in my PBS 4.2 S3 cloud storage cost comparison.

  4. No monitoring on the sync job. PBS will log failures, but you won't see them unless you check. Set up a notification (email or webhook) for job failures. A sync that fails silently for three days means your "off-site backup" is three days stale.

  5. Forgetting that S3 DELETE is permanent. There's no undo. If PBS prunes a snapshot and you realize you needed it, it's gone. This is another reason to set retention deliberately rather than "just in case."

A Note on Local Storage as the Primary Target

The S3 target is your off-site copy. Your primary PBS datastore should still be on local, fast storage — a ZFS pool on NVMe or a good SSD array. The PBS-to-S3 sync is the second leg of your backup strategy, not the first. If your primary datastore is on S3 directly, every restore reads from S3 and you're paying data transfer-out charges on top of the restore time. I've written about building the local ZFS side of this in my Proxmox NAS and ZFS pool guide.

Conclusion

The S3 target in PBS 4.2 removes the network and hardware complexity from off-site backup, but it shifts the risk to configuration: throttling, retention math, and endpoint correctness. Run the pre-flight checks above — validate the endpoint with curl, set a bandwidth limit before the first sync, calculate your retention cost before you save the job — and you'll avoid the three failure modes that show up in week one. Your next step: add the S3 datastore, run a single test backup of a small VM, verify the data lands in the bucket, and only then point your production jobs at it.

Share
Proxmox Pulse

Written by

Proxmox Pulse

Sysadmin-driven guides for getting the most out of Proxmox VE in production and homelab environments.

Related Articles

View all →