↓ Skip to main content

The Update Routine

·9 mins
Table of Contents
Most homelab setups break in one of two ways: hardware fails, or software rots. The second one is slower and quieter: a package left unpatched for six months, a container image with a known CVE still running because “it works, I’ll update it later.”

My setup spans a three-node Proxmox cluster, a Talos Kubernetes cluster, a Synology NAS, pfSense, and Docker services on Oracle Cloud. Different release cadences, different failure modes if ignored. This is what keeps it current without becoming a second job.

Updated since

The hardware underneath has changed. The three-node Proxmox cluster is now a single R730xd, and the Synology is cold storage. The current shape is in the rack tour. The rolling-upgrade method below still applies to any multi-node cluster; the cadence table describes the layout as it was in May 2026.

The checklist
#

The foundation is a checklist in Joplin, reset monthly, covering every layer that needs a human. Its value isn’t the content. It forces a decision. You either tick it, or you consciously skip it.

LayerComponentCadenceMethod
NetworkpfSense, TP-Link WiFiMonthlyManual
HypervisorsProxmox cluster (3 nodes)2× / monthManual, rolling
HypervisorsDell R720On bootManual
StorageSynology DS223+AutomatedDSM scheduler
VMsTalos K8s nodesAutomatedRenovate
CloudOracle Ubuntu + Docker2× / monthapt + Portainer

The split between automated and manual is intentional. Automation handles what’s low-risk and high-frequency. Manual attention goes to the layers where a bad update can take down the whole environment.

Network
#

pfSense CE
#

pfSense Community Edition moves slowly. Major releases are infrequent, and Netgate increasingly nudges CE users toward pfSense Plus, but the CE branch is still maintained for security patches.

The standard update path (System → Update) covers full version upgrades. For between-version patches, the System Patches package adds a second layer: it pulls individual patches directly from the pfSense source repository and applies them without a full upgrade cycle.

To install it: System → Package Manager → Available Packages → search "System Patches".

Once installed, System → Patches shows any recommended patches for your running version. Most months the list is empty. When there is something, it’s typically a targeted fix, applied in under two minutes with no reboot.

Tip

pfSense patches don’t require a reboot unless the kernel is touched. Security patches to the PHP web layer, firewall rules engine, or service configs apply live. Full version upgrades are the only thing that needs a maintenance window.

TP-Link WiFi#

Firmware for the AX3000 ships infrequently, a few times a year at most. Monthly check via the admin interface is more than enough. The update applies in under two minutes with a brief disconnection; all connected clients reconnect automatically.

Hypervisors: the Proxmox cluster
#

The part most guides skip: how do you update a multi-node Proxmox cluster without taking everything offline?

My cluster is three nodes: Beelink GTi 13 (px-0, i9-13900H, 64GB DDR5), and two Dell OptiPlex 6500T units (px-1, px-2, each 32GB DDR4). Full specs in the rack tour. Each node runs one Talos VM for Kubernetes, so losing a node during an update means losing a K8s control plane node. Not catastrophic with three nodes, but still worth doing cleanly.

Rolling, one node at a time
#

The rule: update one node at a time, verify it’s back healthy before touching the next. Kernel updates require a reboot, and on a cluster node, that reboot needs to be planned.

  1. 1

    Check the cluster first

    pvecm status should show every node online. If one is already degraded, fix that first; never start a rolling upgrade on an unhealthy cluster.
  2. 2

    Move the VMs off

    qm migrate <vmid> <target-node> --online live-migrates each VM while it keeps running. For px-0, the Talos control plane VM goes to px-1.
  3. 3

    Upgrade

    apt update && apt dist-upgrade -y. On Proxmox it has to be dist-upgrade: plain upgrade won’t install new packages or remove obsolete ones, and kernel and pve package changes need both.
  4. 4

    Reboot and verify

    After the reboot, pvecm status for membership, uname -r for the new kernel, pveversion to confirm the version matches the other nodes.
  5. 5

    Move the VMs back

    Then the next node.
Proxmox node px-2 completing apt dist-upgrade

Why px-0 goes first
#

px-0 (Beelink) is the highest-spec node (64GB DDR5, i9-13900H). The Talos VMs temporarily land on px-1 and px-2 during the Beelink upgrade. Both OptiPlexes together have 64GB combined, enough to absorb the load short-term. Once px-0 is back online and clean, the VMs migrate back to it, and the OptiPlexes update one at a time with nothing critical running on them.

Note

In over a year of this routine, I haven’t had a single issue from a Proxmox upgrade. The dist-upgrade path is reliable. The rolling order matters not because Proxmox upgrades are risky, but because it keeps K8s healthy throughout: you never drop below two control plane nodes.

The R720
#

The R720 is nearly decommissioned and boots only occasionally for benchmarking or one-off experiments. The rule is simple: boot → update → do the work → shut down. Never leave it running on a stale package tree.

apt update && apt dist-upgrade -y

Storage: Synology DS223+
#

Synology DSM has two auto-update modes worth understanding:

  • Automatically install important updates: security patches and critical stability fixes only. Doesn’t touch major version upgrades or optional package updates.
  • Automatically install the latest updates: everything, including feature releases that can change behavior or require service restarts.

I use the first option. It runs on a schedule (Friday 02:30) so the NAS is never mid-update during active use.

Synology DSM Update Settings: important updates, Friday 02:30

Major DSM version upgrades I handle manually after reading the release notes. These occasionally change NFS/SMB behavior or package compatibility, so they are worth a few minutes of reading before applying.

Tip

The “important updates” mode is the right default for most homelab NAS setups. It covers CVEs and crash fixes automatically, while keeping you in control of anything that could disrupt your Plex, Synology Drive, or ARR stack mounts.

The biweekly manual check isn’t for applying updates. It verifies that the automated jobs ran. DSM logs every update under Control Panel → Update & Restore → Update History. One glance confirms nothing silently failed.

VMs: Talos, on Renovate
#

The three Talos VMs are the only layer I don’t touch manually at all. This is where automation actually pays off.

Renovate Bot monitors the infrastructure repo (github.com/meroxdotdev/infrastructure) and opens pull requests whenever it finds an updated version of anything: container images, Helm charts, Flux components, GitHub Actions, Talos OS itself, Kubernetes version.

The settings that matter
#

From .renovaterc.json5, trimmed:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
{
  extends: [
    "config:recommended",
    ":automergeBranch",
    ":dependencyDashboard",
    ":disableRateLimiting",
  ],
  schedule: ["every weekend"],
  packageRules: [
    {
      matchManagers: ["github-actions"],
      matchUpdateTypes: ["minor", "patch", "digest"],
      automerge: true,
      minimumReleaseAge: "3 days",
    },
    {
      matchManagers: ["mise"],
      matchUpdateTypes: ["minor", "patch"],
      automerge: true,
    },
  ],
}
  • schedule: every weekend: every PR lands in one weekend session instead of trickling in daily. One review cycle a week, not twenty.
  • :disableRateLimiting: by default Renovate caps how many PRs it opens at once, and in a repo with many dependencies some updates queue silently for weeks. Without the cap, everything surfaces on the weekend.
  • :automergeBranch: when something is allowed to automerge, it merges straight from its branch, without a PR.
  • minimumReleaseAge: 3 days: a GitHub Action doesn’t merge the day it’s published, which catches the occasional release that gets yanked within 48 hours.

What merges itself
#

Only two things: GitHub Actions (minor, patch and digest bumps) and the mise tool versions. Container images, Helm charts, Flux components and Talos all arrive as PRs that I review on the weekend. Major versions are always manual.

Two rules in the config exist because an update merged that shouldn’t have:

Some pairs are constraints, not questions

Talos and the kubelet are one decision: a Talos minor only serves a fixed range of Kubernetes versions, so Renovate groups the two images, and the kubelet is capped with allowedVersions below the first version the running Talos can’t serve. Immich’s Postgres image leads its tag with the Postgres major, so a “routine” 15 → 17 bump breaks the database on disk; major updates to it are disabled outright. In both cases a dashboard approval was tried first, and a batch merge got past it.

Things Renovate can’t see on its own
#

Some versions live in files Renovate has no manager for. A customManagers regex rule in the config matches an annotation comment, and the line under it becomes a dependency:

# renovate: datasource=docker depName=ghcr.io/siderolabs/kubelet
KUBERNETES_VERSION=v1.31.1

Renovate resolves the version against the Docker registry and opens a PR when a new tag appears. The rule’s file patterns cover .yaml, .env and .sh files, so the annotation works in any of them.

Tip

Renovate’s dependency dashboard (:dependencyDashboard) is a single GitHub issue listing every pending update, rate-limited item, and ignored package. It’s the single source of truth for what’s waiting, useful when you come back after a vacation and want to know what built up.

The weekly merge session for K8s takes about 10 minutes: scan the Renovate dashboard issue, check any major version PRs for breaking changes, merge. Everything else Renovate handled while I wasn’t looking.

Cloud: the Oracle instance
#

The Oracle instance runs Ubuntu 24.04 on an Ampere A1 (4 vCPU, 24GB RAM, the free tier). All Docker services run through Portainer.

Container images
#

Portainer detects outdated images automatically: when a container’s image has a newer digest available upstream, it shows a marker in the container list.

Portainer containers on the Oracle Cloud instance

The update flow: pull the new image → recreate the container. In Portainer’s UI: container detail → Recreate → enable “Re-pull image”. All volume mounts and environment variables carry over from the existing container config, so there is nothing to re-enter.

Warning

“Recreate” in Portainer preserves bind mounts and named volumes, but not anonymous volumes. If any of your containers use anonymous volumes for persistent data, those will be lost on recreate. Check before pulling.

The host itself
#

apt update && apt upgrade -y
apt autoremove -y

Twice a month, same as the Proxmox nodes. It’s internet-facing (Traefik and Pi-hole sit on it), so staying current matters more here, not less. A kernel update only takes effect after a reboot; /var/run/reboot-required exists when one is pending.

What to automate, and what not to
#

The pattern across all of this:

Automate when: the update is low-risk, high-frequency, and the failure mode is contained. Synology security patches, GitHub Actions and tool bumps are safe to automate because they’re scoped, reversible, and the blast radius is small. Container images sit on the line: Renovate finds them, a human merges them.

Stay manual when: the update touches the infrastructure layer itself: Proxmox host OS, pfSense, major DSM versions, K8s major upgrades. These can take down multiple services at once and benefit from a human checking cluster health before and after.

The checklist isn’t overhead; it’s the audit trail. It turns “I think I updated everything recently” into “I know I updated everything on these dates.”

Full infrastructure: the rack tour. Kubernetes GitOps setup: github.com/meroxdotdev/infrastructure.