↓ Skip to main content

Built to be rebuilt.

Three Proxmox hosts, one Kubernetes cluster, and nothing that answers the internet. All of it comes back from a git repository and three keys in about thirty-five minutes — which is the only reason it is safe to keep changing.

Hypervisor
Proxmox VE
Kubernetes
Talos, Flux
Network
pfSense, Cilium
Storage
Synology, TrueNAS
Power
~65 W at the wall
The office corner: a standing desk running monitoring dashboards, and the open rack beside it — R730xd, mini PCs, NAS and UPS

Network

No way in.

Nothing in the rack answers the internet. One UDP port is forwarded, for Tailscale’s direct WireGuard connections, and it ignores anything without a key. Every public name arrives through a tunnel that dials out.

pfSenseFirewall and gateway on an XCY X44. Physical segmentation — one NIC per segment, not VLAN tags.
TailscaleA subnet router puts the LAN on the tailnet, so the rack and the cloud VPS share one flat mesh. This is the way in.
Cloudflare TunnelOutbound-only, one per place that needs it: the cluster, the VPS, and the NAS for Synology Drive — so Drive stays up when Kubernetes does not.
UnboundRecursive resolution, no upstream forwarder, with overrides for the internal names.

Grafana, Internet row: 936 Mb/s down, 657 Mb/s up, 21 ms round trip, and six hours of flat throughput and latency history
Grafana · 4 Oct 2026

Compute

Three small machines. One big one, asleep.

Standalone hosts, joined through Proxmox Datacenter Manager. One Kubernetes node each, on three different chassis. The R730xd that used to carry most of this is now the vault — off most of the day.

Beelink GTi13 Ultra

CPUCore i9-13900HK · 20 threads
Memory64 GB DDR5
Nodekubernetes-1 · 14 cores / 32 GB
WhyIris Xe passed through — every transcode happens here

Dell OptiPlex 3050

CPUCore i5 · 4 cores
Memory32 GB DDR4
Nodekubernetes-2 · 3 cores / 16 GB
Why960 GB D3-S4510 passed through raw — power-loss protection, the right disk for etcd

Dell OptiPlex 3050

CPUCore i5-6500T · 4 cores
Memory32 GB DDR4
Nodekubernetes-3 · plus PDM and Garage, Longhorn’s backup target
WhyThe same D3-S4510, and the host that wakes the vault

Dell PowerEdge R730xd

CPUXeon E5-2630 v4 · 10C/20T
Memory251 GB DDR4 ECC
RunsTrueNAS, nothing else
WhyTwelve SAS bays for the offline copy. Off about 22 hours a day

Grafana, Power row: battery at 100%, 36 minutes of runtime, the estate drawing 74 W, and six hours of draw between 51 and 94 W
Grafana · 4 Oct 2026

Storage

One live store. One vault that sleeps.

Everything that changes lives on a two-disk Synology: the film library, my files, my photos, and the landing spot of every backup. The cluster keeps only its own app volumes, on Longhorn, three replicas.

WhereLayoutHolds
NAS · DS2232 × 2 TB, mirroredthe live store — 1.1 TB of 1.8 TB used
vault · R730xd12 × 600 GB SAS, one RAIDZ3, encryptedhistory: daily and monthly ZFS snapshots
LonghornNVMe on pve-1, D3-S4510 on pve-2 and pve-3app volumes, one replica per node
Grafana storage panels: the NAS-backed Garage storage at 60%, every other Proxmox storage under 23%, and SSD wear from 35% on one pve-1 disk down to zero on the D3-S4510s
Grafana · 4 Oct 2026
Tip

I made twelve SAS disks park whenever nothing read them, and it worked — the story is in Spinning Down SAS Disks. Then I moved the live data to the NAS and let the whole R730xd power off instead. Parking a disk saves a few watts. Not running the server saves all of them.

Kubernetes

git push is the deploy.

Talos has no shell and no package manager — it is configured by API and by nothing else. Flux watches the repository and makes the cluster match it. There is no third way to change anything, which is what makes the rebuild time believable.

kubernetes/apps/
├── default/         jellyfin, the *arr stack, n8n
├── kube-system/     Cilium, CoreDNS, Intel GPU plugin
├── network/         gateway, Cloudflare tunnel, netboot.xyz
├── observability/   Prometheus, Grafana, Loki, Alloy
└── storage/         Longhorn
Headlamp’s map of the cluster: kube-system, default, longhorn-system and observability collapsed; flux-system open on its five controllers; network on cloudflare-dns, cloudflare-tunnel, netboot-xyz and k8s-gateway; cert-manager on its three deployments
Headlamp · 11 Sep 2026

Grafana, Kubernetes row: cluster CPU at 5%, memory at 19%, and eleven app tiles, every one green and Up
Grafana · 4 Oct 2026

Services

What it actually runs.

Media and automation in the cluster, files and photos on the NAS, identity and remote access on the VPS. Everything with data behind it is declared in the repo.

Media

  • Jellyfin: 4K, hardware transcoded
  • Jellyseerr: requests, the one route that faces out
  • Radarr · Sonarr: film and TV automation
  • Prowlarr · qBittorrent: indexing and download, VPN-routed

Files & photos

  • Synology Drive: files, synced to my Mac; in a browser from Romania only
  • Synology Photos: the photo library, filed by country
  • Joplin: notes sync

Platform

  • Talos + Flux: the cluster and what reconciles it
  • Longhorn: block storage, three replicas
  • Cilium: eBPF CNI, Gateway API, L2
  • cert-manager: certificates, renewed without asking

Edge & identity

  • Traefik: reverse proxy on the VPS
  • Authentik: single sign-on
  • Guacamole: remote desktop in a tab
  • Pi-hole + Unbound: filtering DNS, own resolver

Watching it

  • Prometheus: metrics, hosts and cluster
  • Grafana: dashboards
  • Loki + Alloy: logs
  • Alertmanager: straight to Telegram

Odds

  • n8n: one workflow, the daily news digest
  • Headlamp: Kubernetes, when a UI is faster
  • Portainer: the VPS containers
  • netboot.xyz: installers over the network

Portainer on the VPS: the Authentik server, worker, Postgres and Redis, Guacamole, the public Homepage, Pi-hole, Joplin and its database, Portainer itself, the restic REST server and Traefik; every container healthy or running
Portainer · the VPS · 11 Sep 2026

Alerting

Silence is the alarm.

Alertmanager, Flux, Proxmox and TrueNAS go straight to Telegram — critical and warning only, tagged by layer. Every scheduled job also pings healthchecks.io on its way out, so a job that stops running raises its own hand. Nobody has to notice.

A message gets read once, though, and then it scrolls away. So there is one page for looking on purpose: sixteen numbers that sit neutral on a good day and turn amber or red only when one needs me, and below them every console in the house, each with a dot that is its own health check. It answers only on the tailnet.

The internal dashboard: four status cards — Attention with four warnings, Backups with every job up and the Longhorn backup 43 hours old in red, Network, Hosts — and below them the consoles grouped as Home, Cluster, Media and Cloud, every dot green
Homepage · 6 Oct 2026

The first evening its Backups card had real data, it showed a job down: the monthly restore drill had been failing since August — Authentik one month, Joplin the next — while both dumps were intact. The drill was racing Postgres’s own start-up. The red in this picture is just as real: one night’s Longhorn backup failed, and the card stays that colour until a good one replaces it.

One of the four warnings never clears, and it is honest: if pve-1 died, the other two nodes could not hold everything it runs. I know, and I chose three small machines anyway.

Backup

Three copies. Nothing can delete them.

Every producer writes the one thing it owns to the NAS. Once a day the vault wakes, pulls the NAS read-only, snapshots it, pushes it to Oracle, and powers itself off. Nothing holds a credential into the vault, and Oracle’s restic server is append-only — its retention runs on the VPS, never from home.

flowchart LR
  lh["Longhorn · Garage"] --> nas["Synology NAS
latest only"] pf["pfSense config"] --> nas vps["VPS services"] --> nas files["Files · Photos"] --> nas nas -- "daily, pulled read-only" --> vault["R730xd vault
30 days + 12 months"] vault -- "daily, append-only" --> oc["Oracle Cloud
cannot be pruned from home"]

A new producer gets both further copies for free by writing to the NAS. A compromised one can overwrite its own latest copy there and nothing further: the vault keeps the old version beside the new, and a pull that would delete more than 500 files stops before it snapshots anything.

Losing any single copy is a non-event. On the first of every month the vault restores a sample from Oracle and compares checksums, rather than trusting that it could.

Recovery

Drilled, then done for real.

cluster restore
~35 min
onto different hardware, from git and object storage
VPS rebuild
~15 min
from nothing
found by the drill
4 + 1
tooling bugs, and one library excluded from every restore

Cluster restored onto different hardware. Drilled on 29 August, then done for real on 1 September — from git and object storage, in about thirty-five minutes. The VPS rebuilds from nothing in about fifteen.

The drill is the point. It found four tooling bugs and one library that had been quietly excluded from every restore for months — none of which a backup report would have shown.

Questions

The ones every rack post gets.

Why no Proxmox cluster?
The hosts are standalone, joined through Proxmox Datacenter Manager and never corosync — no HA at the hypervisor, deliberately. Kubernetes carries the redundancy: one node per physical machine, so the three etcd votes sit in three different chassis, and losing a machine costs a workload rather than the cluster.
What does it draw?
About 65 W for everything on the UPS, averaged over two days with the vault’s daily run included. It was about 165 W while the R730xd ran around the clock; switching it off did more than any tuning I did to it.
Why keep the R730xd at all?
Twelve bays and ECC memory make a good vault, and a vault that is off cannot be reached. It wakes once a day, does its job, and powers itself off. When it does run, a daemon holds the fans at 1680 RPM, about 18 dB under stock.

Tech specs

All of it, in order.

HypervisorProxmox VE on three standalone hosts, joined by Proxmox Datacenter Manager. No corosync, no HA — deliberate.
KubernetesTalos Linux, three control planes, one per physical machine. Flux reconciles from GitHub.
NetworkingpfSense on an XCY X44. Physical segmentation, one NIC per segment. Cilium for CNI, Gateway API and L2 announcement.
StorageSynology DS223, two disks mirrored — the live store for media, files, photos and backups. NFS to the cluster nodes only, one rule per client. Longhorn for app volumes, three replicas.
Off-siteOracle Cloud Free Tier, Ampere A1 — 4 vCPU ARM, 24 GB, 200 GB. Traefik, Authentik, Guacamole, Joplin, Pi-hole, the public dashboard and an append-only restic server.
Offline copyDell R730xd on TrueNAS. 12 × 600 GB SAS in one encrypted RAIDZ3. Wakes once a day, pulls the NAS read-only, snapshots, pushes to Oracle, powers off.
PowerCyberPower VP700ELCD, monitored by NUT. ~65 W total, averaged over two days with the vault’s daily run included — down from ~165 W while the R730xd ran around the clock.
NoiseThe R730xd is off most of the day. While it runs, a daemon takes the fans off iDRAC and holds a 1680 RPM floor, about 18 dB under stock.
SecretsSOPS/age for Kubernetes, Ansible Vault for the VPS. Three keys cannot be lost; nothing else is irreplaceable.

Everything above is in one repository, and inside.merox.dev shows it running.