↓ Skip to main content

Built to be rebuilt.

Three Proxmox hosts, one Kubernetes cluster, and nothing that answers the internet. All of it comes back from a git repository and three keys in about thirty-five minutes — which is the only reason it is safe to keep changing.

Overview
#

The office corner: a standing desk running monitoring dashboards, and the open rack beside it — R730xd, mini PCs, NAS and UPS

3

Proxmox hosts

standalone, joined by PDM

Talos

kubernetes

3 nodes on 3 machines, run by Flux

~6 TB

pools

ZFS — NVMe, SSD, SAS

~165 W

draw

three hosts, one UPS

0

inbound ports

tunnel + tailnet only

3

backup copies

one off-site, one offline

Next Network

Network
#

No way in.

Not one port is forwarded. Nothing in the rack answers the internet directly, and the two paths that do exist both dial outward: a mesh VPN, and a tunnel that opens itself.

pfSenseFirewall and gateway on an XCY X44. Physical segmentation — one NIC per segment, not VLAN tags.
TailscaleA subnet router puts the LAN on the tailnet, so the rack and the cloud VPS share one flat mesh. This is the way in.
Cloudflare TunnelOutbound-only, three routes. A router that does not exist on a port cannot be reached through it.
UnboundRecursive resolution, no upstream forwarder, with overrides for the internal names.

Grafana, Internet row: 935 Mb/s down, 660 Mb/s up, 21 ms round trip, and five minutes of flat throughput and latency history

Next Compute

Compute
#

Three machines. Three jobs.

Standalone hosts, joined through Proxmox Datacenter Manager. One Kubernetes node each, on three different chassis.

Beelink GTi13 Ultra

CPUCore i9-13900HK · 20 threads
Memory64 GB DDR5
Nodekubernetes-1 · 14 cores / 32 GB
WhyIris Xe passed through — every transcode happens here

Dell PowerEdge R730xd

CPUXeon E5-2630 v4 · 10C/20T
Memory251 GB DDR4 ECC
Nodekubernetes-2 · NFS, Garage S3, Nextcloud
WhyTwelve SAS bays, and the backup hub everything drains into

Dell OptiPlex 3050

CPUCore i5-6500T · 4 cores
Memory32 GB DDR4
Nodekubernetes-3 · one Longhorn replica
Why960 GB D3-S4510 — power-loss protection, the best etcd disk here

Proxmox Datacenter Manager, resource usage across the three hosts: CPU 12%, 5.48 of 44 cores with 37 allocated; memory 27%, 93 of 345 GiB; storage 4%, 367 GiB of 9 TiB
Proxmox Datacenter Manager · 11 Sep 2026

Next Storage

Storage
#

Six terabytes, mostly asleep.

Twelve SAS disks hold the library. They park whenever nothing is reading them — forty to fifty watts that would otherwise be spent spinning platters for nobody, in a room somebody sleeps in.

PoolHostLayoutUsed
mediapve-212 × SAS in two raidz2 vdevs1.31 TB of ~4.2 TB
rpoolpve-2SSD mirror — everything touched daily159 GB of 861 GB
cluster-storagepve-1Single NVMe, one Longhorn replica105 GB of 899 GB
Grafana storage panels: every Proxmox storage under 17% full, and SSD wear from 34% on one pve-1 disk down to zero on the rest
Tip

Waking the disks used to stall etcd’s write-ahead log on the shared HBA. The fix was not more hardware — it was raising the election timeout from one second to five, and parking the pool outside the backup window. The whole story is in Spinning Down SAS Disks.

Next Kubernetes

Kubernetes
#

git push is the deploy.

Talos has no shell and no package manager — it is configured by API and by nothing else. Flux watches the repository and makes the cluster match it. There is no third way to change anything, which is what makes the rebuild time believable.

kubernetes/apps/
├── default/         jellyfin, immich, the *arr stack, n8n
├── kube-system/     Cilium, CoreDNS, NFS CSI, Intel GPU plugin
├── network/         gateway, Cloudflare tunnel, netboot.xyz
├── observability/   Prometheus, Grafana, Loki, Alloy
└── storage/         Longhorn
Headlamp’s map of the cluster: kube-system, default, longhorn-system and observability collapsed; flux-system open on its five controllers; network on cloudflare-dns, cloudflare-tunnel, netboot-xyz and k8s-gateway; cert-manager on its three deployments
Headlamp · 11 Sep 2026

Grafana, Kubernetes row: cluster CPU at 6%, memory at 19%, and fourteen app tiles, every one green and Up

Next Services

Services
#

What it actually runs.

Disk-bound things in the cluster, identity and remote access on the VPS. Everything with data behind it is declared in the repo.

Media

  • Jellyfin: 4K, hardware transcoded
  • Jellyseerr: requests, the one route that faces out
  • Radarr · Sonarr: film and TV automation
  • Prowlarr · qBittorrent: indexing and download, VPN-routed

Files & photos

  • Immich: photo library, own Postgres
  • Nextcloud: files and office, AIO on its own VM
  • Collabora: documents edited in the browser
  • Joplin: notes sync

Platform

  • Talos + Flux: the cluster and what reconciles it
  • Longhorn: block storage, three replicas
  • Cilium: eBPF CNI, Gateway API, L2
  • cert-manager: certificates, renewed without asking

Edge & identity

  • Traefik: reverse proxy on the VPS
  • Authentik: single sign-on
  • Guacamole: remote desktop in a tab
  • Pi-hole + Unbound: filtering DNS, own resolver

Watching it

  • Prometheus: metrics, hosts and cluster
  • Grafana: dashboards
  • Loki + Alloy: logs
  • Alertmanager: straight to Telegram

Odds

  • n8n: one workflow, the daily news digest
  • Headlamp: Kubernetes, when a UI is faster
  • Portainer: the VPS containers
  • netboot.xyz: installers over the network

Portainer on the VPS: the Authentik server, worker, Postgres and Redis, Guacamole, the public Homepage, Pi-hole, Joplin and its database, Portainer itself, the restic REST server and Traefik; every container healthy or running
Portainer · the VPS · 11 Sep 2026

Next Alerting

Alerting
#

Silence is the alarm.

Alertmanager, Flux and Proxmox go straight to Telegram — critical and warning only, tagged by layer. Every scheduled job also pings healthchecks.io on its way out, so a job that stops running raises its own hand. Nobody has to notice.

The whole Homelab Overview board in Grafana: zero alerts, on mains, every host, node and pool online; backups hours old; fourteen apps up; then the internet, host, storage and UPS rows, the estate drawing 164 W
Homelab Overview · 11 Sep 2026

Next Backup

Backup
#

Three copies. It can delete none of them.

The machine holding the data can add to both off-site targets and remove from neither. Oracle runs an append-only restic server. The NAS pulls with a read-only key and leaves no credential behind it. Retention runs on each target, not on the source.

flowchart LR
  lh["Longhorn · cluster"] -- nightly --> hub["R730xd
Garage S3"] vps["Oracle VPS"] -- nightly --> hub pf["pfSense config"] --> hub hub -- "weekly, pulled" --> nas["Synology
asleep between wake windows"] hub -- "nightly, append-only" --> oc["Oracle Cloud
cannot be pruned from home"]
Grafana, Backups & certificates row: last local backup 8.1 hours ago, off-site 7.8 hours, VM image 2.1 days, certificates renewed 5.8 days ago

Losing any single copy is a non-event. That is the whole point of the arrangement, and it is checked by restoring a five-percent sample every week rather than by trusting it.

Next Recovery

Recovery
#

Drilled, then done for real.

~35 min

cluster restore

onto different hardware, from git and object storage

~15 min

VPS rebuild

from nothing

4 + 1

found by the drill

tooling bugs, and one library excluded from every restore

Cluster restored onto different hardware. Drilled on 29 August, then done for real on 1 September — from git and object storage, in about thirty-five minutes. The VPS rebuilds from nothing in about fifteen.

The drill is the point. It found four tooling bugs and one library that had been quietly excluded from every restore for months — none of which a backup report would have shown.

Next Questions

Questions
#

The ones every rack post gets.

Why no Proxmox cluster?
The hosts are standalone, joined through Proxmox Datacenter Manager and never corosync — no HA at the hypervisor, deliberately. Kubernetes carries the redundancy: one node per physical machine, so the three etcd votes sit in three different chassis, and losing a machine costs a workload rather than the cluster.
What does it draw?
About 165 W for the whole estate, roughly 110 W of it the R730xd — around 119 kWh a month, on a CyberPower VP700ELCD watched by NUT. Parking the SAS pool saves another forty to fifty watts whenever nothing reads it.
Is the R730xd loud?
Not any more. A daemon takes the fans off iDRAC and holds a 1680 RPM floor, about 18 dB under stock — it shares a room somebody sleeps in.

Next Tech specs

Tech specs
#

All of it, in order.

HypervisorProxmox VE on three standalone hosts, joined by Proxmox Datacenter Manager. No corosync, no HA — deliberate.
KubernetesTalos Linux, three control planes, one per physical machine. Flux reconciles from GitHub.
NetworkingpfSense on an XCY X44. Physical segmentation, one NIC per segment. Cilium for CNI, Gateway API and L2 announcement.
StorageZFS. media = 12 × SAS in 2 × raidz2, parked when idle. rpool = SSD mirror. cluster-storage = NVMe. NFS to the cluster nodes only, ACL per client IP.
Off-siteOracle Cloud Free Tier, Ampere A1 — 4 vCPU ARM, 24 GB, 200 GB. Traefik, Authentik, Guacamole, Joplin, Pi-hole, the public dashboard and an append-only restic server.
Cold copySynology DS223+. Pulls weekly over a read-only rrsync key, asleep the rest of the week.
PowerCyberPower VP700ELCD, monitored by NUT. ~165 W total, ~110 W of it the R730xd. Roughly 119 kWh a month.
NoiseFans off iDRAC — a daemon holds a 1680 RPM floor, about 18 dB under stock.
SecretsSOPS/age for Kubernetes, Ansible Vault for the VPS. Three keys cannot be lost; nothing else is irreplaceable.

Everything above is in one repository, and inside.merox.dev shows it running.