Platforms that stay up.
I design, build and run the Linux, virtualisation and Kubernetes platforms other work depends on. Here is what I can take on, and three places I already have.
What I offer
-
Linux fleets
Provisioning, patching and monitoring, from a handful of hosts to the more than 250 I look after in HPC today.
-
Virtualisation
Proxmox clusters planned around failure domains, with storage and networking that survive losing a host.
-
Kubernetes and GitOps
Talos, Flux and Terraform: clusters declared in git and rebuilt from it, with updates validated before they land.
-
Backups and recovery
Backups that have been restored from at least once, with the restore written down before it is needed.
-
Hosting and automation
Billing, provisioning and DNS wired together so an order becomes a working service with no human step.
Selected work
Industrial SCADA platform
Redundant control platform, re-architected to fit the site it landed on.
Two Ignition gateways run as a redundant pair, so the control system keeps running when either one goes down. Under them, SQL Server log shipping keeps a standby copy of the historian current on a second server, built to work on the standalone Windows servers the site already had. It was handed over with a runbook for the team that operates it.
Ignition · SQL Server · Windows Server · PowerShell
Dolphost
A hosting business, from the billing page down to the servers.
I ran a web hosting company selling shared hosting, VPS and domains to paying customers. WHMCS and cPanel were wired together so a paid order became a working account with no human step in the middle, and I ran the fleet it landed on. Billing, provisioning and uptime were all mine, which made the failure modes commercial as well as technical.
WHMCS · cPanel · LiteSpeed · MySQL
Homelab platform
A production platform in three machines, declared in git.
Three Proxmox hosts run one Talos Kubernetes node each, so the three etcd votes sit in three separate chassis. Flux reconciles the cluster from a public repository and there is no second way to change anything, which is what makes the rebuild believable. Terraform provisions it, Renovate opens the updates, and CI validates them before Flux is allowed to see them.
Proxmox · Talos · Flux · Terraform
Working together
Available for infrastructure work.
Linux fleets, virtualisation, Kubernetes, backups that have been restored from at least once — the kind of platform someone notices when it stops. Tell me what you are running and what is going wrong with it.