15 PARTS
Running Production Solo
My k3s high-availability journey — every incident, diagnosis, and fix from building out a self-hosted production cluster alone.
01
The Moment a Single Node Couldn’t Keep Up: It Started With a Capacity Report
k3s Series 1 — An automated risk report turned my “it’s probably fine” instinct into a red Critical label A single node is a single point of failure — the fi...
Aug 6
02
From 1 Node to 3: The Full Story of Building Out Longhorn
k3s Series 2 — Every volume in the cluster said “degraded.” Not one of them was actually broken. Three connected nodes, finally enough room for every replica...
Aug 8
03
Longhorn PVC Operations: Shrinking and Growing Storage Without Losing Data
k3s Series 3 — Kubernetes won’t let you shrink a PVC. Once you understand why, the same trick fixes two completely different problems. You can’t resize a PVC...
Aug 9
04
RWO→RWX: Down the Rabbit Hole to a Corrupted Instance-Manager
k3s Series 4 — The error message said “Multi-Attach.” The actual problem was a container that had quietly stopped being able to write to its own disk. What l...
Aug 11
05
From SQLite to a 3-Server etcd Cluster: The Full HA Upgrade
k3s Series 5 — There’s a one-way door in this migration. Cross it, and there’s no going back to the way things were. Three nodes were already running. Only o...
Aug 12
06
Off-Site Backup: etcd Snapshots and Longhorn’s Double Insurance
k3s Series 6 — HA answers “what if one machine dies.” This post is about the question HA can’t answer: what if all three die at once? The cluster’s brain and...
Aug 13
07
Assume All 3 Machines Die: A Full Disaster Recovery Drill
k3s Series 7 — A backup you’ve never restored from isn’t a backup. It’s a hypothesis. Three empty machines, two backup files, and a procedure I need to...
Aug 15
08
Prometheus + External Grafana: Wiring Up Monitoring for Production
k3s Series 8 — HA and backups don’t matter if you’re the last to know something’s wrong. This is the plumbing that closes that gap. The metrics were always...
Aug 15
09
Alerting Isn't Just Adding Rules: The PromQL Traps I Hit
k3s Series 9 — A PromQL query reads like a sentence. It doesn’t behave like one. Every rule here looked right on the first read. Most of them weren’t. This is...
Aug 17
10
One IP to Rule the Control Plane: Adding a VIP to k3s HA
k3s Series 10 — The control plane finally had no single point of failure. My kubectl config still pointed at one. Three servers, one address that doesn’t care...
Aug 17
11
Multiple Nodes Isn't the Same as Highly Available: A Full HA Audit
k3s Series 11 — I had 3 servers, replicated storage, a VIP, and no idea whether any of it would actually survive losing a node under real load. So I did the...
Aug 18
12
Your k3s HA Might Be Fake: How a Traditional HDD Quietly Undermines etcd
k3s Series 12–3 servers, verified quorum, a VIP, audited N-1 capacity. None of it matters if the disk underneath etcd can’t keep up. Every fix in this series...
Aug 18
13
When You Can't Replace the HDD: Buying etcd More Time
k3s Series #13 — I assumed etcd and Longhorn were fighting over the same disk. They weren't even on the same physical drive.
Aug 19
14
Containerd Filled the System Disk: A Two-Phase, Minimal-Downtime Migration
k3s Series #14 — Three identical machines. One small decision at join time. Three completely different maintenance bills, years later.
Aug 20
15
The Discipline Behind All of This
k3s Series #15 — Fourteen posts of incidents. One habit shows up in almost every single one of them.
Aug 21