JC_

← All series

15 PARTS

Running Production Solo

My k3s high-availability journey — every incident, diagnosis, and fix from building out a self-hosted production cluster alone.

01

The Moment a Single Node Couldn’t Keep Up: It Started With a Capacity Report

k3s Series 1 — An automated risk report turned my “it’s probably fine” instinct into a red Critical label A single node is a single point of failure — the fi...

Aug 6

02

From 1 Node to 3: The Full Story of Building Out Longhorn

k3s Series 2 — Every volume in the cluster said “degraded.” Not one of them was actually broken. Three connected nodes, finally enough room for every replica...

Aug 8

03

Longhorn PVC Operations: Shrinking and Growing Storage Without Losing Data

k3s Series 3 — Kubernetes won’t let you shrink a PVC. Once you understand why, the same trick fixes two completely different problems. You can’t resize a PVC...

Aug 9

04

RWO→RWX: Down the Rabbit Hole to a Corrupted Instance-Manager

k3s Series 4 — The error message said “Multi-Attach.” The actual problem was a container that had quietly stopped being able to write to its own disk. What l...

Aug 11

05

From SQLite to a 3-Server etcd Cluster: The Full HA Upgrade

k3s Series 5 — There’s a one-way door in this migration. Cross it, and there’s no going back to the way things were. Three nodes were already running. Only o...

Aug 12

06

Off-Site Backup: etcd Snapshots and Longhorn’s Double Insurance

k3s Series 6 — HA answers “what if one machine dies.” This post is about the question HA can’t answer: what if all three die at once? The cluster’s brain and...

Aug 13

07

Assume All 3 Machines Die: A Full Disaster Recovery Drill

k3s Series 7 — A backup you’ve never restored from isn’t a backup. It’s a hypothesis. Three empty machines, two backup files, and a procedure I need to...

Aug 15

08

Prometheus + External Grafana: Wiring Up Monitoring for Production

k3s Series 8 — HA and backups don’t matter if you’re the last to know something’s wrong. This is the plumbing that closes that gap. The metrics were always...

Aug 15

09

Alerting Isn't Just Adding Rules: The PromQL Traps I Hit

k3s Series 9 — A PromQL query reads like a sentence. It doesn’t behave like one. Every rule here looked right on the first read. Most of them weren’t. This is...

Aug 17

10

One IP to Rule the Control Plane: Adding a VIP to k3s HA

k3s Series 10 — The control plane finally had no single point of failure. My kubectl config still pointed at one. Three servers, one address that doesn’t care...

Aug 17

11

Multiple Nodes Isn't the Same as Highly Available: A Full HA Audit

k3s Series 11 — I had 3 servers, replicated storage, a VIP, and no idea whether any of it would actually survive losing a node under real load. So I did the...

Aug 18

12

Your k3s HA Might Be Fake: How a Traditional HDD Quietly Undermines etcd

k3s Series 12–3 servers, verified quorum, a VIP, audited N-1 capacity. None of it matters if the disk underneath etcd can’t keep up. Every fix in this series...

Aug 18

13

When You Can't Replace the HDD: Buying etcd More Time

k3s Series #13 — I assumed etcd and Longhorn were fighting over the same disk. They weren't even on the same physical drive.

Aug 19

14

Containerd Filled the System Disk: A Two-Phase, Minimal-Downtime Migration

k3s Series #14 — Three identical machines. One small decision at join time. Three completely different maintenance bills, years later.

Aug 20

15

The Discipline Behind All of This

k3s Series #15 — Fourteen posts of incidents. One habit shows up in almost every single one of them.

Aug 21