Running Production Solo: My k3s High-Availability Journey — Series Overview

Running Production Solo: My k3s High-Availability Journey — Series Overview

Originally published on Medium

Everything I broke, diagnosed, and eventually fixed while self-hosting a k3s production cluster — alone.

I self-host a k3s production environment on 3 bare Ubuntu 24.04 machines — 180+ pods, a dozen-plus namespaces, stateful services like Postgres/Mongo/Redis, Longhorn for storage, Synology NAS for off-site backup, and an external Grafana for monitoring. There's no team behind this. Just me, and whatever mistakes I was willing to make and then write down.

This series is the full, chronological record of that build: what looked fine and wasn't, what I broke on purpose to learn how it failed, and what I had to fix twice because the first fix was wrong. Every post follows the same shape — the problem, the exact commands and error messages, and what I'd tell someone about to hit the same wall.

Below is every part of the series, in the order things actually happened.

Part 1 — The Moment a Single Node Couldn't Keep Up: It Started With a Capacity Report

A single-node cluster that felt fine — until a capacity report flagged Critical-level OOM risk. Quantifying the danger, then hardening the kubelet.

Part 2 — From 1 Node to 3: The Full Story of Building Out Longhorn

Every Longhorn volume was "degraded" because there was only one node to put replicas on. Scaling to 3 nodes, and the three traps along the way.

Part 3 — Longhorn PVC Operations: Shrinking and Growing Storage Without Losing Data

Kubernetes won't let you shrink a PVC. Two unrelated incidents, one shared fix: --cascade=orphan.

Part 4 — RWO→RWX: Down the Rabbit Hole to a Corrupted Instance-Manager

A textbook Multi-Attach error, until the known fix stopped working. Five layers down: a corrupted instance-manager, debugged from inside a distroless container.

Part 5 — From SQLite to a 3-Server etcd Cluster: The Full HA Upgrade

Three nodes were already running — only one was actually in charge. The irreversible migration from single-server SQLite to full 3-server etcd HA.

Part 6 — Off-Site Backup: etcd Snapshots and Longhorn's Double Insurance

HA survives losing one node. This closes the gap HA can't: two independent backup legs for etcd and Longhorn, shipped off-site.

Part 7 — Assume All 3 Machines Die: A Full Disaster Recovery Drill

A backup you've never restored from is a hypothesis, not a guarantee. The documented — but honestly not yet fire-drilled — full recovery procedure.

Part 8 — Prometheus + External Grafana: Wiring Up Monitoring for Production

HA and backups don't matter if you're the last to know something's wrong. Wiring Prometheus into k3s and pointing it at an existing external Grafana.

Part 9 — Alerting Isn't Just Adding Rules: The PromQL Traps I Hit

A PromQL query reads like a sentence but doesn't behave like one. Six alert rules that looked correct and fired anyway — or worse, stayed silent.

Part 10 — One IP to Rule the Control Plane: Adding a VIP to k3s HA

The cluster survived losing a node. My kubectl config didn't — it still pointed at one. Adding a VIP with kube-vip closed that gap.

Part 11 — Multiple Nodes Isn't the Same as Highly Available: A Full HA Audit

3 nodes and "highly available" aren't the same claim. Running the actual N-1 math, checking replica placement, and auditing hidden single points.

Part 12 — Your k3s HA Might Be Fake: How a Traditional HDD Quietly Undermines etcd

3 servers, VIP, audited N-1 capacity — none of it checks disk latency. A read-only test found etcd's fsync averaging 14ms against a 10ms ceiling.

Part 13 — When You Can't Replace the HDD: Buying etcd More Time

etcd and Longhorn weren't fighting over the same disk — they weren't even on the same drive. What actually helps when you can't replace an HDD yet.

Part 14 — Containerd Filled the System Disk: A Two-Phase, Minimal-Downtime Migration

Same batch of 3 nodes, one small decision at join time: 13% vs 65% vs 80% disk usage, years later. A two-phase, minimal-downtime fix.

Part 15 — The Discipline Behind All of This

Fourteen posts of incidents. Looking back, one habit — trust the number, not the vibe — shows up in almost every single one of them. The series closer.

More parts will be added to this list as they're published — check back, or follow along for updates.

Site d'Outils propose une collection d'outils en ligne gratuits et faciles à utiliser, incluant calculatrice, convertisseur d'unités, liste de tâches, générateur de mots de passe, générateur de QR code, minuteur, chronomètre, minuteur pomodoro et horloge. Améliorez votre productivité et simplifiez vos tâches quotidiennes.