Single-node Kubernetes cluster that ran at home, managed with Flux.
One Talos Linux node, everything in this repository, nothing configured by hand. If the machine died I wanted to rebuild it from a clean disk and this Git history, which is the whole reason it is laid out this way.
Decommissioned on 2026-09-29. The cluster no longer runs, and this repository is kept
as a historical record. Selected services (Umami, Cairn, the Gatus checks, the DNS
fallback and slskd) are being rebuilt on a VPS platform in the private fobiat/vps-ops
repository, tracked there as issue #43. Changes here are limited to documentation and
the decommission itself.
Version 3 ran on Talos with real workloads: monitoring (kube-prometheus-stack, alerting
to Discord), CNPG-backed Postgres, self-hosted GitHub Actions runners for several private
repos, Umami analytics, Gatus health checks, and
Cairn, a UK live-incident lookup service deployed
straight from its own repository via Flux and served internally at cairn.lab.fobiat.dev.
Everything with state was backed up nightly and the restore path is written down, though the restic repository lived on the same machine. See Backups.
The rest of this page describes the cluster as it ran before the decommission.
The history here goes back to January 2021. Version 1 was Kubernetes on a Dell PowerEdge
and lived in this repository until electricity prices made a full rack unappealing.
Version 2 was k3s on a Dell Optiplex and lives in
-DEPRECIATED-k3s-homelab at the v2
tag. Version 3 starts here, on Talos.
| Node | Dell Optiplex 3050 SFF, 4 cores, 8 threads, 32GB |
| Ran as | Talos VM on Hyper-V, external virtual switch |
| Storage | NVMe boot, second SSD for persistent volumes |
| Was planned | Minisforum MS-03 class, at which point the Optiplex would have become the spare |
One node meant no high availability. Upgrades took the cluster down, because there was nowhere to drain to. The docs say so wherever it matters rather than pretending otherwise.
| Tool | Job |
|---|---|
| Talos Linux | The OS. Immutable, no SSH, configured by API |
| Flux | Reconciles this repository into the cluster |
| Cilium | CNI, kube-proxy replacement, and the Gateway API implementation |
| cert-manager | Wildcard certificates over DNS-01 |
| external-dns | DNS records from HTTPRoutes |
| SOPS and age | Secrets, encrypted in this repository |
| Tailscale | How I reach the node and the cluster |
| VolSync | Restic backups of every persistent volume |
| Gatus | Health checks, and the status page |
| Renovate | Keeps everything current |
Nothing is public by default. Services live on *.lab.fobiat.dev and resolve only over
Tailscale or the local network. Anything that genuinely needs to be reachable by someone
without my tailnet gets its own name and goes through a Cloudflare Tunnel, one service at
a time, as a deliberate decision rather than a default.
Exactly one thing was public before the decommission: insights.fobiat.dev, which is Umami's collector. The
route matches two exact paths, /script.js and /api/send, and nothing else. A request
for / returns 404 because there is no rule for it, which is the intended surface rather
than a fault. The dashboard itself stays inside the tailnet.
Tailscale runs as a Talos system extension rather than in the cluster, so the node is reachable before Kubernetes starts. That matters on the day the cluster is the thing that is broken.
Private repos run their GitHub Actions on this cluster rather than on GitHub's hosted runners. One actions-runner-controller scale set per repo, each scaling from zero, so an idle repo costs one small listener pod and nothing else.
They run in containerMode: kubernetes, which means every job needs a
container:. Talos has no Docker daemon and cannot load workload kernel
modules, so the usual docker-in-docker mode is not available here. See
ADR 0015.
maxRunners is 1 per repo. The real ceiling is the namespace ResourceQuota,
which is what actually stops several repos building at once from taking the node
down. scripts/sync-actions-runners.sh adds a scale set for any private repo
that has CI and does not have one yet.
This repository is deliberately not one of them. It is public, and a self-hosted runner on a public repo lets a fork's pull request run its own code on the cluster. home-ops uses GitHub-hosted runners for that reason alone.
A side effect worth knowing: because self-hosted runners do not consume hosted minutes, they keep working when the account's hosted minutes are not available.
Details, including how to add a repo and what the shared GitHub App has to have access to, are in the runbook.
Three things get backed up, on their own schedules, through restic:
| What | When (UTC) | Kept |
|---|---|---|
| etcd snapshot | 03:15 | 14 |
| Talos machine config | 03:30 | 30 |
| Persistent volumes | 04:00, 04:10, 04:40 | 7 daily, 5 weekly, 6 monthly |
The order matters. Both cluster-level jobs finish well before VolSync copies the volume they write into, so one restic snapshot holds a consistent set.
The honest caveat: the restic repository is a second SSD on the same machine. That
survives a bad upgrade, a corrupted etcd, or a fat-fingered kubectl delete. It does not
survive the machine. Moving the repository to object storage is the open piece of work,
and until it lands this is a rollback mechanism rather than disaster recovery.
Restoring is the part people skip, so it has its own runbook: Restore for the whole node, and Etcd snapshot restore for the control plane on its own.
talos/ Machine configuration, encrypted secrets, Image Factory schematic
bootstrap/ Cilium and Flux, installed once before GitOps takes over
kubernetes/
flux/ Root Kustomization and cluster-wide variables
components/ Shared Kustomize components
apps/ Everything else, by namespace
docs/ Runbooks, decision records, and how to rebuild this from nothing
Published at fobiat.github.io/home-ops, and
was mirrored inside the tailnet at docs.lab.fobiat.dev until the decommission. The source is in
docs/. Worth reading first:
- Bootstrap, bare disk to running cluster
- Restore, what to do when the disk is gone
- Decision records, why things are the way they are, including the two choices that go against what most people do
This borrows heavily. Worth your time if you are building something similar:
- onedr0p/home-ops and cluster-template
- buroa/k8s-gitops, for the namespace component trick
- carpenike/k8s-gitops, for the flux-local diff workflow
- home-operations, particularly
tuppr, which is the only thing I found that properly handles upgrading a cluster with one node in it - kubesearch.dev, for finding who else runs a given chart
The k8s-at-home organisation was archived in May 2026. The community moved to the Home Operations Discord.
MIT. Take whatever is useful.