14 - Backup & Restore (etcd, Velero)
Why this matters
Everything in Kubernetes lives in etcd. If etcd is gone with no backup, your entire cluster state (not workloads' data — the cluster's definition of itself) is gone. This is a CKA exam topic and a real disaster-recovery skill.
Read this first — Definitions & Explanations
Why etcd backup matters
etcd holds API objects. No good etcd backup ≈ you can lose cluster state (Deployments, RBAC, etc.). App data in disks/DBs is a separate concern.
etcd snapshot
Point-in-time backup of etcd. Restores must be done carefully with matching versions and procedures.
Velero
Popular tool for backing up Kubernetes resources and, with plugins, persistent volume data. Useful for namespace/cluster disaster recovery drills.
App backup vs cluster backup
- Cluster/API backup: objects in etcd (+ optionally PVs)
- Application backup: database dumps, object storage, etc.
Production usually needs both.
Test restores
A backup you’ve never restored is a hope, not a plan.
Official docs (read for detail)
Key Concepts
- etcd snapshot save/restore (
etcdctl snapshot save/restore) - What a restore actually does (restores to a new data dir, requires re-pointing etcd)
- Velero: backs up actual Kubernetes objects + can snapshot PV data (via cloud provider snapshot APIs)
- Backup scope decisions: whole cluster vs namespace vs label-selected resources
- Disaster recovery drill mentality: backups are worthless until you've tested a restore
YouTube search terms
- "etcd snapshot backup restore Kubernetes"
- "Velero Kubernetes backup restore tutorial"
- "Kubernetes disaster recovery etcd"
Hands-on lab (on prod-sim)
# etcd snapshot save + restore drill
minikube ssh -p prod-sim
sudo ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 \
--cacert=/var/lib/minikube/certs/etcd/ca.crt \
--cert=/var/lib/minikube/certs/etcd/server.crt \
--key=/var/lib/minikube/certs/etcd/server.key \
snapshot save /tmp/etcd-backup.db
sudo ETCDCTL_API=3 etcdctl --write-out=table snapshot status /tmp/etcd-backup.db
exit
# Create something, "accidentally" delete it, understand what a restore would recover
kubectl create namespace important-stuff
kubectl -n important-stuff create configmap must-not-lose --from-literal=key=value
# (In a real DR drill: stop kube-apiserver, restore the snapshot to a new data-dir,
# update etcd manifest to point at it, restart. Full walkthrough:
# https://kubernetes.io/docs/tasks/administer-cluster/configure-upgrade-etcd/#restoring-an-etcd-cluster)
# Velero (application-level backup, easier to reason about for this exercise)
# Install Velero CLI, then with a local MinIO as the backup target (since no cloud provider here):
# https://velero.io/docs/main/contributions/minio/
# Once installed:
velero backup create important-stuff-backup --include-namespaces important-stuff
velero backup describe important-stuff-backup
kubectl delete namespace important-stuff
velero restore create --from-backup important-stuff-backup
kubectl get configmap -n important-stuff must-not-lose # should be back
Notes
(fill in your own words after watching + labbing)