⚡ ~/naveed k8s
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Phase 2 — Cluster Administration Module 14 of 24 Free & Open Access

Backup & Restore (etcd, Velero)

Complete production curriculum breakdown. Learn core architectural mechanics, study definitions in plain language, practice hands-on labs with the local minikube prod-sim cluster, and test active recall.

14 - Backup & Restore (etcd, Velero)

Why this matters

Everything in Kubernetes lives in etcd. If etcd is gone with no backup, your entire cluster state (not workloads' data — the cluster's definition of itself) is gone. This is a CKA exam topic and a real disaster-recovery skill.

Read this first — Definitions & Explanations

Why etcd backup matters

etcd holds API objects. No good etcd backup ≈ you can lose cluster state (Deployments, RBAC, etc.). App data in disks/DBs is a separate concern.

etcd snapshot

Point-in-time backup of etcd. Restores must be done carefully with matching versions and procedures.

Velero

Popular tool for backing up Kubernetes resources and, with plugins, persistent volume data. Useful for namespace/cluster disaster recovery drills.

App backup vs cluster backup

Test restores

A backup you’ve never restored is a hope, not a plan.

Official docs (read for detail)

Key Concepts

YouTube search terms

Hands-on lab (on prod-sim)

# etcd snapshot save + restore drill
minikube ssh -p prod-sim
  sudo ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 \
    --cacert=/var/lib/minikube/certs/etcd/ca.crt \
    --cert=/var/lib/minikube/certs/etcd/server.crt \
    --key=/var/lib/minikube/certs/etcd/server.key \
    snapshot save /tmp/etcd-backup.db
  sudo ETCDCTL_API=3 etcdctl --write-out=table snapshot status /tmp/etcd-backup.db
  exit

# Create something, "accidentally" delete it, understand what a restore would recover
kubectl create namespace important-stuff
kubectl -n important-stuff create configmap must-not-lose --from-literal=key=value
# (In a real DR drill: stop kube-apiserver, restore the snapshot to a new data-dir,
#  update etcd manifest to point at it, restart. Full walkthrough:
#  https://kubernetes.io/docs/tasks/administer-cluster/configure-upgrade-etcd/#restoring-an-etcd-cluster)

# Velero (application-level backup, easier to reason about for this exercise)
# Install Velero CLI, then with a local MinIO as the backup target (since no cloud provider here):
# https://velero.io/docs/main/contributions/minio/
# Once installed:
velero backup create important-stuff-backup --include-namespaces important-stuff
velero backup describe important-stuff-backup
kubectl delete namespace important-stuff
velero restore create --from-backup important-stuff-backup
kubectl get configmap -n important-stuff must-not-lose   # should be back

Notes

(fill in your own words after watching + labbing)

📋 Self-Assessment Mastery Checklist (4 Competencies)
🧠 Practice Exam Questions (Module 14 MCQs)
⚡ Take Quiz & Save Progress in Tracker

Review these sample exam questions out loud, test your retrieval, and then unlock official scoring in the interactive tracker.

Question 1: etcd backup is critical because etcd holds:
  • A. Cluster state (API objects)
  • B. Only container images
  • C. Only node kernels
  • D. Only Ingress logs
✓ Correct Answer: A (Cluster state (API objects))
Option A ('Cluster state (API objects)') is the standard production architectural best practice.
Question 2: Velero is commonly used for:
  • A. Cluster backup/restore including PVs (with plugins)
  • B. Replacing kube-proxy
  • C. Building container images
  • D. Managing Helm repos only
✓ Correct Answer: A (Cluster backup/restore including PVs (with plugins))
Option A ('Cluster backup/restore including PVs (with plugins)') is the standard production architectural best practice.
Question 3: Restoring etcd incorrectly can:
  • A. Corrupt or mismatch cluster state — handle with care
  • B. Only affect DNS
  • C. Never affect Pods
  • D. Only reset Helm
✓ Correct Answer: A (Corrupt or mismatch cluster state — handle with care)
Option A ('Corrupt or mismatch cluster state — handle with care') is the standard production architectural best practice.
Question 4: A good backup strategy includes:
  • A. Regular tested restores, not just backups
  • B. Backing up once ever
  • C. Only ConfigMaps
  • D. Only worker node disks
✓ Correct Answer: A (Regular tested restores, not just backups)
Option A ('Regular tested restores, not just backups') is the standard production architectural best practice.
Question 5: Application-level backups vs etcd backups:
  • A. Both may be needed depending on stateful apps
  • B. etcd alone always backups PVC file contents everywhere
  • C. App backups replace control plane certs
  • D. Neither is useful
✓ Correct Answer: A (Both may be needed depending on stateful apps)
Option A ('Both may be needed depending on stateful apps') is the standard production architectural best practice.
← Previous Module (13) Troubleshooting & Debugging Next Module (15) → Monitoring & Logging