Home
/
All 24 Modules Curriculum
24 Production Modules
4 Structured Phases
Free & Open Access
Complete 24-Module Kubernetes Learning Curriculum
A structured learning path engineered for production reality. No random video hopping. Each module pairs clear theoretical definitions with hands-on CLI labs on your local 3-node minikube cluster and 5-question mastery quizzes.
💻 Prerequisites: Spin Up Your Local 3-Node minikube Lab (prod-sim)
Every hands-on command throughout the 24 modules runs seamlessly on a local multi-node minikube profile called prod-sim. Run this once in your terminal:
minikube start -p prod-sim --nodes 3 --driver docker
minikube addons enable metrics-server -p prod-sim
kubectl get nodes -o wide
Kubernetes Architecture & Core Concepts
I can draw the architecture diagram from memory (control plane + node components) I can explain why etcd is only ever touched by the API server
Pods & Workload Controllers
I can explain when to use Deployment vs StatefulSet vs DaemonSet I've done a rolling update and a rollback
Services & Cluster Networking Basics
I can explain ClusterIP vs NodePort vs LoadBalancer and when to use each I've traced a Service down to its EndpointSlice and iptables rule
ConfigMaps & Secrets
I can explain why Secrets are not "secure" by default without etcd encryption I've mounted a ConfigMap both as env vars and as a volume
Storage: Volumes, PV, PVC, StorageClasses
I can explain the PV/PVC/StorageClass relationship end to end I've proven data survives a pod restart via a PVC
Scheduling: Affinity, Taints, Resources, PDBs
I can explain requests vs limits and why memory limits can OOMKill a pod I've tainted a node and scheduled a pod onto it with a toleration
Ingress & Ingress Controllers
I understand an Ingress resource does nothing without a controller I've routed two different hostnames to two different services through one Ingress
Namespaces, RBAC & Security Basics
I can explain why namespaces are isolation, not a hard security boundary I've created a Role + RoleBinding scoped to one namespace and one verb set
Helm & Package Management
I've installed, upgraded, and rolled back a real chart from a public repo I've written and installed a minimal custom chart with `helm create`
Control Plane Deep Dive
I can explain Raft/quorum and why etcd clusters use odd numbers I've queried etcd directly with etcdctl and found a pod's key
Installing & Upgrading Clusters (kubeadm)
I can explain the drain -> upgrade -> uncordon sequence and why order matters I understand why control plane is upgraded before worker nodes
Networking Deep Dive (CNI, NetworkPolicy)
I know which CNI my cluster uses and whether it enforces NetworkPolicy I've implemented default-deny + a specific allow rule and proven both states
Troubleshooting & Debugging
I can diagnose each of the 5 classic failure states above without looking anything up I've used `kubectl debug` on a pod with no shell
Backup & Restore (etcd, Velero)
I've taken and validated an etcd snapshot (`snapshot status`) I understand a restore requires a new data-dir + restarting etcd pointed at it
Monitoring & Logging
I have Prometheus + Grafana running and can view a dashboard I've written at least 3 PromQL queries from scratch
Cluster Hardening & CIS Benchmarks
I've run kube-bench against a real cluster and read through the FAIL findings I've enabled audit logging and confirmed events are being recorded
Pod & Workload Security
I've enforced `restricted` PSA on a namespace and proven a privileged pod gets rejected I've written a pod spec that passes `restricted` (non-root, no priv-escalation, capabilities dropped, seccomp)
Supply Chain Security & Admission Control
I've scanned an image with Trivy and read real CVE findings I've written and enforced a Kyverno policy that blocks a real anti-pattern
Runtime & Network Security
I've installed Falco and triggered at least 2 real runtime alerts myself I've implemented an egress-deny-by-default NetworkPolicy with a DNS exception
CRDs & Custom Controllers/Operators
I've done the full kubebuilder CronJob tutorial from the book, end to end I've written my own minimal CRD + controller that reconciles a real Deployment
GitOps (ArgoCD/Flux)
I've deployed a real app via ArgoCD from a git repo I've proven drift detection: manual change gets flagged/reverted
Autoscaling (HPA, VPA, Cluster Autoscaler, KEDA)
I've watched HPA scale a deployment up under load and back down when load stops I understand why VPA and HPA on CPU/memory together is risky (they can fight each other)
Service Mesh & Multi-Cluster
I've deployed a mesh-enabled app and confirmed every pod has a sidecar I've done a weighted canary traffic split and watched requests distribute
Certification Exam Prep (CKA/CKAD/CKS)
I've registered for killer.sh practice sessions and completed at least one full timed run I can generate 90% of my YAML via imperative commands + `--dry-run=client -o yaml`, not typing from scratch
// ECOSYSTEM CROSS-LINKING & KUBERNETES KNOWLEDGE GRAPH
Kubernetes Complete Curriculum Interlinking Matrix
Explore production failure runbooks, deep-dive architectural blog posts, interactive troubleshooting games, and interview prep.