22 - Autoscaling (HPA, VPA, Cluster Autoscaler, KEDA)
Why this matters
Static replica counts and static node pools waste money or fall over under load. Real prod clusters scale both pods and nodes dynamically based on demand.
Read this first — Definitions & Explanations
HPA (Horizontal Pod Autoscaler)
Scales replica count based on metrics (CPU, memory, or custom metrics). Needs metrics pipeline (often metrics-server for resource metrics).
VPA (Vertical Pod Autoscaler)
Recommends or applies better requests/limits for Pods. Different problem than HPA.
Cluster Autoscaler
Adds/removes nodes when Pods are unschedulable or nodes are underused. Works with cloud node groups / autoscaling groups.
KEDA
Event-driven autoscaling — scale on queue depth, Kafka lag, etc., including scale to zero patterns.
Autoscaling hygiene
Garbage in, garbage out: if requests are wrong, HPA/Cluster Autoscaler behave badly. Set realistic requests first.
Official docs (read for detail)
- Horizontal Pod Autoscaling
- HorizontalPodAutoscaler Walkthrough
- Vertical Pod Autoscaler
- Cluster Autoscaler
- KEDA Documentation
Key Concepts
- HPA (Horizontal Pod Autoscaler): scales replica count based on metrics (CPU/memory by default, custom metrics via adapters)
- VPA (Vertical Pod Autoscaler): adjusts requests/limits automatically (careful: restarts pods to resize)
- Cluster Autoscaler: adds/removes nodes based on unschedulable pods / underutilized nodes
- KEDA: event-driven autoscaling — scale on queue depth, Kafka lag, cron schedule, etc. (not just CPU)
- HPA needs metrics-server (or custom/external metrics API) to function at all
- Scaling behavior tuning: stabilization windows, scale-up/down policies (avoid flapping)
YouTube search terms
- "Kubernetes HPA horizontal pod autoscaler tutorial"
- "Kubernetes VPA vertical pod autoscaler explained"
- "Cluster Autoscaler Kubernetes explained"
- "KEDA event driven autoscaling tutorial"
Hands-on lab (on prod-sim)
# HPA — you already have metrics-server enabled
kubectl create deployment php-apache --image=k8s.gcr.io/hpa-example --requests='cpu=200m'
kubectl expose deployment php-apache --port=80
kubectl autoscale deployment php-apache --cpu-percent=50 --min=1 --max=5
kubectl get hpa -w &
# Generate load and watch it scale up
kubectl run load-generator --image=busybox --restart=Never -- \
sh -c "while true; do wget -q -O- http://php-apache; done"
sleep 60
kubectl get hpa php-apache
kubectl get pods -l app=php-apache
kubectl delete pod load-generator # stop load, watch it scale back down after stabilization window
kill %1
# VPA (install the components first: https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler)
cat <<EOF | kubectl apply -f -
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: php-apache-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: php-apache
updatePolicy:
updateMode: "Auto"
EOF
kubectl describe vpa php-apache-vpa # see recommended requests
# KEDA
helm repo add kedacore https://kedacore.github.io/charts
helm install keda kedacore/keda -n keda --create-namespace
# scale a deployment based on a cron schedule as the simplest possible example
cat <<EOF | kubectl apply -f -
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: cron-scaledobject
spec:
scaleTargetRef:
name: php-apache
minReplicaCount: 1
maxReplicaCount: 5
triggers:
- type: cron
metadata:
timezone: UTC
start: "0 * * * *"
end: "5 * * * *"
desiredReplicas: "3"
EOF
Notes
(fill in your own words after watching + labbing)