⚡ ~/naveed k8s
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Phase 4 — Advanced / Mastery Module 22 of 24 Free & Open Access

Autoscaling (HPA, VPA, Cluster Autoscaler, KEDA)

Complete production curriculum breakdown. Learn core architectural mechanics, study definitions in plain language, practice hands-on labs with the local minikube prod-sim cluster, and test active recall.

22 - Autoscaling (HPA, VPA, Cluster Autoscaler, KEDA)

Why this matters

Static replica counts and static node pools waste money or fall over under load. Real prod clusters scale both pods and nodes dynamically based on demand.

Read this first — Definitions & Explanations

HPA (Horizontal Pod Autoscaler)

Scales replica count based on metrics (CPU, memory, or custom metrics). Needs metrics pipeline (often metrics-server for resource metrics).

VPA (Vertical Pod Autoscaler)

Recommends or applies better requests/limits for Pods. Different problem than HPA.

Cluster Autoscaler

Adds/removes nodes when Pods are unschedulable or nodes are underused. Works with cloud node groups / autoscaling groups.

KEDA

Event-driven autoscaling — scale on queue depth, Kafka lag, etc., including scale to zero patterns.

Autoscaling hygiene

Garbage in, garbage out: if requests are wrong, HPA/Cluster Autoscaler behave badly. Set realistic requests first.

Official docs (read for detail)

Key Concepts

YouTube search terms

Hands-on lab (on prod-sim)

# HPA — you already have metrics-server enabled
kubectl create deployment php-apache --image=k8s.gcr.io/hpa-example --requests='cpu=200m'
kubectl expose deployment php-apache --port=80
kubectl autoscale deployment php-apache --cpu-percent=50 --min=1 --max=5
kubectl get hpa -w &

# Generate load and watch it scale up
kubectl run load-generator --image=busybox --restart=Never -- \
  sh -c "while true; do wget -q -O- http://php-apache; done"
sleep 60
kubectl get hpa php-apache
kubectl get pods -l app=php-apache
kubectl delete pod load-generator   # stop load, watch it scale back down after stabilization window
kill %1

# VPA (install the components first: https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler)
cat <<EOF | kubectl apply -f -
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: php-apache-vpa
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: php-apache
  updatePolicy:
    updateMode: "Auto"
EOF
kubectl describe vpa php-apache-vpa   # see recommended requests

# KEDA
helm repo add kedacore https://kedacore.github.io/charts
helm install keda kedacore/keda -n keda --create-namespace
# scale a deployment based on a cron schedule as the simplest possible example
cat <<EOF | kubectl apply -f -
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: cron-scaledobject
spec:
  scaleTargetRef:
    name: php-apache
  minReplicaCount: 1
  maxReplicaCount: 5
  triggers:
  - type: cron
    metadata:
      timezone: UTC
      start: "0 * * * *"
      end: "5 * * * *"
      desiredReplicas: "3"
EOF

Notes

(fill in your own words after watching + labbing)

📋 Self-Assessment Mastery Checklist (4 Competencies)
🧠 Practice Exam Questions (Module 22 MCQs)
⚡ Take Quiz & Save Progress in Tracker

Review these sample exam questions out loud, test your retrieval, and then unlock official scoring in the interactive tracker.

Question 1: HPA scales:
  • A. Replicas based on metrics (e.g. CPU/custom)
  • B. Node count only
  • C. etcd size
  • D. PVC capacity
✓ Correct Answer: A (Replicas based on metrics (e.g. CPU/custom))
Option A ('Replicas based on metrics (e.g. CPU/custom)') is the standard production architectural best practice.
Question 2: Cluster Autoscaler scales:
  • A. Nodes in the node pool/group
  • B. Pods inside a Deployment only
  • C. Ingress rules
  • D. Helm releases
✓ Correct Answer: A (Nodes in the node pool/group)
Option A ('Nodes in the node pool/group') is the standard production architectural best practice.
Question 3: VPA recommends or applies:
  • A. Pod resource requests/limits adjustments
  • B. Service types
  • C. NetworkPolicies
  • D. DNS TTLs
✓ Correct Answer: A (Pod resource requests/limits adjustments)
Option A ('Pod resource requests/limits adjustments') is the standard production architectural best practice.
Question 4: KEDA is known for:
  • A. Event-driven autoscaling
  • B. Replacing CoreDNS
  • C. Backing up etcd
  • D. Signing images
✓ Correct Answer: A (Event-driven autoscaling)
Option A ('Event-driven autoscaling') is the standard production architectural best practice.
Question 5: Autoscaling works best when:
  • A. Metrics and requests are set thoughtfully
  • B. All Pods use BestEffort with no requests
  • C. You disable metrics-server
  • D. You scale manually only
✓ Correct Answer: A (Metrics and requests are set thoughtfully)
Option A ('Metrics and requests are set thoughtfully') is the standard production architectural best practice.
← Previous Module (21) GitOps (ArgoCD/Flux) Next Module (23) → Service Mesh & Multi-Cluster