⚡ ~/naveed k8s
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Phase 1 — Fundamentals Module 06 of 24 Free & Open Access

Scheduling: Affinity, Taints, Resources, PDBs

Complete production curriculum breakdown. Learn core architectural mechanics, study definitions in plain language, practice hands-on labs with the local minikube prod-sim cluster, and test active recall.

06 - Scheduling: Affinity, Taints, Resources, PDBs

Why this matters

This is how you control where pods run and what happens when nodes are under pressure or being drained. Critical for prod stability and multi-tenant clusters.

Read this first — Definitions & Explanations

Scheduling

The process of choosing which node a Pod should run on. Done by kube-scheduler (usually).

Requests vs Limits

nodeSelector

Simple “must have this label” rule for nodes.

Affinity / anti-affinity

Richer rules: prefer/require certain nodes or relative placement to other Pods (spread replicas across zones/nodes).

Taints and Tolerations

PodDisruptionBudget (PDB)

Limits how many Pods in a set can be voluntarily disrupted at once (drains, upgrades). Protects availability during maintenance.

Priority and preemption

Higher-priority Pods can preempt lower-priority ones when the cluster is full (advanced scheduling behavior).

Official docs (read for detail)

Key Concepts

YouTube search terms

Hands-on lab (on prod-sim)

# Taint a worker node, prove pods don't schedule there without toleration
kubectl taint nodes prod-sim-m03 dedicated=gpu:NoSchedule
kubectl create deployment plain --image=nginx --replicas=3
kubectl get pods -o wide   # none should land on prod-sim-m03

# Add toleration + nodeAffinity to force it there
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
  name: gpu-pod
spec:
  tolerations:
  - key: dedicated
    operator: Equal
    value: gpu
    effect: NoSchedule
  affinity:
    nodeAffinity:
      requiredDuringSchedulingIgnoredDuringExecution:
        nodeSelectorTerms:
        - matchExpressions:
          - key: kubernetes.io/hostname
            operator: In
            values: ["prod-sim-m03"]
  containers:
  - name: gpu-pod
    image: nginx
EOF
kubectl get pod gpu-pod -o wide   # should be on prod-sim-m03
kubectl taint nodes prod-sim-m03 dedicated=gpu:NoSchedule-   # cleanup

# Resource requests/limits + OOMKill demo
kubectl run oom --image=polinux/stress --restart=Never \
  --requests='memory=50Mi' --limits='memory=100Mi' \
  -- stress --vm 1 --vm-bytes 200M --vm-hang 1
kubectl get pod oom -w   # watch it go OOMKilled
kubectl describe pod oom | grep -A3 "Last State"

# PodDisruptionBudget — protect a deployment during a drain
kubectl create deployment protected --image=nginx --replicas=3
cat <<EOF | kubectl apply -f -
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: protected-pdb
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app: protected
EOF
kubectl drain prod-sim-m02 --ignore-daemonsets --delete-emptydir-data
# watch it respect the PDB (won't evict below minAvailable)
kubectl uncordon prod-sim-m02

Notes

(fill in your own words after watching + labbing)

📋 Self-Assessment Mastery Checklist (4 Competencies)
🧠 Practice Exam Questions (Module 06 MCQs)
⚡ Take Quiz & Save Progress in Tracker

Review these sample exam questions out loud, test your retrieval, and then unlock official scoring in the interactive tracker.

Question 1: The scheduler's main job is to:
  • A. Run containers
  • B. Choose a node for unbound Pods
  • C. Store objects in etcd
  • D. Terminate Nodes
✓ Correct Answer: B (Choose a node for unbound Pods)
Option B ('Choose a node for unbound Pods') is the standard production architectural best practice.
Question 2: A taint on a node prevents Pods from scheduling unless they have a matching:
  • A. Label
  • B. Toleration
  • C. Annotation only
  • D. ServiceAccount
✓ Correct Answer: B (Toleration)
Option B ('Toleration') is the standard production architectural best practice.
Question 3: nodeSelector / affinity influence:
  • A. Which nodes a Pod can land on
  • B. Service load balancing algorithm
  • C. Ingress TLS
  • D. etcd compaction
✓ Correct Answer: A (Which nodes a Pod can land on)
Option A ('Which nodes a Pod can land on') is the standard production architectural best practice.
Question 4: Requests vs Limits: requests mainly affect:
  • A. Scheduling / Guaranteed resources
  • B. Image pull policy
  • C. DNS TTL
  • D. API authorization
✓ Correct Answer: A (Scheduling / Guaranteed resources)
Option A ('Scheduling / Guaranteed resources') is the standard production architectural best practice.
Question 5: A PodDisruptionBudget (PDB) helps:
  • A. Limit voluntary disruptions to keep availability
  • B. Encrypt Secrets
  • C. Speed up image pulls
  • D. Replace HPA
✓ Correct Answer: A (Limit voluntary disruptions to keep availability)
Option A ('Limit voluntary disruptions to keep availability') is the standard production architectural best practice.
← Previous Module (05) Storage: Volumes, PV, PVC, StorageClasses Next Module (07) → Ingress & Ingress Controllers