06 - Scheduling: Affinity, Taints, Resources, PDBs
Why this matters
This is how you control where pods run and what happens when nodes are under pressure or being drained. Critical for prod stability and multi-tenant clusters.
Read this first — Definitions & Explanations
Scheduling
The process of choosing which node a Pod should run on. Done by kube-scheduler (usually).
Requests vs Limits
- requests: what the scheduler uses to place the Pod (“needs 250m CPU”)
- limits: hard cap at runtime (exceed CPU → throttled; exceed memory → often OOMKilled)
nodeSelector
Simple “must have this label” rule for nodes.
Affinity / anti-affinity
Richer rules: prefer/require certain nodes or relative placement to other Pods (spread replicas across zones/nodes).
Taints and Tolerations
- Taint on a node: “don’t schedule here unless you tolerate me”
- Toleration on a Pod: “I accept that taint”
Used to reserve nodes (GPU, control-plane, dedicated teams).
PodDisruptionBudget (PDB)
Limits how many Pods in a set can be voluntarily disrupted at once (drains, upgrades). Protects availability during maintenance.
Priority and preemption
Higher-priority Pods can preempt lower-priority ones when the cluster is full (advanced scheduling behavior).
Official docs (read for detail)
- Assigning Pods to Nodes
- Taints and Tolerations
- Resource Management for Pods and Containers
- Pod Disruption Budgets
- Scheduling Framework
Key Concepts
- Resource
requestsvslimits(CPU is compressible, memory is not — OOMKill risk) - QoS classes: Guaranteed, Burstable, BestEffort (drives eviction order)
- nodeSelector (simple) vs nodeAffinity (expressive, required/preferred)
- podAffinity / podAntiAffinity — co-locate or spread pods relative to other pods
- Taints & Tolerations — repel pods from nodes unless they tolerate the taint
- Topology spread constraints — spread pods evenly across zones/nodes
- PriorityClass — which pods get scheduled/evicted first under pressure
- PodDisruptionBudget (PDB) — protects availability during voluntary disruptions (drains, upgrades)
YouTube search terms
- "Kubernetes resource requests and limits explained"
- "Kubernetes taints and tolerations vs node affinity"
- "Kubernetes PodDisruptionBudget explained"
- "Kubernetes QoS classes explained"
Hands-on lab (on prod-sim)
# Taint a worker node, prove pods don't schedule there without toleration
kubectl taint nodes prod-sim-m03 dedicated=gpu:NoSchedule
kubectl create deployment plain --image=nginx --replicas=3
kubectl get pods -o wide # none should land on prod-sim-m03
# Add toleration + nodeAffinity to force it there
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: gpu-pod
spec:
tolerations:
- key: dedicated
operator: Equal
value: gpu
effect: NoSchedule
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/hostname
operator: In
values: ["prod-sim-m03"]
containers:
- name: gpu-pod
image: nginx
EOF
kubectl get pod gpu-pod -o wide # should be on prod-sim-m03
kubectl taint nodes prod-sim-m03 dedicated=gpu:NoSchedule- # cleanup
# Resource requests/limits + OOMKill demo
kubectl run oom --image=polinux/stress --restart=Never \
--requests='memory=50Mi' --limits='memory=100Mi' \
-- stress --vm 1 --vm-bytes 200M --vm-hang 1
kubectl get pod oom -w # watch it go OOMKilled
kubectl describe pod oom | grep -A3 "Last State"
# PodDisruptionBudget — protect a deployment during a drain
kubectl create deployment protected --image=nginx --replicas=3
cat <<EOF | kubectl apply -f -
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: protected-pdb
spec:
minAvailable: 2
selector:
matchLabels:
app: protected
EOF
kubectl drain prod-sim-m02 --ignore-daemonsets --delete-emptydir-data
# watch it respect the PDB (won't evict below minAvailable)
kubectl uncordon prod-sim-m02
Notes
(fill in your own words after watching + labbing)