13 - Troubleshooting & Debugging
Why this matters
This is the actual job. Anyone can kubectl apply a working manifest — mastery is
diagnosing a broken cluster/app fast, under pressure, with incomplete information.
Read this first — Definitions & Explanations
Debugging order that works
kubectl get/kubectl describe(Events!)kubectl logs/ previous logs- exec only if running
- node/runtime checks if it’s a node problem
CrashLoopBackOff
Container starts, exits, kubelet restarts it with backoff. Causes: bad command, missing config, failing dependency, failing liveness probe.
ImagePullBackOff
kubelet can’t pull the image (typo, private registry auth, network, rate limits).
Pending
Scheduler can’t place the Pod: insufficient CPU/memory, affinity/taints, PVC unbound, etc. Read Events.
describe Events
Often the fastest truth. Don’t skip them.
Useful tools
kubectl get events --sort-by=.lastTimestamp, kubectl top, kubectl get endpoints, node conditions (Ready, MemoryPressure).
Official docs (read for detail)
- Debug Running Pods
- Debug Services
- Determine the Reason for Pod Failure
- Debug Nodes / Troubleshoot Clusters
- Ephemeral Containers
Key Concepts
- The debugging funnel: Pod status → Events → Logs → Exec → Node → Network
kubectl describeEvents section is almost always the first real cluekubectl logs --previousfor crashed containers- Ephemeral debug containers (
kubectl debug) for distroless/minimal images with no shell - Common failure signatures:
CrashLoopBackOff,ImagePullBackOff,Pending(scheduling failure),Evicted,OOMKilled,CreateContainerConfigError - Node-level debugging: kubelet logs,
crictl, disk/memory pressure conditions kubectl get events --sort-by=.lastTimestamp -Aas a cluster-wide first look
YouTube search terms
- "Kubernetes troubleshooting CrashLoopBackOff step by step"
- "kubectl debug ephemeral containers explained"
- "Kubernetes Pending pod troubleshooting scheduling failures"
- "Kubernetes node NotReady troubleshooting"
Hands-on lab (on prod-sim)
# Manufacture and diagnose each classic failure
# 1. ImagePullBackOff
kubectl run bad-image --image=nginx:doesnotexist12345
kubectl describe pod bad-image | tail -15 # read the Events
# 2. CrashLoopBackOff
kubectl run crasher --image=busybox -- sh -c "exit 1"
kubectl logs crasher --previous
kubectl describe pod crasher | grep -A5 "Last State"
# 3. Pending (unschedulable — request more than any node has)
kubectl run huge --image=nginx --requests='cpu=100,memory=500Gi'
kubectl describe pod huge | grep -A5 Events # "Insufficient cpu/memory"
kubectl delete pod huge
# 4. CreateContainerConfigError (missing configmap/secret ref)
kubectl run missingcfg --image=nginx --env-from=configmap/doesnotexist
kubectl describe pod missingcfg | tail -10
# 5. Debug a "distroless" pod with no shell using ephemeral containers
kubectl run distroless --image=gcr.io/distroless/static-debian12 -- /nonexistent
kubectl debug -it distroless --image=busybox --target=distroless -- sh
# Cluster-wide triage habit
kubectl get events -A --sort-by=.lastTimestamp | tail -30
kubectl get pods -A | grep -v Running | grep -v Completed
# Node-level: simulate disk pressure awareness
kubectl describe node prod-sim-m02 | grep -A10 Conditions
Notes
(fill in your own words after watching + labbing)