⚡ ~/naveed k8s
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Phase 2 — Cluster Administration Module 13 of 24 Free & Open Access

Troubleshooting & Debugging

Complete production curriculum breakdown. Learn core architectural mechanics, study definitions in plain language, practice hands-on labs with the local minikube prod-sim cluster, and test active recall.

13 - Troubleshooting & Debugging

Why this matters

This is the actual job. Anyone can kubectl apply a working manifest — mastery is diagnosing a broken cluster/app fast, under pressure, with incomplete information.

Read this first — Definitions & Explanations

Debugging order that works

  1. kubectl get / kubectl describe (Events!)
  2. kubectl logs / previous logs
  3. exec only if running
  4. node/runtime checks if it’s a node problem

CrashLoopBackOff

Container starts, exits, kubelet restarts it with backoff. Causes: bad command, missing config, failing dependency, failing liveness probe.

ImagePullBackOff

kubelet can’t pull the image (typo, private registry auth, network, rate limits).

Pending

Scheduler can’t place the Pod: insufficient CPU/memory, affinity/taints, PVC unbound, etc. Read Events.

describe Events

Often the fastest truth. Don’t skip them.

Useful tools

kubectl get events --sort-by=.lastTimestamp, kubectl top, kubectl get endpoints, node conditions (Ready, MemoryPressure).

Official docs (read for detail)

Key Concepts

YouTube search terms

Hands-on lab (on prod-sim)

# Manufacture and diagnose each classic failure

# 1. ImagePullBackOff
kubectl run bad-image --image=nginx:doesnotexist12345
kubectl describe pod bad-image | tail -15   # read the Events

# 2. CrashLoopBackOff
kubectl run crasher --image=busybox -- sh -c "exit 1"
kubectl logs crasher --previous
kubectl describe pod crasher | grep -A5 "Last State"

# 3. Pending (unschedulable — request more than any node has)
kubectl run huge --image=nginx --requests='cpu=100,memory=500Gi'
kubectl describe pod huge | grep -A5 Events   # "Insufficient cpu/memory"
kubectl delete pod huge

# 4. CreateContainerConfigError (missing configmap/secret ref)
kubectl run missingcfg --image=nginx --env-from=configmap/doesnotexist
kubectl describe pod missingcfg | tail -10

# 5. Debug a "distroless" pod with no shell using ephemeral containers
kubectl run distroless --image=gcr.io/distroless/static-debian12 -- /nonexistent
kubectl debug -it distroless --image=busybox --target=distroless -- sh

# Cluster-wide triage habit
kubectl get events -A --sort-by=.lastTimestamp | tail -30
kubectl get pods -A | grep -v Running | grep -v Completed

# Node-level: simulate disk pressure awareness
kubectl describe node prod-sim-m02 | grep -A10 Conditions

Notes

(fill in your own words after watching + labbing)

📋 Self-Assessment Mastery Checklist (4 Competencies)
🧠 Practice Exam Questions (Module 13 MCQs)
⚡ Take Quiz & Save Progress in Tracker

Review these sample exam questions out loud, test your retrieval, and then unlock official scoring in the interactive tracker.

Question 1: First place to check for a failing Pod often is:
  • A. kubectl describe pod / kubectl logs
  • B. Deleting the node immediately
  • C. Rebooting etcd blindly
  • D. Removing CNI
✓ Correct Answer: A (kubectl describe pod / kubectl logs)
Option A ('kubectl describe pod / kubectl logs') is the standard production architectural best practice.
Question 2: CrashLoopBackOff usually means:
  • A. Container starts then exits repeatedly
  • B. Scheduler is offline permanently
  • C. PVC is ReadOnlyMany
  • D. Ingress TLS is perfect
✓ Correct Answer: A (Container starts then exits repeatedly)
Option A ('Container starts then exits repeatedly') is the standard production architectural best practice.
Question 3: ImagePullBackOff indicates:
  • A. Cannot pull the container image
  • B. DNS works but Service is missing
  • C. RBAC denied list nodes
  • D. etcd is full always
✓ Correct Answer: A (Cannot pull the container image)
Option A ('Cannot pull the container image') is the standard production architectural best practice.
Question 4: kubectl get events helps with:
  • A. Recent cluster/object lifecycle messages
  • B. Compressing logs
  • C. Creating StorageClasses
  • D. Upgrading kubeadm
✓ Correct Answer: A (Recent cluster/object lifecycle messages)
Option A ('Recent cluster/object lifecycle messages') is the standard production architectural best practice.
Question 5: Pending Pods are often caused by:
  • A. Scheduling failures (resources/affinity/taints)
  • B. Successful Running state
  • C. Completed Jobs
  • D. Healthy ReplicaSets only
✓ Correct Answer: A (Scheduling failures (resources/affinity/taints))
Option A ('Scheduling failures (resources/affinity/taints)') is the standard production architectural best practice.
← Previous Module (12) Networking Deep Dive (CNI, NetworkPolicy) Next Module (14) → Backup & Restore (etcd, Velero)