⚡ ~/naveed k8s
⚡ Portfolio Home ✍️ Engineering Blog Deep Dives 🎯 Interview Hub 1,000+ Scenarios ☸️ Kubernetes Mastery Hub 24 Modules 🎮 DevOps Arcade & Quizzes Subnet Blitz ⚡ 🗺️ DevOps Roadmaps PDFs & Guides 🤖 Morpheus Analysis AI Quant ↗ 🛠️ Developer Tools Utilities 🧪 Labs & Experiments 📄 Interactive CV & Certs 🔗 All Links & Socials ⚡ Join The Dispatch (Weekly SRE Newsletter) →
Phase 2 — Cluster Administration Module 15 of 24 Free & Open Access

Monitoring & Logging

Complete production curriculum breakdown. Learn core architectural mechanics, study definitions in plain language, practice hands-on labs with the local minikube prod-sim cluster, and test active recall.

15 - Monitoring & Logging

Why this matters

You can't operate what you can't observe. This is the standard stack you'll meet at almost every company running Kubernetes.

Read this first — Definitions & Explanations

metrics-server

Provides resource metrics for kubectl top and resource-based HPA. It is not long-term historical monitoring.

Prometheus

Pull-based metrics system that scrapes targets (/metrics). Stores time series for alerting and dashboards.

Grafana

Visualization layer — dashboards on top of Prometheus (and other data sources).

Logging patterns

Apps write logs → node agents collect → aggregator/store (ELK/EFK, Loki, cloud logging). kubectl logs is for live debugging, not retention.

Alerting

Alert on symptoms and SLOs (error rate, latency, saturation), not every debug line. Actionable alerts only.

Official docs (read for detail)

Key Concepts

YouTube search terms

Hands-on lab (on prod-sim)

# You already have metrics-server. Add the full observability stack via Helm.
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo add grafana https://grafana.github.io/helm-charts
helm repo update

helm install kps prometheus-community/kube-prometheus-stack -n monitoring --create-namespace
kubectl get pods -n monitoring

# Access Grafana
kubectl -n monitoring port-forward svc/kps-grafana 3000:80
# open http://localhost:3000, default user admin, get password:
kubectl -n monitoring get secret kps-grafana -o jsonpath='{.data.admin-password}' | base64 -d

# Access Prometheus directly, run a PromQL query
kubectl -n monitoring port-forward svc/kps-kube-prometheus-stack-prometheus 9090
# open http://localhost:9090, try query: rate(container_cpu_usage_seconds_total[5m])

# Add Loki for logs
helm install loki grafana/loki-stack -n monitoring --set grafana.enabled=false
# Add Loki as a data source in Grafana (http://loki:3100), then explore logs by pod label

# Generate some load to see metrics move
kubectl run loadtest --image=busybox --restart=Never -- sh -c \
  "while true; do wget -qO- http://kubernetes.default 2>/dev/null; done"

Notes

(fill in your own words after watching + labbing)

📋 Self-Assessment Mastery Checklist (4 Competencies)
🧠 Practice Exam Questions (Module 15 MCQs)
⚡ Take Quiz & Save Progress in Tracker

Review these sample exam questions out loud, test your retrieval, and then unlock official scoring in the interactive tracker.

Question 1: metrics-server provides:
  • A. Resource metrics for kubectl top / HPA (resource)
  • B. Long-term Prometheus storage
  • C. Log aggregation
  • D. Ingress metrics only
✓ Correct Answer: A (Resource metrics for kubectl top / HPA (resource))
Option A ('Resource metrics for kubectl top / HPA (resource)') is the standard production architectural best practice.
Question 2: Prometheus typically scrapes:
  • A. Metrics endpoints from targets
  • B. etcd raw keys as logs
  • C. Container filesystem layers
  • D. Helm values
✓ Correct Answer: A (Metrics endpoints from targets)
Option A ('Metrics endpoints from targets') is the standard production architectural best practice.
Question 3: Grafana is commonly used to:
  • A. Visualize metrics/dashboards
  • B. Schedule Pods
  • C. Issue client certs
  • D. Replace CoreDNS
✓ Correct Answer: A (Visualize metrics/dashboards)
Option A ('Visualize metrics/dashboards') is the standard production architectural best practice.
Question 4: Centralized logging often uses:
  • A. agents → aggregator → store (e.g. EFK/PLG patterns)
  • B. Only kubectl logs forever on one node
  • C. etcd as the log DB
  • D. PVCs as the only log shipper
✓ Correct Answer: A (agents → aggregator → store (e.g. EFK/PLG patterns))
Option A ('agents → aggregator → store (e.g. EFK/PLG patterns)') is the standard production architectural best practice.
Question 5: Alerting should be based on:
  • A. SLOs/symptoms + actionable signals
  • B. Every debug log line
  • C. Only CPU always
  • D. Only Ingress 200s
✓ Correct Answer: A (SLOs/symptoms + actionable signals)
Option A ('SLOs/symptoms + actionable signals') is the standard production architectural best practice.
← Previous Module (14) Backup & Restore (etcd, Velero) Next Module (16) → Cluster Hardening & CIS Benchmarks