15 - Monitoring & Logging
Why this matters
You can't operate what you can't observe. This is the standard stack you'll meet at almost every company running Kubernetes.
Read this first — Definitions & Explanations
metrics-server
Provides resource metrics for kubectl top and resource-based HPA. It is not long-term historical monitoring.
Prometheus
Pull-based metrics system that scrapes targets (/metrics). Stores time series for alerting and dashboards.
Grafana
Visualization layer — dashboards on top of Prometheus (and other data sources).
Logging patterns
Apps write logs → node agents collect → aggregator/store (ELK/EFK, Loki, cloud logging). kubectl logs is for live debugging, not retention.
Alerting
Alert on symptoms and SLOs (error rate, latency, saturation), not every debug line. Actionable alerts only.
Official docs (read for detail)
Key Concepts
- metrics-server: lightweight, powers
kubectl topand HPA — NOT for long-term storage - Prometheus: pull-based metrics, PromQL, ServiceMonitor/PodMonitor (via Prometheus Operator)
- Grafana: dashboards on top of Prometheus (and Loki)
- Loki: log aggregation, label-based (like Prometheus but for logs)
- kube-state-metrics: cluster object state as metrics (separate from metrics-server)
- Node-level logging: where container logs actually live on disk, log rotation
- Alerting: Alertmanager, routing, silence windows
YouTube search terms
- "Prometheus Grafana Kubernetes monitoring stack tutorial"
- "kube-state-metrics vs metrics-server explained"
- "Loki Kubernetes logging tutorial"
- "PromQL tutorial for beginners"
Hands-on lab (on prod-sim)
# You already have metrics-server. Add the full observability stack via Helm.
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo add grafana https://grafana.github.io/helm-charts
helm repo update
helm install kps prometheus-community/kube-prometheus-stack -n monitoring --create-namespace
kubectl get pods -n monitoring
# Access Grafana
kubectl -n monitoring port-forward svc/kps-grafana 3000:80
# open http://localhost:3000, default user admin, get password:
kubectl -n monitoring get secret kps-grafana -o jsonpath='{.data.admin-password}' | base64 -d
# Access Prometheus directly, run a PromQL query
kubectl -n monitoring port-forward svc/kps-kube-prometheus-stack-prometheus 9090
# open http://localhost:9090, try query: rate(container_cpu_usage_seconds_total[5m])
# Add Loki for logs
helm install loki grafana/loki-stack -n monitoring --set grafana.enabled=false
# Add Loki as a data source in Grafana (http://loki:3100), then explore logs by pod label
# Generate some load to see metrics move
kubectl run loadtest --image=busybox --restart=Never -- sh -c \
"while true; do wget -qO- http://kubernetes.default 2>/dev/null; done"
Notes
(fill in your own words after watching + labbing)