Memory hook: Context first; events next; change one cause.
Must remember
Start with kubectl config current-context and the requested namespace. Use get, describe and sorted events to identify the failing layer before editing. Verify the target context again when switching between tasks.
| Symptom | First useful evidence |
|---|---|
| Pod Pending | Scheduling events: requests, affinity, taints, quota and PVC binding. |
| ImagePullBackOff | Image reference, registry reachability and image-pull credentials. |
| CrashLoopBackOff | kubectl logs POD --previous, exit code, command and probes. |
| Running but unavailable | Readiness, listener/targetPort, selectors and EndpointSlices. |
| Node NotReady | Node conditions, kubelet/runtime health, disk pressure and network plugin. |
| API unavailable | API endpoint, control-plane static Pods, certificates and etcd health. |
kubectl logs -c CONTAINER selects the right container; --previous reads a prior terminated instance. kubectl exec is useful when tools exist in the container; minimal images may lack a shell. Use an approved debug container or diagnostic Pod when appropriate, without assuming it has identical policy/identity to the failing workload.
On a node, systemctl status kubelet, journalctl -u kubelet and runtime tools such as crictl help when the API path is broken. Static control-plane Pod manifests are managed by the kubelet; keep backups outside its watched manifest directory to avoid accidental extra Pods.
kubectl top requires a metrics API and shows recent usage; it is not a full historical monitoring system. Compare usage with requests, limits and node allocatable capacity. Check DNS, Service endpoints, policies and routing in sequence for network failures.
Timed drill: diagnose a deliberately broken selector, an oversized request and a wrong image in a disposable cluster. State the observed cause, make the smallest correction, then verify the requested application behavior. A green Pod phase alone is not proof of success.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Application recently restarted | Previous container logs and termination status. |
| Cluster API unavailable | Node-level kubelet/runtime and control-plane diagnostics. |
| Service DNS resolves but request fails | Check endpoints and application ports, then traffic policy. |
Traps
- Blind restarts can erase evidence and fail to fix the cause.
- Editing the wrong context can complete the wrong task perfectly.
Active recall
1. What should you confirm before every cluster task?
The required kubeconfig context and namespace.
2. Where are logs from the last crashed container?
kubectl logs with --previous, selecting the correct container.
3. Why might kubectl top fail on a healthy cluster?
The metrics API may be absent or unavailable.
4. What first explains an unschedulable Pod?
Its events and scheduling constraints, not application logs from a container that never started.
5. How do you prove a Service fix?
Verify ready endpoints and an actual request through the intended Service path.