certslothcertsloth
KCNA/Topic 07

CNCF / Associate

Troubleshooting and application debugging

4 min read5 recall promptsReviewed 2026-10-10

Memory hook: Find the failing layer: placement, image, process, readiness, Service, network or dependency.

Must remember

  • Start with scope and evidence: current context, namespace, affected objects, recent changes and the intended outcome. kubectl get gives the overview; describe and events explain lifecycle failures; logs explain application behaviour.
  • A Pending Pod has not completed setup. Events may reveal insufficient resources, incompatible node selection, a missing volume or other scheduling/setup problems. Do not assume that an application exception causes every Pending state.
  • ImagePullBackOff points to repeated image-pull failures: check the image name/tag/digest, registry access, credentials and connectivity. CrashLoopBackOff means a container repeatedly exits or fails and restart attempts are backing off; it describes a symptom, not a unique root cause.
  • For crashes, inspect the current and previous container logs, exit reason, command, configuration, Secrets and resource limits. OOMKilled makes memory behaviour a priority. In multi-container Pods, select the correct container with -c.
  • Readiness asks whether a container should receive application traffic. Liveness asks whether a running container needs restarting. Startup gives a slow-starting application time to initialise before liveness and readiness checks take over.
  • A Pod can be Running without being Ready. A failed readiness check can remove its endpoint from normal Service traffic without restarting the container. Repeated liveness or startup failures can trigger container restart according to the configured behaviour.
  • For connectivity, trace name resolution → Service → EndpointSlices → Pod readiness → target port → process listener → policy and dependencies. A mismatched Service selector or wrong target port can look like an application outage.
  • For node-level issues, check node readiness, pressure conditions, resource availability and the kubelet/runtime/networking layer. A central application log alone cannot explain every infrastructure failure.
  • kubectl exec runs a command in an existing container; kubectl debug can introduce a diagnostic environment when supported and authorised. Minimal images may lack shells or debugging tools. Debug actions require appropriate permissions and can change or expose a workload.
  • Change one plausible cause at a time, observe the result and verify user-facing recovery. Repeatedly deleting Pods can erase useful evidence while a controller reproduces the same broken specification.

Read-only command recognition, using example names:

kubectl get pods -n demo -o wide
kubectl describe pod web -n demo
kubectl logs web -n demo -c app --previous
kubectl get events -n demo --sort-by=.metadata.creationTimestamp
kubectl get endpointslices -n demo -l kubernetes.io/service-name=web

Choose under exam pressure

Symptom First useful evidence
Pod is Pending Events, resource requests, node eligibility and claims
Container cannot download its image Image reference, registry connectivity and pull credentials
Container restarts repeatedly Previous logs, exit reason, configuration and probe behaviour
Service exists but has no ready backends Selector, EndpointSlices and readiness conditions
Slow startup causes repeated restarts Startup probe and probe timing, then the actual startup dependency

Traps

  • A readiness failure is not the same action as a liveness failure.
  • Exit status and events are evidence; status labels alone do not establish a root cause.
  • The relevant error may be in a previous container instance or another container in the Pod.
  • A liveness check that fails whenever an external dependency is temporarily unavailable can create unnecessary restart storms.

Active recall

1. A container is running, but its database-dependent readiness check fails. Should Kubernetes restart it solely for that readiness failure?

No. Readiness controls eligibility to receive traffic. Liveness and startup probes have restart implications; readiness failure on its own is not a restart instruction.

2. A new release immediately enters ImagePullBackOff. What should you check before debugging application code?

The image reference and whether the node can retrieve it: registry reachability, authentication, image existence and compatible artifact availability. The application may never have started.

3. The current logs are empty after a crash. Which kubectl option can retrieve the last terminated container instance's logs?

--previous, with the correct Pod and container selection. Combine that evidence with termination state, exit reason and events.

4. A Service selects app=web, but all intended Pods are labelled app=frontend. What result should you expect?

The selector does not match those Pods, so the Service will not obtain the intended backends through normal selector-based endpoint management. Correct the intended labels or selector rather than changing DNS first.

5. A healthy application needs two minutes to initialise, but a liveness check restarts it every twenty seconds. What is the design correction?

Use an appropriate startup probe and timings so initialisation can complete before liveness takes over. Also confirm that the application is actually progressing and that the probe tests the intended condition.

Sources

CLOSE THE NOTES. EXPLAIN THE CHOICE.

How well could you recall it?

Your next review is based on this answer. Progress stays in this browser.

Search across every published topic.