Kubernetes Troubleshooting Tools Your DevOps Team Should Use

A pod that keeps dying has already written down why. The exit code, the event list, and the container’s last line of output are sitting in the cluster, one command apart.

Kubernetes troubleshooting tools earn their install when they reach evidence the built-in commands cannot, and not before. I broke four workloads on a throwaway cluster to find where that line falls.

Why the evidence comes before the tool

A failure leaves three kinds of record, and each one captures a different part of it. The pod object holds what the control plane decided, the event stream holds what each component tried, and the container’s output holds what your process did with its turn.

They were produced in that order, so read them in that order. A dashboard that shows all three at once hides which question you are asking.

EvidenceWhere it livesFirst commandQuestion it answers
Pod phase and conditionsthe pod object’s statuskubectl get podDid it schedule, and what state is it in
Last exit code and reasoncontainerStatuses lastStatekubectl get pod -o jsonpathWhy the previous attempt stopped
Eventskubelet and control planekubectl describe pod or kubectl eventsWhat the cluster tried and could not finish
Container outputthe container runtime’s log storekubectl logsWhat the process printed before it stopped

Every command in that table ships with kubectl, so a first pass over an incident needs no install at all. A tool earns its place when a question survives all four rows unanswered.

What you need before the first command

The commands here ran against a kubectl v1.37.0 client talking to a v1.37.0 cluster. The client tolerates a server one minor version either side of it, so a v1.36 or v1.38 cluster takes the same sequence.

Beyond a kubeconfig that points at the failing cluster, you need these permissions:

  • read access to the namespace, which covers the pod object and the event list
  • pods/log for the container output
  • pods/exec for anything that opens a shell inside a container
  • pods/ephemeralcontainers for the debug container path, which many locked-down clusters never grant to application developers
  • cluster-admin or an equivalent role for the node-level path

If your permissions stop at the namespace edge, the pod paths still work and the node path is the one to hand to whoever owns the cluster.

Work a Kubernetes failure from the symptom to the cause

The symptoms that land on a DevOps on-call rotation come down to a handful of shapes, and each one has a first command that answers it. Match the symptom, run the command, and stop as soon as the output names a cause.

A pod that is restarting

CrashLoopBackOff means the container started, exited, and started again, so the cluster already holds the exit code.

kubectl get pods -n shop
NAME                        READY   STATUS             RESTARTS         AGE
api-gateway                 1/1     Running            0                4m12s
batch-import                0/1     Pending            0                10m
checkout-66f4c7c867-hcrtn   0/1     Error              7 (5m13s ago)    11m
checkout-66f4c7c867-qsgfq   0/1     Error              7 (5m44s ago)    11m
checkout-debug              0/2     CrashLoopBackOff   12 (3m30s ago)   9m21s
payments                    0/1     CrashLoopBackOff   6 (2m29s ago)    9m59s
report-worker               0/1     ImagePullBackOff   0                10m
kubectl get pods output for the shop namespace showing CrashLoopBackOff, Error, ImagePullBackOff, Pending and Running pods
CrashLoopBackOff, Error, ImagePullBackOff and Pending in one namespace. The status column names each one before you open a log.

The status column splits the cases before you read anything else. CrashLoopBackOff says the process ran and stopped, ImagePullBackOff says it never started, and Pending says the scheduler never placed it on a node.

The exit code from the last attempt is the single fact that decides where you look next.

kubectl get pod -n shop payments -o jsonpath='{.status.containerStatuses[0].lastState.terminated}'
{"containerID":"containerd://9c96bce0862a7032309e3c57cc17f5df9936e242697cdaaf2e7f0a82d66ed962","exitCode":1,"finishedAt":"2026-09-16T14:13:33Z","reason":"Error","startedAt":"2026-09-16T14:13:21Z"}

Exit 1 is your program reporting a failure. Exit 137 is a SIGKILL, which in a container usually means the kernel enforced a memory limit, and that same field then reads OOMKilled instead of Error. Exit 143 is a SIGTERM from a shutdown, and 139 is a segmentation fault.

That single value sends you to four different places, so the crash loop path starts here rather than at a tool. When you want the wider flag surface for those reads, the kubectl command reference on this site covers it.

I expected the previous container’s output to hold the error message, and that expectation was wrong.

kubectl logs -n shop payments --previous
payments: warming cache
payments: panic: ledger lock timeout

Previous output is served by the container runtime rather than by the pod object, and the runtime removes it along with the dead container. The same command answered with the lines above on one attempt and with nothing but this line on another: unable to retrieve container logs for containerd://2023bc59d85c04ac9caf9a2c7e76cb6b3f99f9f75b867e5332adc27e418011db.

Because that store is temporary, the exit code and the event reason are the parts of the record that survive, and they are usually enough to route the incident.

The event list is the cluster’s own account of what it tried.

kubectl events --for pod/payments -n shop --types=Warning
LAST SEEN            TYPE      REASON    OBJECT         MESSAGE
27s (x11 over 11m)   Warning   BackOff   Pod/payments   Back-off restarting failed container payments in pod payments_shop(c0878e60-5038-4649-b3ad-1a81dd252bb5)

BackOff paired with a rising restart count is the signature of a crash loop rather than a one-off exit. Read it next to the exit code and you have the shape of the failure before you open a single log line.

An image that will not pull

ImagePullBackOff reads straight off the event message, because the kubelet copies the registry response into it. I put a deliberately broken tag on a pod to see which reason appeared first.

kubectl events --for pod/report-worker -n shop
LAST SEEN             TYPE      REASON      OBJECT              MESSAGE
10m                   Normal    Scheduled   Pod/report-worker   Successfully assigned shop/report-worker to cfg-control-plane
7m22s (x5 over 10m)   Normal    Pulling     Pod/report-worker   Pulling image "busybox:1.36.1-not-a-real-tag"
7m21s (x5 over 10m)   Warning   Failed      Pod/report-worker   Failed to pull image "busybox:1.36.1-not-a-real-tag": rpc error: code = NotFound desc = failed to pull and unpack image "docker.io/library/busybox:1.36.1-not-a-real-tag": failed to resolve reference "docker.io/library/busybox:1.36.1-not-a-real-tag": docker.io/library/busybox:1.36.1-not-a-real-tag: not found
7m21s (x5 over 10m)   Warning   Failed      Pod/report-worker   Error: ErrImagePull
30s (x42 over 10m)    Normal    BackOff     Pod/report-worker   Back-off pulling image "busybox:1.36.1-not-a-real-tag"
30s (x42 over 10m)    Warning   Failed      Pod/report-worker   Error: ImagePullBackOff

The kubelet reports this in three states for one underlying problem. ErrImagePull is the first failure, ImagePullBackOff is the retry back-off, and in practice the same two warnings repeat with a growing age until someone changes the tag.

The cause sits in the last line of the message. Here the tag does not exist in the registry, and the response says so in as many words.

A pod that never schedules

A Pending pod never reached a container runtime, so there are no logs and no exit code to read. The scheduler leaves its reasoning on the pod condition instead.

kubectl get pod -n shop batch-import -o jsonpath='{.status.conditions[?(@.type=="PodScheduled")].message}'
0/1 nodes are available: 1 Insufficient cpu, 1 Insufficient memory. preemption: 0/1 nodes are available: 1 Preemption is not helpful for scheduling.

I sized that pod at 8 CPUs and 32Gi on purpose to make the scheduler refuse it. The message names every node and the reason it rejected the pod, so it can report two shortages at once.

kubectl describe node shows the same numbers as allocatable and allocated, which is where you check whether the request is unrealistic or the node is genuinely full.

A container with no shell

Distroless and scratch images carry no shell, so kubectl exec has nothing to open inside them. The debug container path adds a second container to a running pod instead, in the target’s process namespace.

kubectl debug -n shop api-gateway --image=busybox:latest --target=api --attach=false -- sh -c 'echo "from the debug container"; ps -o pid,comm | head -6'
Targeting container "api". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
Defaulting debug container name to debugger-g55nh.

The generated name matters, because kubectl logs rejects the plain debugger alias and accepts only the suffixed one.

kubectl logs -n shop api-gateway -c debugger-sv4mf
from the debug container
PID   COMMAND
    1 pause
    7 sh
   21 sleep
   22 sh
   35 ps

That process list is the proof the debug container sees the target’s processes rather than its own, and it depends on shareProcessNamespace being set on the pod.

The same command failed when I aimed it at a crash-looping pod:

failed to generate container "7c3b69c4f19f3f99de69a267cbc6712ee40b7d06fd4caa265ac944bc69a1ead3" spec: failed to generate spec opts: invalid target container: container "a3084a2db1fea280d65f1586c3b6c583cfa1246f518a34fda0ef44b473a4bd6e" is not running - in state CONTAINER_EXITED

An ephemeral container needs a target container that is still running. It suits a live process that is misbehaving, so a pod that keeps dying belongs in the crash loop path above instead.

When the misbehaving process is a JVM, the reach of a debug container changes again, and the walkthrough on debugging a Java application in Kubernetes production covers the flags, dumps and attach permissions that apply.

The node is the problem

A NotReady node, or one whose kubelet stopped posting, needs the same trick one level down.

kubectl debug node/cfg-control-plane --image=busybox:latest --attach=false -- sleep 3600
Creating debugging pod node-debugger-cfg-control-plane-jpsqs with container debugger on node cfg-control-plane.

That pod starts in the host namespaces with the node’s filesystem mounted at /host, which is what puts the kubelet and the container runtime within reach.

kubectl exec -n default node-debugger-cfg-control-plane-jpsqs -- ls /host/etc/kubernetes/manifests
etcd.yaml
kube-apiserver.yaml
kube-controller-manager.yaml
kube-scheduler.yaml

Those four manifests are the static pods that make up the control plane, and editing one in place is how you repair a control plane that will not start at all.

The debug container needs a long-running command, because the default one exits immediately and leaves a completed pod that refuses exec, and a non-root process cannot read the host root without extra capabilities. Each one blocks a different repair.

kubectl describe node is the read-only path when you cannot create a debug pod at all. Its conditions block reports memory, disk and PID pressure plus the kubelet’s own heartbeat, which is what separates a busy node from a dead one.

The tools that take over when kubectl stops answering

Each tool below answers a question the commands above raise and cannot settle, so I checked the release date of every one before writing it down, because a recommendation without that check ages badly.

ToolJob during an incidentVersion testedLast release
sternstream logs from several pods at oncev1.34.02026-05-02
k9sbrowse cluster objects in a terminalv0.51.02026-06-06
metrics-serversupply the metrics API behind kubectl topv0.9.02026-07-13
popeyescan objects for problems before they breakv0.22.12025-01-28
Helmhold the release history and roll backv4.3.02026-09-09
kube-prometheusstore and query the time seriesv0.18.02026-06-18
kube-state-metricsexpose object state as seriesv2.20.02026-08-18
Komodorassemble a change timelinecommercialn/a
Kubewatcharchived, do not adoptnone2020-07-03
Gitkubeunmaintained since 2020, do not adoptnone2020-04-09

The releases API for a project gives the last release date and whether the repository is archived. That is a few seconds of work and the difference between a working pointer and a dead end.

Stream every pod’s log at once

stern tails several pods and containers with one command, which matters when a request crosses four services and only one of them is unhappy.

stern -n shop --tail 2 --no-follow --selector app=checkout
+ checkout-66f4c7c867-qsgfq › checkout
+ checkout-66f4c7c867-hcrtn › checkout
checkout-66f4c7c867-hcrtn checkout checkout: connecting to redis:6379
- checkout-66f4c7c867-hcrtn › checkout
checkout-66f4c7c867-qsgfq checkout checkout: connecting to redis:6379
- checkout-66f4c7c867-qsgfq › checkout
stern streaming logs from two restarting checkout pods in the shop namespace
stern follows every pod matching a selector, which a single kubectl logs call cannot do.

I pointed it at two pods that restart every few seconds, and both streams arrived in one output. Selector matching is the reason to use it over a loop across pod names, since the same selector that defines a deployment defines its log set.

Browse the cluster in a terminal

k9s is a terminal interface over the same API, and its value during an incident is speed rather than new information.

k9s version
 ____  __ ________       
|    |/  /   __   \______
|       /\____    /  ___/
|    \   \  /    /\___  \
|____|\__ \/____//____  /
         \/           \/ 

Version:    v0.51.0
Commit:     558caafe7ba067467de46b320cc22ef11fef9c34
Date:       2026-06-06T14:04:25Z
k9s version output showing v0.51.0 with the project banner
The terminal interface at version v0.51.0.

A terminal interface assumes a person at a keyboard, so it earns its place on a workstation rather than in a script. Resource views, port forwards and log tails are the three jobs it shortens.

Get numbers for the container eating memory

kubectl top is part of the client, but it reads the metrics API rather than the pod objects, so a cluster with no metrics source answers with an error.

kubectl top pods -n shop
error: Metrics API not available

metrics-server is the component that supplies it. On a kind cluster it needs one extra argument, because the kubelet certificate it scrapes is not signed for the node address.

curl -sSLo metrics-server.yaml https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
kubectl apply -f metrics-server.yaml
kubectl -n kube-system patch deployment metrics-server --type=json \
  -p '[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'

I installed it on this cluster and the same command that failed a minute earlier started returning numbers. The latest download resolves to version v0.9.0.

kubectl top nodes
NAME                CPU(cores)   CPU(%)   MEMORY(bytes)   MEMORY(%)   
cfg-control-plane   173m         8%       725Mi           9%          
kubectl top nodes output showing CPU and memory for the cfg-control-plane node
kubectl top returns numbers only once a metrics source exists in the cluster.

The metrics path carries its own boundaries. A pod younger than the first scrape returns Metrics not available for that pod together with its age, because a Pending pod has no metrics to report at all.

Scan for what has not broken yet

popeye reads cluster objects and grades them against a set of checks, which is a different job from reading one failure. It scored this namespace and reported three things about the same broken deployment.

popeye -n shop -s deploy
DEPLOYMENTS (2 SCANNED)                                                        💥 1 😱 1 🔊 0 ✅ 0 0٪
  · shop/checkout..................................................................................💥
    💥 [POP-501] Unhealthy 2 desired but have 0 available.
    🐳 checkout
      😱 [POP-101] Image tagged "latest" in use.
      😱 [POP-106] No resources requests/limits defined.
  · shop/demo-release..............................................................................😱
    🐳 demo-release
      😱 [POP-106] No resources requests/limits defined.
SUMMARY
Your cluster score: F (0)

None of those three findings appear in the crash loop output, and the floating tag and the missing resource requests are the kind of problem worth reporting before an incident rather than during one.

Roll back the release instead of diagnosing it

Helm appears on every list of Kubernetes tools because it deploys things, and its use during an incident is the release history. I installed a release on this cluster to see what that history holds.

helm history demo-release -n shop
REVISION	UPDATED                 	STATUS  	CHART             	APP VERSION	DESCRIPTION     
1       	Wed Sep 16 14:20:25 2026	deployed	demo-release-0.1.0	1.16.0     	Install complete

An upgrade adds a revision, and the rollback puts the older chart back without a diagnosis.

helm upgrade demo-release ./demo-release -n shop --set replicaCount=2
helm rollback demo-release 1 -n shop
helm history demo-release -n shop
Rollback was a success! Happy Helming!
REVISION	UPDATED                 	STATUS    	CHART             	APP VERSION	DESCRIPTION     
1       	Wed Sep 16 14:20:25 2026	superseded	demo-release-0.1.0	1.16.0     	Install complete
2       	Wed Sep 16 14:20:26 2026	superseded	demo-release-0.1.0	1.16.0     	Upgrade complete
3       	Wed Sep 16 14:20:27 2026	deployed  	demo-release-0.1.0	1.16.0     	Rollback to 1   

The version that ran here is v4.3.0 from September 2026, so Helm 4 is the release family to install rather than Helm 3.

Tool lists go stale quickly. Kubewatch still turns up in recommendations, and its repository is archived with the last release in 2020. Gitkube has shipped nothing since April 2020, so neither belongs in a list you hand to a team today.

Checking the releases API before recommending a tool costs seconds and saves the reader a dead end.

Keep a timeline of what changed

The commands above tell you what the cluster is doing now, and nothing about what changed an hour ago. That second half is where many incidents get solved.

Prometheus stores time series and answers questions in PromQL, and the kube-prometheus operator bundle at v0.18.0 deploys it with Kubernetes dashboards and alerting rules. kube-state-metrics at v2.20.0 adds the object-level series that describe deployment and replica state rather than resource use.

The commercial option is a timeline product more than a dashboard. Komodor’s current positioning is an agentic operations platform for production, and what helps during an incident is the change history it assembles across deployments, configuration and infrastructure.

Headlamp, the Kubernetes SIG web interface at v0.45.0, covers the browse-the-cluster job for anyone who would rather not learn a terminal interface.

What each command does not cover

Every path above has a boundary that turns a one minute command into a blocked one, and knowing them in advance is the difference between escalating and flailing.

  • The metrics API is absent on a fresh cluster, so kubectl top fails until a metrics source exists.
  • Previous container output lives in the runtime’s log store and is removed with the container, so crash loop evidence can vanish between attempts.
  • An ephemeral container needs a target container that is still running, plus the pods/ephemeralcontainers verb.
  • A node debug pod needs a long-running command, and a non-root process cannot reach the host root without extra capabilities.
  • A dead kubelet writes no events at all, so a NotReady node has nothing to read until you reach the node itself.
  • Locked-down clusters refuse exec, debug containers and node pods by policy, and no flag changes that.

Some of those boundaries are permissions rather than technical limits, so confirm your own verbs before planning a path the cluster will refuse.

Read the events before you install anything

The built-in commands turn a red pod into a cause, and the tools that follow read the same evidence at a scale a person cannot. That makes the order fixed: describe the object, read the events, read the container output, and install something only when a question survives all three.

Start with the events for the namespace, sorted so the newest change sits at the bottom.

kubectl get events -n shop --sort-by=.lastTimestamp

Everything after that follows the evidence, which is the only sequence that keeps a tool from becoming the thing you are debugging.

FAQ

How do I troubleshoot a Kubernetes pod that keeps restarting?

Read the last exit code from the pod’s container status first, then the Warning events for that pod, then the container output. Exit 1 points at the application, exit 137 at an enforced memory limit, and 143 at a shutdown signal.

Why does kubectl logs –previous sometimes say there are no logs?

Previous output comes from the container runtime’s log store, and the runtime removes it along with the dead container. The exit code and the event reason stay on the pod object, so read those first and treat the log as a bonus.

Does kubectl top work without metrics-server?

No. kubectl top reads the metrics API rather than the pod objects, and a cluster with no metrics source answers with Metrics API not available. metrics-server is the component that provides that API.

Is kubectl debug safe to run in production?

The command adds a container to a running pod and shares the target’s namespaces without restarting it. The permission it needs, pods/ephemeralcontainers, is the control worth reviewing, because any process you start there runs with the pod’s access.

Pankaj Kumar
Pankaj Kumar

Pankaj Kumar is the founder and CEO of CodeForGeek, with more than 14 years in IT. He is an open-source enthusiast who enjoys sharing what he learns through CodeForGeek and YouTube, with a focus on Python, data analytics, machine learning, Angular, Node.js, and Kafka.

Articles: 336