2026-09-03

Kubernetes Ephemeral Containers: Safe Production Debugging

Safely inspect a live Pod without baking debugging tools into production images.

Ephemeral containers let you attach a temporary debugging container to an existing Pod.

This is useful when:

  • your app image is minimal (distroless/scratch)
  • you need curl, dig, tcpdump, strace, etc.
  • you want to keep production images clean

Requirements

  • Ephemeral Containers are Stable starting with Kubernetes 1.25; do not copy these commands to an older cluster unchanged
  • You need RBAC permission to use pods/ephemeralcontainers

Basic workflow

Attach a debug container:

kubectl debug -n <ns> -it pod/<pod> --image=busybox:1.36 --target=<container-name>

--target requests the process namespace of another container. The container runtime must support this capability; otherwise the debug container may start without visibility into the target processes. DNS and Service checks do not need it because containers in one Pod already share the network namespace.

Choosing the right debug image

Different debugging tasks need different tools. A few common patterns:

  • BusyBox: tiny, good for basic networking (nslookup, wget, nc).
  • Alpine: a bit more flexible, can add packages if needed (but beware of network egress restrictions).
  • Netshoot-style images: loaded with curl, dig, tcpdump, mtr, etc. Great for networking, but heavier and potentially riskier.

In production, consider maintaining a blessed debug image:

  • pinned by digest (immutable)
  • regularly scanned
  • minimal but sufficient tools

That gives you consistent, auditable behavior.

Understanding --target (process namespace sharing)

When you specify:

kubectl debug -n <ns> -it pod/<pod> --image=<img> --target=<container>

Kubernetes asks the runtime to place the ephemeral container in the target container’s process namespace. This helps when you want to:

  • inspect processes
  • run tools like strace (where permitted)
  • understand what the application is doing in real time

If the runtime does not support target process namespaces, or policy blocks the required privileges, the debug container may run without showing target processes. You still share the Pod network namespace, which is usually enough for:

  • DNS checks
  • Service connectivity tests
  • HTTP probing from “inside the Pod”

Common production use cases

1) Check DNS and Service resolution

kubectl debug -n <ns> -it pod/<pod> --image=busybox:1.36
nslookup kubernetes.default.svc.cluster.local
nslookup <service>.<ns>.svc.cluster.local
cat /etc/resolv.conf

This separates a wrong Service name from blocked DNS egress, a CoreDNS failure, or a node/CNI problem.

2) Test the Service path from inside the Pod

wget -S -O- http://<service>.<ns>.svc.cluster.local:8080/readyz
nc -vz <service>.<ns>.svc.cluster.local 8080

The test runs in the same Pod network namespace as the application, so it exercises the application’s real network boundary.

3) Readiness fails but the distroless image has no curl

kubectl debug -n <ns> -it pod/<pod> --image=curlimages/curl:8.5.0
curl -sS -i http://127.0.0.1:8080/readyz

If the loopback request succeeds but the external probe fails, inspect the Service selector, endpoints, and node network before changing the application.

4) Check whether NetworkPolicy blocks egress

Ephemeral containers remain subject to the Pod’s NetworkPolicy rules, which makes the result representative of the application:

nslookup <dependency>.<ns>.svc.cluster.local
nc -vz <dependency>.<ns>.svc.cluster.local 5432

Compare the result from another namespace. A namespace-specific failure usually points to labels or policy selectors.

5) Inspect environment and mounts without changing state

Common configuration failures include missing environment variables, ConfigMap or Secret mounts at the wrong path, and stale files. Prefer read-only commands such as env, ls -la /path, and cat /etc/hosts.

Limitations you should know

Ephemeral containers are intentionally limited:

  • They don’t become part of your deployment spec, so you can’t “fix” a Pod by leaving one around.
  • They are not meant to expose ports for traffic (treat them as internal tools).
  • They are not restarted if they exit.

Also, some environments restrict ephemeral containers for security reasons (policy engines, admission control, managed platforms).

Make debugging safe: RBAC + audit + process

An ephemeral container effectively brings an operator into a production Pod. Separate permission to read Pods from permission to update the pods/ephemeralcontainers subresource:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: ephemeral-debug
  namespace: app
rules:
  - apiGroups: [""]
    resources: ["pods"]
    verbs: ["get", "list"]
  - apiGroups: [""]
    resources: ["pods/ephemeralcontainers"]
    verbs: ["get", "update"]

Before an incident, decide which namespaces allow debugging, which images are approved, and when node-level debugging requires additional approval. Record who ran kubectl debug, when, why, and which incident or ticket authorized it. Treat the session as read-only unless the incident procedure explicitly says otherwise.

High-risk debugging: copy first

If the investigation requires installing packages, restarting processes, or attaching heavier tooling, copy the Pod when your kubectl version and policy allow it. Higher-risk investigations should favor “copy, then debug” over experimenting inside the live production Pod.

kubectl debug -n <ns> -it pod/<pod> \
  --image=ubuntu:24.04 \
  --share-processes \
  --copy-to=<pod>-debug

kubectl delete pod -n <ns> <pod>-debug

The copied Pod may still reach production dependencies. Ensure no Service selector targets it, and do not start side-effecting consumers or scheduled work during the investigation.

Bonus: Debug the node (when the problem is below Kubernetes)

Sometimes the issue is located on the node rather than inside the Pod:

  • CNI problems (routes/iptables/eBPF)
  • disk pressure
  • kubelet or container runtime issues
  • DNS problems from the node’s perspective

kubectl debug node/<node> creates a debug Pod, but the default profile may not have enough privilege for chroot /host. When host-level access is genuinely required, select the sysadmin profile explicitly and treat the session as a privileged operation.

Conceptually, it looks like this:

kubectl debug node/<node> -it \
  --image=ubuntu:24.04 \
  --profile=sysadmin \
  -- chroot /host

Once inside (and if permitted), you can run:

  • ip a, ip r
  • ss -lntp
  • journalctl (on some systems)
  • check /etc/resolv.conf and CNI config

Important: this is powerful and should be restricted even more tightly than pod debugging.

Close the debugging session properly

Ephemeral container status remains attached to the Pod until the Pod is deleted. Leaving the shell is only the first step:

  • record the commands, relevant output, and conclusion
  • link the investigation to an incident or ticket
  • deliver the real fix through configuration, image, or deployment workflow

Pre-rollout checklist

  • Production images remain minimal and do not bundle general-purpose debug tools.
  • Approved ephemeral containers can reproduce the Pod’s real DNS and network path.
  • RBAC limits pods/ephemeralcontainers, and audit logs record each use.
  • Investigation is read-only by default; fixes return through the release process.

Three operational boundaries

Q: Are ephemeral containers safe in production? A: They are intended for debugging and do not restart automatically. Use RBAC to restrict who can add them.

Q: Why can I not add volumes or ports? A: Ephemeral containers are intentionally limited to avoid mutating the workload. Use them for inspection, not for changes.

Q: How do I view logs? A: Use kubectl logs <pod> -c <ephemeral-container-name> or kubectl describe pod to confirm status.

References