Kubernetes Ephemeral Containers: Safe Production Debugging
Safely inspect a live Pod without baking debugging tools into production images.
Ephemeral containers let you attach a temporary debugging container to an existing Pod.
This is useful when:
- your app image is minimal (distroless/scratch)
- you need
curl,dig,tcpdump,strace, etc. - you want to keep production images clean
Requirements
- Ephemeral Containers are Stable starting with Kubernetes 1.25; do not copy these commands to an older cluster unchanged
- You need RBAC permission to use
pods/ephemeralcontainers
Basic workflow
Attach a debug container:
kubectl debug -n <ns> -it pod/<pod> --image=busybox:1.36 --target=<container-name>
--target requests the process namespace of another container. The container runtime must support this capability; otherwise the debug container may start without visibility into the target processes. DNS and Service checks do not need it because containers in one Pod already share the network namespace.
Choosing the right debug image
Different debugging tasks need different tools. A few common patterns:
- BusyBox: tiny, good for basic networking (
nslookup,wget,nc). - Alpine: a bit more flexible, can add packages if needed (but beware of network egress restrictions).
- Netshoot-style images: loaded with
curl,dig,tcpdump,mtr, etc. Great for networking, but heavier and potentially riskier.
In production, consider maintaining a blessed debug image:
- pinned by digest (immutable)
- regularly scanned
- minimal but sufficient tools
That gives you consistent, auditable behavior.
Understanding --target (process namespace sharing)
When you specify:
kubectl debug -n <ns> -it pod/<pod> --image=<img> --target=<container>
Kubernetes asks the runtime to place the ephemeral container in the target container’s process namespace. This helps when you want to:
- inspect processes
- run tools like
strace(where permitted) - understand what the application is doing in real time
If the runtime does not support target process namespaces, or policy blocks the required privileges, the debug container may run without showing target processes. You still share the Pod network namespace, which is usually enough for:
- DNS checks
- Service connectivity tests
- HTTP probing from “inside the Pod”
Common production use cases
1) Check DNS and Service resolution
kubectl debug -n <ns> -it pod/<pod> --image=busybox:1.36
nslookup kubernetes.default.svc.cluster.local
nslookup <service>.<ns>.svc.cluster.local
cat /etc/resolv.conf
This separates a wrong Service name from blocked DNS egress, a CoreDNS failure, or a node/CNI problem.
2) Test the Service path from inside the Pod
wget -S -O- http://<service>.<ns>.svc.cluster.local:8080/readyz
nc -vz <service>.<ns>.svc.cluster.local 8080
The test runs in the same Pod network namespace as the application, so it exercises the application’s real network boundary.
3) Readiness fails but the distroless image has no curl
kubectl debug -n <ns> -it pod/<pod> --image=curlimages/curl:8.5.0
curl -sS -i http://127.0.0.1:8080/readyz
If the loopback request succeeds but the external probe fails, inspect the Service selector, endpoints, and node network before changing the application.
4) Check whether NetworkPolicy blocks egress
Ephemeral containers remain subject to the Pod’s NetworkPolicy rules, which makes the result representative of the application:
nslookup <dependency>.<ns>.svc.cluster.local
nc -vz <dependency>.<ns>.svc.cluster.local 5432
Compare the result from another namespace. A namespace-specific failure usually points to labels or policy selectors.
5) Inspect environment and mounts without changing state
Common configuration failures include missing environment variables, ConfigMap or Secret mounts at the wrong path, and stale files. Prefer read-only commands such as env, ls -la /path, and cat /etc/hosts.
Limitations you should know
Ephemeral containers are intentionally limited:
- They don’t become part of your deployment spec, so you can’t “fix” a Pod by leaving one around.
- They are not meant to expose ports for traffic (treat them as internal tools).
- They are not restarted if they exit.
Also, some environments restrict ephemeral containers for security reasons (policy engines, admission control, managed platforms).
Make debugging safe: RBAC + audit + process
An ephemeral container effectively brings an operator into a production Pod. Separate permission to read Pods from permission to update the pods/ephemeralcontainers subresource:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: ephemeral-debug
namespace: app
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["get", "list"]
- apiGroups: [""]
resources: ["pods/ephemeralcontainers"]
verbs: ["get", "update"]
Before an incident, decide which namespaces allow debugging, which images are approved, and when node-level debugging requires additional approval. Record who ran kubectl debug, when, why, and which incident or ticket authorized it. Treat the session as read-only unless the incident procedure explicitly says otherwise.
High-risk debugging: copy first
If the investigation requires installing packages, restarting processes, or attaching heavier tooling, copy the Pod when your kubectl version and policy allow it. Higher-risk investigations should favor “copy, then debug” over experimenting inside the live production Pod.
kubectl debug -n <ns> -it pod/<pod> \
--image=ubuntu:24.04 \
--share-processes \
--copy-to=<pod>-debug
kubectl delete pod -n <ns> <pod>-debug
The copied Pod may still reach production dependencies. Ensure no Service selector targets it, and do not start side-effecting consumers or scheduled work during the investigation.
Bonus: Debug the node (when the problem is below Kubernetes)
Sometimes the issue is located on the node rather than inside the Pod:
- CNI problems (routes/iptables/eBPF)
- disk pressure
- kubelet or container runtime issues
- DNS problems from the node’s perspective
kubectl debug node/<node> creates a debug Pod, but the default profile may not have enough privilege for chroot /host. When host-level access is genuinely required, select the sysadmin profile explicitly and treat the session as a privileged operation.
Conceptually, it looks like this:
kubectl debug node/<node> -it \
--image=ubuntu:24.04 \
--profile=sysadmin \
-- chroot /host
Once inside (and if permitted), you can run:
ip a,ip rss -lntpjournalctl(on some systems)- check
/etc/resolv.confand CNI config
Important: this is powerful and should be restricted even more tightly than pod debugging.
Close the debugging session properly
Ephemeral container status remains attached to the Pod until the Pod is deleted. Leaving the shell is only the first step:
- record the commands, relevant output, and conclusion
- link the investigation to an incident or ticket
- deliver the real fix through configuration, image, or deployment workflow
Pre-rollout checklist
- Production images remain minimal and do not bundle general-purpose debug tools.
- Approved ephemeral containers can reproduce the Pod’s real DNS and network path.
- RBAC limits
pods/ephemeralcontainers, and audit logs record each use. - Investigation is read-only by default; fixes return through the release process.
Three operational boundaries
Q: Are ephemeral containers safe in production? A: They are intended for debugging and do not restart automatically. Use RBAC to restrict who can add them.
Q: Why can I not add volumes or ports? A: Ephemeral containers are intentionally limited to avoid mutating the workload. Use them for inspection, not for changes.
Q: How do I view logs?
A: Use kubectl logs <pod> -c <ephemeral-container-name> or kubectl describe pod to confirm status.