NetworkPolicy is one of the best tools to reduce blast radius, but it can be painful if you flip “default deny” too early.

Start with rollout risk

NetworkPolicy is not hard because the YAML is long. It is hard because one effective policy can immediately look like “the service is down.” The common failure shapes are:

  • default deny blocks traffic from the ingress controller or gateway
  • egress policy blocks DNS, so every external hostname starts timing out
  • unstable labels make the policy select the wrong Pods, or no Pods at all
  • multiple policies select the same Pod and the combined allow-list is wider than expected

Do not roll this out as one large change. Pick one namespace, start with ingress default deny, then add entry points, DNS, telemetry, and external APIs as explicit dependencies.

Evidence before packet captures

kubectl get netpol -n <ns>
kubectl describe netpol -n <ns> <policy>
kubectl get pod -n <ns> --show-labels
kubectl get endpointslice -n <ns>
kubectl exec -n <ns> <pod> -- nslookup kubernetes.default

First confirm which Pods the policy selects. Then confirm the direction of traffic. Only after that should you reach for CNI logs or packet captures.

Prerequisite

NetworkPolicy enforcement depends on your CNI:

  • Some CNIs enforce it by default
  • Some require explicit enablement

Verify in your environment before relying on policies.

Understand the default behavior: policies are allow-lists

NetworkPolicy is not a firewall rule engine that “adds blocks on top of allows”. It behaves more like this:

  • If no policy selects a Pod for a given direction (Ingress/Egress), that direction is allowed by default.
  • Once a Pod is selected by a policy of a given type, traffic is denied by default and only traffic explicitly allowed by policies is permitted.

When Pod A connects to Pod B, and A is isolated for Egress while B is isolated for Ingress, both the source Egress rules and destination Ingress rules must allow the connection. Changing only one side still leaves the flow blocked. Reply traffic for an allowed connection does not require a separate reverse rule.

This is why “default deny” uses podSelector: {} to select every Pod in a namespace:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-ingress
  namespace: app
spec:
  podSelector: {}
  policyTypes:
    - Ingress

This isolates incoming traffic only. Egress stays open until a policy selects those Pods for Egress.

Ingress design: start with “who can talk to me”

A safe rollout sequence:

  1. Default deny ingress
  2. Allow ingress from known entry points (ingress controller / gateway)
  3. Allow intra-namespace traffic if needed (service-to-service)
  4. Add tighter allow lists per workload over time

For example, allow only the ingress controller namespace to reach Pods labeled app=api on port 8080:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-from-ingress
  namespace: app
spec:
  podSelector:
    matchLabels:
      app: api
  policyTypes: [Ingress]
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: ingress-nginx
      ports:
        - protocol: TCP
          port: 8080

Confirm the namespace label exists before relying on it.

When one from or to item contains both a namespaceSelector and a podSelector, they are ANDed: only matching Pods inside matching namespaces are selected. Splitting them into two list items makes them OR alternatives and can allow the entire namespace plus matching Pods elsewhere. One indentation change can materially widen access.

Allow traffic within the namespace (common microservice need)

If services talk to each other inside the same namespace, add:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-same-namespace
  namespace: app
spec:
  podSelector: {}
  policyTypes: [Ingress]
  ingress:
    - from:
        - podSelector: {}

Then tighten with labels later (for example, only allow app=frontend to reach app=api).

Egress design: treat it as a migration project

Egress is where you break things you didn’t know you depended on:

  • DNS
  • metrics / tracing exporters
  • time sync, certificate fetching, external APIs
  • cloud metadata endpoints (in some environments)

A concrete approach:

  1. Turn on egress default deny for a single namespace (or one workload).
  2. Allow DNS (CoreDNS) explicitly.
  3. Add specific egress rules for known dependencies.
  4. Observe failures and iterate.

Better DNS allow rule (target CoreDNS pods)

Allowing egress to the whole kube-system namespace is broader than necessary. A tighter approach is to allow egress to CoreDNS pods by label.

The label differs across clusters; a common one is k8s-app=kube-dns:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-dns-egress
  namespace: app
spec:
  podSelector: {}
  policyTypes: [Egress]
  egress:
    - to:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: kube-system
          podSelector:
            matchLabels:
              k8s-app: kube-dns
      ports:
        - protocol: UDP
          port: 53
        - protocol: TCP
          port: 53

CoreDNS labels vary by cluster. With NodeLocal DNSCache, Pods may send queries to a node-local address instead of directly to CoreDNS Pods in kube-system. Confirm the real destination from /etc/resolv.conf, the DNS Service or DaemonSet, and the CNI documentation before enforcing egress.

External egress: start with allowlists

If your services need to call external APIs, you can allow by CIDR using ipBlock:

egress:
  - to:
      - ipBlock:
          cidr: 203.0.113.0/24
    ports:
      - protocol: TCP
        port: 443

Notes:

  • ipBlock is IP-based, not DNS-based. Many public APIs change IPs, so consider whether you need a stable egress proxy or NAT with fixed ranges.
  • Some CNIs have limitations around ipBlock and certain traffic patterns.
  • Services, ingresses, and cloud load balancers may apply SNAT or DNAT before or after policy evaluation, so different implementations can observe different source or destination IPs. Do not use ipBlock to select Pod IPs.

Testing and debugging (what to do when traffic is blocked)

When a request fails after applying policies:

  1. Confirm which policies select the Pod:
kubectl get netpol -n app
kubectl describe netpol -n app <name>
  1. Test connectivity from inside the Pod network namespace:
  • if you have tools: curl, wget, nc, dig
  • if not: use an ephemeral container to run tooling safely
  1. Check whether your CNI provides policy logs/flow logs.

Flow logs can turn “it’s blocked” into “it’s blocked because rule X doesn’t match label Y”.

Label strategy: your policies are only as good as your labels

NetworkPolicy selectors are label-based, so treat labels as API:

  • avoid constantly changing labels used by security policies
  • standardize label keys (app, component, tier, etc.)
  • apply namespace labels consistently (so namespaceSelector rules keep working)

If you rely on:

namespaceSelector:
  matchLabels:
    kubernetes.io/metadata.name: ingress-nginx

make sure those labels exist in your cluster (most modern Kubernetes versions apply them automatically).

“Why is it still allowed?”: multiple policies can combine

Remember that multiple policies can select the same Pod. The effective allowed traffic is the union of all allowed rules for that direction.

This is useful (layered policies), but it can also confuse debugging:

  • an “allow all from namespace X” policy can widen the final union and make a newer restriction appear ineffective

When troubleshooting, always list all policies that select the Pod.

Day-2 operating checklist

  • Change one namespace or workload at a time, observe traffic, then expand.
  • Exercise real traffic in staging instead of testing only health endpoints.
  • Maintain baseline policies for DNS, metrics, tracing, log collection, and approved entry points.
  • Keep reducing unnecessary connections without making the service impossible to operate.

Pre-rollout checklist

  • Confirm that the CNI actually enforces NetworkPolicy.
  • Apply ingress default deny first, then add known entry points.
  • Roll out egress separately, starting with DNS and known dependencies.
  • Keep labels stable because policy selectors are part of the security boundary.
  • Pilot the change instead of applying one namespace-wide cutover everywhere.

Four checks before enforcement

Q: Why is my policy ignored? A: NetworkPolicies require a compatible CNI plugin. If the cluster has no provider, policies are inert.

Q: How do I allow DNS? A: Read the actual DNS address from the Pod’s /etc/resolv.conf, then allow TCP/UDP 53 to the corresponding CoreDNS Pods or NodeLocal DNS path.

Q: Does every Pod get isolated? A: Only Pods selected by at least one policy are isolated; others remain default-allow.

Q: Why did the application fail immediately after adding an egress policy? A: DNS, telemetry, and external APIs all need egress. If those dependencies are not explicitly allowed first, name resolution and outbound calls begin timing out immediately.

References