Skip to content
Cloud Efficiency Hub

Cross-AZ Data Transfer from Kubernetes Service Routing in EKS

The short version

Multi-AZ EKS clusters are the recommended default, but Kubernetes Service routing does not take Availability Zones into account unless told to.

PointFive Research

Cloud cost research at PointFive

AWS service
AWS EKS
Category
Networking
Reference
CER-0500
Type
Inefficient Configuration

Explanation

Why the waste happens and who it affects.

On EKS, kube-proxy distributes traffic for a Service across all of its endpoints in the cluster regardless of which node or AZ they run on, so a large share of Pod-to-Pod calls cross an AZ boundary. Load balancers managed by the AWS Load Balancer Controller in instance mode add another hop: traffic lands on a NodePort on any node and kube-proxy may forward it to a Pod in a different AZ.

For chatty microservices, service meshes, and high-throughput internal APIs, this default behavior generates large volumes of inter-AZ data transfer, which AWS bills per GB in both directions. The cost is spread across node-level data transfer line items and is rarely attributed to the Services that cause it. The EKS Best Practices Guide for cost optimization calls out both the kube-proxy default and instance-mode load balancing as sources of inter-AZ charges and recommends same-zone routing and ip target mode.

Billing model

The pricing dimensions that drive this cost.

Inter-AZ data transfer
Traffic between Pods or nodes in different Availability Zones of the same Region is billed per GB in each direction (the EC2 pricing page lists $0.01/GB each way)
Same-AZ traffic
Pod-to-Pod traffic on the same node or within the same Availability Zone has no data transfer charge
Instance target mode
Load balancer traffic to a NodePort that kube-proxy forwards to a Pod in another AZ incurs inter-AZ charges
IP target mode
The load balancer sends traffic directly to the Pod, so there is no extra cross-AZ hop from kube-proxy

How to detect

5 checks to find it in your estate.

  • In the Cost and Usage Report, track regional data transfer usage (usage types ending in DataTransfer-Regional-Bytes) on EKS worker node instances and look for clusters where it is a large share of node cost
  • Measure cross-AZ Pod-to-Pod bytes with VPC Flow Logs enriched with Pod and zone metadata, as described in the AWS post on visibility into EKS cross-AZ Pod-to-Pod network bytes, or with service mesh telemetry
  • List high-traffic Services and check whether they set spec.trafficDistribution (PreferClose or PreferSameZone) or the service.kubernetes.io/topology-mode: Auto annotation; Services with neither use cluster-wide endpoint selection
  • Check the target type of load balancers created by the AWS Load Balancer Controller (the alb.ingress.kubernetes.io/target-type annotation on Ingresses and the NLB target type on Services) and flag those using instance mode
  • Inspect EndpointSlices of Services with topology-aware routing enabled to confirm zone hints are actually assigned, since the controller skips hints when capacity across zones is too imbalanced

How to fix

6 ways to remove the waste.

  • Set trafficDistribution: PreferClose (renamed PreferSameZone in Kubernetes 1.35) on high-volume Services so kube-proxy prefers same-zone endpoints and falls back to other zones only when none are available; it is available from Kubernetes 1.30 and generally available in 1.33
  • Where even load across zones matters more than strict locality, use Topology Aware Routing (service.kubernetes.io/topology-mode: Auto), and pair either option with Pod topology spread constraints so every zone has enough replicas
  • When using same-zone routing, consider separate Deployments and HPAs per zone so a busy zone scales its own replicas instead of triggering scale-out in other zones
  • Switch AWS Load Balancer Controller targets to ip mode so load balancers send traffic straight to Pods, and deploy the load balancer across all subnets used by the cluster
  • Co-locate tightly coupled services with pod affinity, or use internalTrafficPolicy: Local for node-local services such as DaemonSet-backed agents, accepting that traffic is dropped when no local endpoint exists
  • Keep workloads multi-AZ for resilience; single-AZ node groups remove these charges but trade away availability, which the EKS guidance only suggests when that tradeoff is acceptable

Documentation

Vendor references for pricing and configuration.