How to fix AKS pod transitioning into completed state due to preemption by another critical pod
We have a service running in our cluster with its deployment configured for 2 replicas. After some time, the pods of this service get preempted to make room for higher-priority pods. When this happens, the pod transitions into a Completed state, and a new pod is created to replace it.
As a result, we continuously receive notifications indicating that the pod is not in a Running state (it shows as Completed in ArgoCD).
This behavior was initially observed only in only one service, but has now started affecting 2 more microservices as well. No other microservices or pods are exhibiting this issue.
Current status (debugging)
We are enabling AKS logs in the Log Analytics workspace to better understand which higher-priority pod(s) are preempting these affected services.
Reason for concern
- The issue has become predictable — only these three services are consistently impacted.
- We suspect a potential misconfiguration, but the root cause has not yet been identified.
Error message:
Pod was preempted by Kubelet to accommodate a critical pod.