Persistent "Failed to allocate pool: Failed to delegate" Errors Across 160 Clusters

Li Zhao 20 Reputation points Microsoft Employee
2026-05-07T06:41:30.7533333+00:00

Summary

We are experiencing a platform-wide, persistent issue where pods fail to start due to the Azure CNI plugin (azure-vnet) being unable to allocate IP addresses. The error message is:

Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "...":
plugin type="azure-vnet" failed (add): Failed to allocate pool: Failed to delegate:

A subset of nodes also report a CNI plugin crash:

Failed to delegate: netplugin failed with no error message: signal: bus error

This issue affects all 160 clusters in our fleet, across all node pool types, and has been running at high volume for ~10 months without resolution.


Impact

Metric Value
Affected clusters 160
-------- --------
Affected clusters 160
Affected cluster-pool combinations 724
Events in last 48 hours 53,626
Affected nodes in last 48 hours 43,479
Events per month (recent) 500,000 – 780,000
Duration ~10 months (since July 2025)

Pods that hit this error fail to start, causing service degradation, scheduling delays, and reduced cluster capacity. The issue affects all pool types including default pools, spot pools, GPU pools, and memory-optimized pools.


Timeline

The issue was first observed at low frequency in August 2024, then escalated dramatically in July 2025:

Period Events/Month Clusters Affected Notes
2024-08 (first seen) 184 31 Sporadic, low impact
-------- -------- -------- --------
2024-08 (first seen) 184 31 Sporadic, low impact
2024-09 – 2025-06 600 – 1,700 60 – 91 Low frequency, occasional
2025-07 (inflection) 73,681 142 46x increase — suspected platform change
2025-08 437,182 133 Further 6x increase
2025-09 – 2026-05 497,000 – 784,000 128 – 160 Sustained high volume, no improvement

Error Details

Primary Error (majority of events)

Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "...":
plugin type="azure-vnet" failed (add): Failed to allocate pool: Failed to delegate:

The delegate error message is empty, indicating the CNI plugin failed silently.

Secondary Error (subset of nodes)

Failed to create pod sandbox: rpc error: code = Unknown desc = failed to setup network for sandbox "...":
plugin type="azure-vnet" failed (add): Failed to allocate pool: Failed to delegate:
netplugin failed with no error message: signal: bus error

The signal: bus error indicates the azure-vnet CNI plugin binary is crashing, possibly due to memory alignment issues or binary incompatibility with the node OS.

Concurrent high volumes of NetworkNotReady events (10,000–30,000/hour) and FailedCreatePodSandBox events were also observed, confirming widespread CNI plugin instability.


Request

  1. Root cause analysis: Investigate the azure-vnet CNI plugin's IP allocation delegation failure, particularly the signal: bus error crash.
  2. Identify the change: Determine what platform change occurred in late June / early July 2025 that caused the 46x increase in these errors.
  3. Fix or workaround: Provide a CNI plugin fix or a recommended workaround to eliminate these allocation failures.
Azure Kubernetes Service
Azure Kubernetes Service

An Azure service that provides serverless Kubernetes, an integrated continuous integration and continuous delivery experience, and enterprise-grade security and governance.


Answer accepted by question author
Anonymous
2026-06-02T05:57:09.3666667+00:00

Thank you for your patience while we worked closely with the AKS engineering team to investigate this issue.

Based on the investigation conducted by the AKS and AzSecPack engineering teams, the issue has now been isolated and the root cause identified. Although the symptoms appeared as Azure CNI failures with errors such as "Failed to allocate pool: Failed to delegate" and "signal: bus error," the underlying cause was not related to IP exhaustion or networking capacity constraints.

The investigation determined that, under specific timing conditions, a race condition within the AzSecPack CertsInUse component could leave the SymCrypt OpenSSL provider configuration file in an invalid state during node initialization. This subsequently impacted OpenSSL initialization on affected Azure Linux 3 FIPS-enabled nodes and manifested as Azure CNI failures during pod startup.

The engineering team has developed a fix that makes the configuration update process atomic, preventing the configuration file from becoming corrupted even if the process is interrupted. In addition, the Azure Linux team is evaluating further improvements to increase resiliency during the FIPS initialization process.

As a mitigation, the engineering team recommends setting AzSecKeysinuseEngineInstall=False prior to node pool creation until the permanent fix is broadly available. This mitigation has been validated by the product group and is currently considered the preferred workaround.

Was this answer helpful?

1 person found this answer helpful.

0 additional answers

Sort by: Most helpful

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.