An Azure service that provides an event-driven serverless compute platform.
When using the elastic premium plan in the function app and automatically scaling, why does the eventCount value decrease, while the GlobalworkerCount remains very high? Will it take a long time (several hours) to come down? What conditions need to be
When using the premium elastic plan in the function app and automatically scaling, why does the eventCount value sometimes decrease while the GlobalworkerCount remains very high? it takes a long time (several hours) to come down? What conditions need to be met?
Azure Functions
-
Anonymous
2025-11-07T16:22:44.7366667+00:00 Hello Gallatin 21V,
Thanks for posting your question in Microsoft Q&A forum
The reason why the eventCount value sometimes decreases while the GlobalworkerCount remains very high in an Azure Functions app on the Elastic Premium plan is due to the scale-in behavior and design for performance stability
why does the eventCount value sometimes decrease while the GlobalworkerCount remains very high?
- The eventCount represents incoming event triggers, and it can drop quickly when the workload decreases. GlobalworkerCount is the number of running function host instances and it does not scale down immediately once events decrease.
- The Elastic Premium plan maintains a minimum number of always ready instances to prevent cold starts and improve responsiveness, so GlobalworkerCount stays high until those instances can be safely shut down.
- Scale-in is gradual and conservative the platform waits to ensure that workload remains low and all ongoing function executions complete before reducing instances. This behavior can lead to several hours delay for GlobalworkerCount to come down after eventCount drops.
- Conditions for GlobalworkerCount to decrease include sustained low eventCount, no pending executions, and completion of any draining or cleanup tasks on instances. The scaling algorithm balances quick scale-out for performance with slower scale-in to avoid cold start latency and unstable scaling oscillations.
eventCount reflects dynamic workload changes, while GlobalworkerCount reflects a more stable state including always-ready instances and ongoing scale-in delays to optimize responsiveness and reliability.
- Grace Periods: Functions can take time to process their current executions before scaling down. There can be a grace period (up to 60 minutes for Premium plans) during which the instances remain active even though events aren’t arriving.
- Current Load Management: The scaling decisions are influenced not only by how many events are coming in but also by how quickly your function app can process these events. If functions are running slower due to coding practices or inefficiencies, it will affect scaling.
it takes a long time (several hours) to come down?
yes, it can take several hours depending on the workload and instance warm-up/cool-down requirements.The delay in scaling down GlobalworkerCount on the Elastic Premium plan occurs because the system waits for sustained low demand, ensures no active executions are running, and maintains a minimum number of always-ready instances to prevent cold starts. This gradual scale-in prioritizes performance, responsiveness, and workload stability.
What conditions need to be met?
- The load (events/invocations) must remain low for a sustained time.
- All active function invocations should complete, and no new events should trigger an increase.
- The platform needs to detect the idle state of instances and safely drain them.
- Scaling down also tends to be more conservative than scaling out to maintain responsiveness and avoid cold starts.
- Minimum instance threshold the scale-in will not reduce instances below the configured minimum pre-warmed or always-ready instance count.
Refer document:
- https://learn.microsoft.com/en-us/azure/azure-functions/event-driven-scaling?tabs=azure-cli
- https://learn.microsoft.com/en-us/azure/azure-functions/functions-premium-plan?tabs=portal
To assist you better, here are some questions that might help clarify your issue further:
- What type of events are you processing (e.g., Queue messages, HTTP requests)?
- Have you recently made changes to your function code or the number of events being processed?
- Could you share your current scaling configuration settings for the function app?
- Are there specific instances or times when you notice higher
GlobalworkerCountdespite low activity?
I hope this helps you understand the scaling behavior better! Let me know if you have any more questions or if there's anything else you'd like to dive into.
-
txsun • 0 Reputation points
2025-11-11T02:32:42.0533333+00:00 hi @Sandhya Kommineni
When we use the Event Hub trigger, since the upstream continuously receives input, the current instance scales out more instances during the upstream peak period. However, because the partitions taken over by the instances continue to receive a small number of messages afterwards, the scaling down behavior is not successful. Sometimes it can last for several to more than ten hours. During horizontal scaling, does the scale controller request the current instance to prepare for shutdown? From the observed phenomena, this seems unreasonable. After the upstream message peak period, our instances cannot return to a reasonable number. -
Gallatin 21V • 301 Reputation points
2025-11-11T03:24:09.0366667+00:00 Thank you for your reply.
- It is the event hub trigger, which processes data from the event hub.
- There was no modification to the Event Hub and Function code.
- The Event Hub has 32 partitions, and the Function App has a minimum of one instance and a maximum of 20 instances.
- The processing speed of event hub messages by the function is at the millisecond level. We have observed that sometimes the globalWorkerCount decreases rapidly as the EventCount decreases, while at other times it takes several hours.
- The time it takes for the globalWorkerCount to decrease from its maximum value of 20 is quite long. However, when it drops to, say, 19, and then further decreases to 15, the process is relatively faster.
-
Anonymous
2025-11-11T05:30:55.1+00:00 Hello Gallatin 21V,
Where an Event Hub triggered Azure Function App on the Premium plan has a fluctuating GlobalWorkerCount (sometimes dropping slowly over hours, other times more rapidly), is influenced by several interconnected platform factors.
Key Behaviors in Your Case:
- Event Hub Partitioning: With 32 partitions, Azure Functions tries to spread workload across multiple host instances to maximize throughput and avoid hot partitions, but partitions can be reassigned quickly, so instance count isn’t strictly 1:1 with partition count.
- Scaling Policy and Burst Queueing: When EventCount drops sharply (indicating burst processing is complete), fast scale-in will occur only if the scale controller detects all instances are idle or mostly idle, and no new events are enqueued. If even a small backlog remains or if partition leases haven't been fully released, instances will linger in draining/warming states.
- Platform Scale-in Logic: At peak (20 instances), Azure’s controller is more conservative about releasing resources all at once, to prevent oscillations and to ensure there’s no unexpected surge in EventCount right after scaling in. When scaling down from lower counts (e.g., 19 to 15), the platform is more aggressive, especially if evidence shows sustained idle periods.
- Concurrency/Throughput Observations: With very fast process times (milliseconds per message) and high concurrency settings, temporarily idle workers can remain allocated just in case, leading to a delayed scale-in while the controller learns there is no more loads.
- Minimum/Maximum Setting Limits: The minimum is respected strictly, while the scale controller tries to maintain a buffer above the minimum for a period even after load drops, influenced by past load spikes and current warm state.
Why GlobalWorkerCount Scale-in Might Be Slow at Maximum
- Cold Start Avoidance: After a burst, Azure is conservative about scaling down from max, keeping workers ready to avoid performance hits if another surge comes.
- Partition Rebalancing: Scaling in too fast can lose lease on event hub partitions, causing message processing imbalance or repeated lease acquisition costs.
- Drain/Warm States: Platform must fully drain and confirm idleness of each worker, synchronize partition state, and is governed by internal eviction and cooldown windows.
- Azure’s Scale Decision Frequency: The platform only evaluates scale-in at fixed intervals (may be every 1-2 minutes), so many cycles are needed to go from 20 to the minimum
Why Mid-Range Scale-in Is Faster
- As soon as peak is breached and bursts seem unlikely, scale-in logic becomes less conservative this is why drops from 19 to 15 (and further) are much quicker, especially if Event Count is flat or dropping and partition rebalance has begun.
I hope this helps, do let me know if you have any questions
-
Gallatin 21V • 301 Reputation points
2025-11-11T06:14:09.0733333+00:00 Hi @
"fast scale-in will occur only if the scale controller detects all instances are idle or mostly idle, and no new events are enqueued. If even a small backlog remains or if partition leases haven't been fully released, instances will linger in draining/warming states."
Could you please explain how to check in the background whether all instances are idle or mostly idle? Is there any table in Kusto that can be queried? Also, how can we determine whether there is any backlog remaining and whether the partition lease has been fully released?Are there any tables available for querying?
-
Gallatin 21V • 301 Reputation points
2025-11-11T06:29:39.1033333+00:00 "fast scale-in will occur only if the scale controller detects all instances are idle or mostly idle, and no new events are enqueued. If even a small backlog remains or if partition leases haven't been fully released, instances will linger in draining/warming states."
Could you please explain how to check in the background whether all instances are idle or mostly idle? Is there any table in Kusto that can be queried?
Through functionlogs, we discovered that the function operates in a drain mode. However, this drain mode occurs when the instance is brought down. What we are concerned about is why the instance does not come down? So that's why the drain mode hasn't arrived? If it is about the backlog remaining, how do we verify it? And if it is about the partition lease not being released, how do we confirm it?
-
Anonymous
2025-11-12T07:36:48.36+00:00 Hello Gallatin 21V, Thanks for follow-up
To check in the background whether all instances of your Azure Functions are idle or mostly idle, and to investigate why an instance in drain mode does not come down, consider these points:
1.Checking Instance Idle Status in Kusto: There isn't a single built-in Kusto table explicitly marking instances as idle or mostly idle. However, you can query logs related to your function's execution and queue processing to infer idle status:
- Analyze the messages or events processed versus queued in your backlog queues (e.g., storage queues or Event Hubs).
- Check function execution counts or durations from the Function App's logs (available in Azure Monitor/Log Analytics). This can be done by querying function invocation logs and the queue metrics where the incoming events are enqueued.
2.Verifying Backlog Remaining: To verify if a backlog remains (i.e., queued events not yet processed), you can query your event source queues directly through:
- Queue length metrics in Azure Storage Queue metrics or Event Hub partitions.
- Any telemetry relating to queue length stored in Application Insights or your telemetry logs ingestible by Log Analytics. A nonzero queue length indicates a backlog that prevents scale-in.
3.Confirming Partition Lease Release: Partition leases are typically managed by the scale controller or host.
Sign in to comment