Container Apps: Serverless GPU container startup failures

Andrew Wood 20 Reputation points
2026-09-28T21:49:22.91+00:00

Problem description

I am experiencing issues with Azure Container Apps that use serverless GPU resources. The container task gets submitted, but never pulls the container image. After 1 min KEDA retries, getting stuck in a loop. (ContainerAppReady -> RollingRevisionCompleted -> RevisionDeactivating -> ContainerAppReady -> ... )

In this state we are billed for this retry loop even though the resource never starts.

We have seen a similar issue before, which was confirmed as an internal hosting issue.

https://learn.microsoft.com/en-us/answers/questions/5572527/container-app-using-serverless-gpu-stuck-assigning

My guess is that during peak times there are low numbers of serverless A100 GPUs available, but this situation doesn't fail in a useful way - such that we could choose to back off / try alternatives rather than being billed for the startup loop.

Environment

Azure Container Apps, A100 Serverless GPU, Australia East.

Current status

I have a support ticket with more information (2609250030000171), however the automated support was unable to escalate what looks to be an internal issue.

This issue started last week and jobs consistently failed to start for ~24. hours. Yesterday it seemed to improve, but we are now seeing 30min-1hr allocation wait times.

Falling back to CPU isn't really an alternative for these workloads, ideally we would get a clear and obvious error, and trigger a service health event for visibility.

Case ID: 5bc1fc7c-89078458-4d8d2ff7-67f7-4ab9-a32c-593173fc1729

Azure Container Apps
Azure Container Apps

An Azure service that provides a general-purpose, serverless container platform.

0 comments No comments

Answer accepted by question author
Allan Solomon Mejia 9,825 Reputation points
2026-09-28T22:02:03.5266667+00:00

Hi @Andrew Wood

The behavior you described- "repeated revision allocation attempts without reaching image pull or container startup" could indicate a platform-side GPU scheduling or capacity issue. However, the state transitions alone don’t prove that A100 capacity is exhausted.

Australia East supports Azure Container Apps serverless GPUs. The service provides GPU resources on demand, automatic scaling, scale-to-zero, and per-second billing for GPU compute in use.

To strengthen support escalation, capture the platform system logs for the affected period. In the Azure portal, open the Container Apps environment or application and select Monitoring > Log stream > System. This distinguishes these platform-generated system logs from container console logs.

If Log Analytics is configured, query the system log table:

ContainerAppSystemLogs_CL
| where TimeGenerated > ago(24h)
| where ContainerAppName_s == "<container-app-name>"
| project TimeGenerated, EnvironmentName_s, RevisionName_s, Log_s
| order by TimeGenerated asc

ContainerAppSystemLogs_CL contains Container Apps platform events.

Add the following to support case 2609250030000171:

  • Exact UTC timestamps and job execution or revision names
  • Container Apps environment resource ID
  • Consumption-GPU-NC24-A100 workload-profile configuration
  • Relevant system-log entries and correlation IDs
  • Evidence that CPU workloads in the same environment start successfully
  • Cost-analysis records showing charges during the failed allocation loop

The billing documentation states that serverless resources are billed based on resources used. That requires a meter-level billing review by Azure Support.

Microsoft Q&A contributors can’t inspect the backend scheduler or access the support case. Given the duration and the prior similar incident, escalate the existing case to the Azure Container Apps engineering team for a regional capacity and billing review.

References:

Using serverless GPUs in Azure Container Apps

View log streams in Azure Container Apps

Monitor logs with Log Analytics

Azure Container Apps billing


Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.

Was this answer helpful?

1 person found this answer helpful.

0 additional answers

Sort by: Most helpful

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.