New Foundry subscription shows 0 TPM quota for all Azure OpenAI models in every region - cannot deploy gpt-4o-mini

Zhenchuan Ren 0 Reputation points
2026-06-13T10:55:51.5966667+00:00

I created a new Microsoft Foundry resource (kind AIServices, region East US) on a subscription created through the Microsoft for Startups program (application approved, Azure credits active).

Problem: I cannot deploy ANY Azure OpenAI model. When deploying gpt-4o-mini (also gpt-4o, gpt-5.x), the deploy dialog shows "insufficient quota" for every single region - all 28 supported regions are greyed out as "insufficient quota". The TPM quota appears to be 0 across all regions for all Azure OpenAI models.

Importantly, a non-OpenAI model (grok-4.3) deployed successfully with 50K TPM. This shows the subscription, billing, and account are all healthy - only the Azure OpenAI model TPM quota has not been granted.

Question: How do I get TPM quota granted for gpt-4o-mini / gpt-4o (GlobalStandard) in East US on a brand-new subscription? Is 0 quota expected by default for new subscriptions, and does it require a quota increase request, or is it provisioned automatically after some time?

Subscription type: Microsoft for Startups sponsorship. Deployment method: official Azure AI Foundry portal.

Azure OpenAI in Foundry Models
0 comments No comments

3 answers

Sort by: Most helpful
  1. Thanmayi Godithi 11,650 Reputation points Microsoft External Staff Moderator
    2026-07-02T16:19:12.8333333+00:00

    Hi Zhenchuan Ren,

    Good news — nothing is broken here. What you're seeing is expected onboarding behavior: for a new subscription, Azure OpenAI quota isn't always provisioned automatically and can legitimately start at 0 TPM per region/model until an allocation is granted. Other Foundry models like grok-4.3 draw from a separate allocation, which is exactly why they deploy while every Azure OpenAI model shows 0 — the subscription just hasn't been granted Azure OpenAI TPM capacity yet.

    Here's how to get unblocked:

    Check you can actually see the quota. In the Foundry portal → Management → Quota, confirm the OpenAI model classes show 0 / 0 in your region. (To view quota you need the Cognitive Services Usages Reader role at the subscription; to request/edit it you need Owner/Contributor — even an Owner sees nothing without the Usages Reader role, so this rules out a pure visibility issue.)

    Request the initial Azure OpenAI quota using the official form: https://aka.ms/oai/stuquotarequest. Include:

    • Subscription ID
      • Region (e.g., East US)
        • Model(s) — e.g., gpt-4o-mini or gpt-4.1-mini
          • Deployment type — Standard / Global Standard
            • Requested TPM (e.g., 10,000 for low-volume testing)
              • A note that this is initial enablement for a small POC
              Want to test immediately while you wait? Foundry provides a shared quota pool for temporary test endpoints (usage-billed), so you can prototype without waiting for the allocation to land. Use it for testing only, not production. Try another region — quota is per-region, so if your preferred region is constrained, another may already have capacity. You can compare regions in Management → Quota.

    A couple of quick questions so I can point you precisely:

    • What subscription offer type is this (Pay-as-you-go, EA, Student/Trial, Sponsorship)? Student/trial offers often start at 0 by design and need approval.
    • Which region and model(s) do you need, and roughly what TPM?
    • In Management → Quota, does the OpenAI row show 0 / 0, or blank/no data at all?

    Was this answer helpful?

    0 comments No comments

  2. Megha Ramakrishnan 415 Reputation points
    2026-06-14T16:06:35.2866667+00:00

    Hi @Zhenchuan Ren

    Yes, getting 0 TPM quota across all regions by default is expected for many non-standard or sponsored subscription types.

    Microsoft explicitly restricts default Azure OpenAI capacities on non-standard plan codes to manage severe global GPU capacity constraints.

    To get your TPM quota granted for ⁠gpt-4o-mini⁠ and ⁠gpt-4o⁠ in East US, you will need to manually request it. It will not auto-provision over time.

    1. Navigate to the Azure AI Foundry Portal.
    2. Select your project and look at the bottom left menu. Under Management, click on Quota.
    3. Filter by the Azure OpenAI provider and look for the ⁠gpt-4o-mini⁠ or ⁠gpt-4o⁠ model classes under the GlobalStandard deployment type.
    4. If you see ⁠0 / 0 TPM⁠, click the Request Quota (or Request Increase) button inline next to the model name.
    5. Fill out the deployment details:

    Region: East US (or East US 2, wherever your resource is located).

    Model Name: ⁠gpt-4o-mini⁠ / ⁠gpt-4o⁠

    Requested TPM: Enter a modest starter amount (e.g., ⁠30,000⁠ or ⁠50,000⁠ TPM).

    Note: Requesting excessively high initial amounts can trigger automatic system rejections.

    Once the quota team reviews your subscription type, they will unlock a non-zero TPM value, and the greyed-out sliders in your deployment window will instantly become active.

    Please 'Upvote'(Thumbs-up) and 'Accept' as answer if the reply was helpful. This will be benefitting other community members who face the same issue.

    Thank you!

    Was this answer helpful?

    0 comments No comments

  3. AI answer

    2026-06-13T10:56:08.91+00:00

    For a new Microsoft Foundry subscription, Azure OpenAI model quota is not always provisioned at a usable level by default. If all Azure OpenAI models show 0 TPM (insufficient quota) in every region while other Foundry models like grok-4.3 deploy successfully, the subscription has simply not been granted Azure OpenAI TPM capacity yet.

    To get TPM quota for gpt-4o-mini / gpt-4o (GlobalStandard) in East US on this new subscription, follow these steps:

    1. Confirm current quota in Foundry portal
      1. Sign in to Microsoft Foundry and ensure the New Foundry toggle is on.
      2. Select the correct project, then go to OperateQuota.
      3. On the Token per minute tab, check the entries for the Azure OpenAI models (gpt-4o, gpt-4o-mini, etc.) and confirm that the Quota allocation is 0 TPM in all regions.
    2. Request Azure OpenAI quota for the models/regions Since 0 TPM is effectively “no capacity,” request quota explicitly:
      1. In the Quota page, select Request quota in the upper-right corner.
      2. In the request, specify:
        • Subscription (the Microsoft for Startups sponsorship subscription ID).
        • Service: Azure OpenAI in Foundry Models.
        • Model(s): gpt-4o-mini / gpt-4o.
        • Region: East US (and any other regions you plan to use).
        • Required TPM per model.
      3. Alternatively, use the quota increase form at https://aka.ms/oai/stuquotarequest (linked from the documentation and the portal) to request additional TPM for the specific models and regions.
    3. After approval, allocate TPM to deployments Once the quota request is approved for the subscription and region:
      1. Return to OperateQuota.
      2. Verify that the Quota type for the relevant model now shows a non-zero total quota approved.
      3. Create or edit deployments and assign TPM to them (for example, 10K TPM for gpt-4o-mini). The assigned TPM will be deducted from the regional quota for that model.
    4. If quota still appears blocked in all regions
      • Use the quota view to confirm there is no model- or region-specific capacity issue.
      • If all Azure OpenAI models remain at 0 TPM after a successful quota request, open a support case so that backend quota allocation for this sponsorship subscription can be checked and corrected.

    In summary, 0 TPM for Azure OpenAI models on a brand-new Foundry subscription is possible and does require a quota request. Quota is not guaranteed to be provisioned automatically at a non-zero level for every subscription and model; it must be requested and then allocated to deployments.


    References:

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.