A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
Hello Romir Chekuri,
Welcome to the Microsoft Q&A and thank you for posting your questions here.
I understand that your DeepSeek-V4-Pro on Foundry: every request fails with 413 "Max size: 4000 tokens" per-request cap, not TPM.
Regarding your questions:
Is the 4000-token per-request limit on DeepSeek-V4-Pro Global Standard deployments intentional, and can it be raised for a subscription/deployment? If it's a backend default, what is the correct channel to request an increase, since the quota form only covers TPM? I've seen the same class of issue reported for other Foundry serverless models (e.g. an 18K cap on Mistral), so a pointer to the right escalation path would help others too.
- It is being enforced by the deployed Foundry endpoint. Whether it is an intentional service limit or a backend defect cannot be determined by the customer and requires Microsoft Support to confirm.
- There is no documented self-service method to increase the per-request input limit for Azure AI Foundry serverless model deployments.
- No. TPM controls throughput and cannot change a 413 request-size validation limit.
- No. It only handles TPM/RPM capacity, not per-request payload limits.
Since you've opened an Azure Support technical case, is the best they will fix 4,000-token rejection.
I hope this is helpful. Please! Do not hesitate to let me know if you have any other questions, steps or clarifications.
Please do not close the thread by upvoting and accepting the answer if any part of it is helpful.