Azure AI Foundry / Azure OpenAI Monitoring Reconciliation Bug
We are reporting a monitoring / billing reconciliation failure in Azure AI Foundry / Azure OpenAI that materially affected our ability to estimate and control cost.
Impact
We relied on Azure monitoring surfaces to understand token consumption and expected spend across our AI Foundry deployments. The numbers exposed by Azure did not reconcile with our real-time usage experience or with Azure billing exports, and our actual cost was materially higher than expected.
This was not a small rounding issue. For several GPT-family deployments, the discrepancy is large enough that it undermines confidence in Azure’s monitoring for operational cost control.
Problem statement
For the same subscription and overlapping May 2026 usage window, we observed a mismatch between:
- real-time operational experience using our Azure AI Foundry deployments,
- Azure Monitor deployment-level usage metrics, and
- Azure billing CSV exports.
These three views do not reconcile.
The discrepancy is especially visible on GPT-family models. Some simpler categories, such as embeddings and Kimi input, align much more closely, which suggests the issue is concentrated in GPT-family monitoring / billing attribution rather than our extraction logic.
Subscription / scope