Dear all, if you’ve been using any GPT 5.6 model, please do yourselves a favour and check your usage! Cost management for this model series wasn’t working until yesterday, and Azure has now adjusted the costs without prior notice. Limits and other controls weren’t in place at all until yesterday, and now people are facing a massive cost trap with inaccurate cost analyses!
The absurd thing is: on Reddit, some people are reporting that their cost analysis is showing significantly higher charges for API usage than they should be. We’re affected too. We were charged several hundred euros a day, but we were able to verify our actual usage via the logs, and it’s significantly lower than that (about 6-8 times)!
On average, we have 7–8 percentage of new input tokens, and the rest are cache tokens. Our software is well cache token optimized. Cached tokens should cost slightly more than input tokens. Cache tokens would therefore have to cost a little bit more than the input tokens, as they are 10 times cheaper but account for over 90 percentage of usage.
We’ve been billed over 3000 euros for input tokens and only 300 euros for cached tokens. The only thing they calculated correctly were the output tokens.
Furthermore Azure has charged for cache write tokens, even though they weren’t returned by the APIs at all. For anyone who charges their customers for these, good luck getting that money back.
It’s outrageous what Microsoft is getting up to here! Please contact your billing support team straight away via the Azure Portal and, if possible, check your actual usage before you end up being charged 5–10 times what you’ve actually used.