Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Fabric data engineering agent (Project Osmos) uses language model input and output tokens during autonomous data engineering tasks. The consumption rate converts reported token usage into Fabric capacity units (CUs).
Important
The data engineering agent (Project Osmos) is in preview. Preview features are released with limited capabilities and are subject to separate supplemental preview terms. They're not intended for production use, aren't subject to service-level agreements, and might be available only in selected regions. For more information, see Microsoft Fabric preview information.
Monitor usage
In the Fabric Capacity Metrics app, language model usage appears under the Project Osmos operation for the task's lakehouse. This Background operation uses the Copilot and AI Azure billing meter.
For other Fabric operations that a task uses, see Fabric operations.
Consumption rates
The following rates are CU seconds per 1,000 tokens. The applicable billing profile depends on the model the request uses. These profiles describe consumption rates, not a list of model options available in every Fabric data engineering agent client.
| Billing profile | Uncached input | Cached input | Output |
|---|---|---|---|
| Generic | 100 | 10 | 400 |
| GPT-5.1 | 34.95762712 | 3.495762712 | 279.6610169 |
| GPT-5-mini | 6.991525424 | 0.699152542 | 55.93220339 |
| GPT-5.4, short context | 69.91525424 | 6.991525424 | 419.4915254 |
| GPT-5.4, long context | 139.8305085 | 13.98305085 | 629.2372881 |
| GPT-5.4-mini | 20.97457627 | 2.097457627 | 125.8474576 |
| GPT-5.4-nano | 5.593220339 | 0.559322034 | 34.95762712 |
| GPT-5.4-pro, short context | 838.9830508 | Not separately priced | 5033.898305 |
| GPT-5.4-pro, long context | 1677.966102 | Not separately priced | 7550.847458 |
| GPT-5.5, short context | 139.8305085 | 13.98305085 | 838.9830508 |
| GPT-5.5, long context | 279.6610169 | 27.96610169 | 1258.474576 |
| GPT-5.5-pro, short context | 838.9830508 | Not separately priced | 5033.898305 |
| GPT-5.5-pro, long context | 1677.966102 | Not separately priced | 7550.847458 |
Context length
For profiles with short-context and long-context rates, a request uses the long-context profile when its total prompt-token count is greater than 272,000. A request with exactly 272,000 prompt tokens uses the short-context profile.
The total prompt-token count determines the context tier before input usage is split into uncached and cached tokens.
Cached input
A cached-input rate applies when cached tokens are reported separately under that billing profile. Don't assume that every request reports a separate cached-input quantity. Not separately priced doesn't mean that cached input is free.
Calculate consumption
When input usage is reported with separately priced uncached and cached quantities, calculate consumption as follows:
CU seconds = (uncached input tokens * input rate
+ cached input tokens * cached rate
+ output tokens * output rate) / 1,000
Use the rates for the request's billing profile and context tier.