Compare billing models across harnesses

Completed

The billing model that determines what an agent costs to build and run depends on the harness it's built on. Each harness meters consumption differently and starts charging at a different point, so managing cost begins with knowing which harness an agent runs on and how that harness bills. In this unit, you compare how the three harnesses meter consumption, including where the GitHub Copilot harness differs from the others.

What each harness is built for

Each harness targets a different kind of work, and that purpose shapes how it meters consumption.

  • The GitHub Copilot harness handles reasoning-heavy, multistep work, where an agent plans, calls tools, and orchestrates several steps to reach an outcome.
  • The standard harness runs rule-based topics and flows, where an agent follows defined conversational paths and structured logic.
  • The Copilot chat harness extends Microsoft 365 Copilot Chat with organizational knowledge, so an agent grounds responses in your content inside the Microsoft 365 Copilot Chat experience.

Billing approaches across harnesses

Across all three harnesses, Copilot Credits are the single unit of consumption. That shared currency can be misleading, because what counts as consumption changes from one harness to the next.

On the GitHub Copilot harness, consumption follows a usage-based model. It reflects the large language model (LLM) tokens an agent consumes, the tools it calls, and the harness runtime that coordinates the work.

On the standard harness, consumption is metered per event at defined rates, with a set number of credits for each interaction type, such as a generative answer or an agent action. So the question is rarely just "how many credits." It's "how does this harness count consumption, and when does it start counting?"

Because the metering model is harness-dependent, you avoid assuming that a rate you read for one harness applies to another. Matching the metering model to the harness keeps your cost reasoning accurate.

When credit accrual starts

The harnesses also differ in when the meter starts, not only in what they meter. On the standard and Copilot chat harnesses, Copilot Credits accrue at runtime, when a published agent handles interactions. Nothing accrues while you author those agents.

The GitHub Copilot harness works differently. It's the only harness that bills from build time. Copilot Credits accrue during LLM-powered creation as well as runtime execution, so authoring, previewing, and evaluating an agent all draw from your capacity before you ever publish. On this harness, the LLM tokens, the tools an agent invokes (including knowledge sources and MCP connections), and the harness runtime that coordinates them all contribute to consumption in both phases.

The authenticated business-to-employee licensing exception

Licensing changes the picture for some users. For the standard and Copilot chat harnesses, the published per-event rates don't apply to licensed Microsoft 365 Copilot users in authenticated business-to-employee (B2E) scenarios within Microsoft 365 channels such as Teams and SharePoint.

In those cases, eligible users draw on entitlements included with their Microsoft 365 Copilot subscription license rather than metered credits, so an agent serving licensed employees can behave differently on your bill than the published rates suggest. The inclusion depends on all three conditions holding together: the same agent used outside a Microsoft 365 channel, such as on a public website, meters credits even for a licensed user.

This exception has a firm boundary. It applies to the standard and Copilot chat harnesses only. It does not apply to the GitHub Copilot harness, which consumes Copilot Credits during both creation and runtime regardless of whether the user holds a Microsoft 365 Copilot license. Keeping that boundary clear prevents you from assuming a licensed audience zeroes out cost on a reasoning-heavy agent when it doesn't.

The three harnesses side by side

Now that you've seen what each harness meters, when accrual starts, and how the B2E exception applies, the following table consolidates those differences.

Harness What it meters When accrual starts B2E Microsoft 365 Copilot exception applies?
GitHub Copilot harness LLM tokens, tools (including knowledge and Model Context Protocol (MCP) connections), and the harness runtime itself During LLM-powered creation and runtime execution No
Standard harness Copilot Credits at defined per-event rates At published runtime Yes
Copilot chat harness Copilot Credits, consumption-based or included in the Microsoft 365 Copilot subscription license At runtime Yes

Note

Credit definitions and rates change over time, so this unit teaches the metering model rather than specific numbers. For the current per-event rates, the reasoning-model premium token meter, and the Microsoft 365 Copilot license inclusion details, see Billing rates and management.

The three harnesses use different metering approaches that differ in what they count, when they start, and who they exempt. Knowing which harness an agent runs on, and that the GitHub Copilot harness starts metering as soon as you begin building, lets you account for cost from the first design decision rather than discovering it at launch.