Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Important
The Windows ML Runtime API is currently experimental and not supported for use in production environments. Apps trying out this API should not be published to the Microsoft Store.
Important
The Windows ML Runtime API is currently available in C++ and Python; C# is currently not supported. Let us know if you're interested in using these APIs from C# by posting an issue in the Windows ML GitHub repo!
The Windows ML Runtime API is a native C/C++ local AI inferencing framework built specifically for Windows. With the Windows ML Runtime API, you can load a model, choose a CPU, GPU, or NPU execution target (including a specific DXCore adapter or D3D12 device), bind tensors to that target, compose model and processor stages into a pipeline, compile the pipeline, and run it to produce output tensors.
The optional Text Generation and Speech Recognition APIs compose these same Runtime objects into higher-level workflows over a model your app supplies: see Generate text with your own language model using Windows ML and Recognize speech with your own Whisper model using Windows ML.
Runtime API vs ONNX Runtime
If your app only targets Windows (and you're willing to use APIs that are currently experimental), use the Runtime API. It gives explicit control over hardware selection (including a specific DXCore adapter or D3D12 device), tensor residency, multi-stage pipeline placement, and compilation.
Otherwise, if your app needs to run on other platforms, or already has working cross-platform ONNX Runtime session code, use the ONNX Runtime APIs included in Windows ML instead.
Core object model
| Concept | Primary API | Description |
|---|---|---|
| Runtime | IWinMLRuntime |
Interface returned by WinMLCreateRuntime; loads model files, creates execution targets, and creates pipeline builders. |
| Model | IWinMLModel |
Immutable, opaque model artifact loaded from a file or buffer. Reflect its declared schema through IWinMLModelSchema. |
| Execution target | IWinMLExecutionTarget |
CPU, GPU, or NPU placement token used to build stages and create tensors. |
| Pipeline builder | IWinMLPipelineBuilder |
Adds model and processor stages, connects them, and builds a prepared pipeline. |
| Pipeline | IWinMLPipeline |
Prepared, immediately runnable graph that executes bound tensors and produces output tensors. |
| Stage | IWinMLStage |
A step in the pipeline that owns positional input and output bindings and its resolved execution target. |
| Tensor | IWinMLTensor |
Typed multi-dimensional data bound to stage inputs and outputs. |
| Text generation / speech recognition | Windows ML Text Generation and Speech Recognition APIs | Higher-level workflow composed over caller-owned Runtime objects, with typed sessions, streams, and results. |
Choose a backend and execution target
LoadModelFromFile selects the backend from the model artifact's file extension.
Your app uses the same Runtime objects — model, execution target, pipeline
builder, stages, bindings, and Run — no matter which backend loads the file.
| Model file | Backend | Hardware acceleration |
|---|---|---|
.onnx, .ort |
ONNX Runtime | Execution providers (EPs) that Windows installs and keeps up to date through the Windows ML EP catalog |
.gguf |
llama.cpp | An app-local CPU module, plus an optional CUDA module for NVIDIA GPUs that you deploy alongside your app. CPU and CUDA are the only llama.cpp backends Windows ML validates for this release. |
An execution target represents a hardware class — CPU, GPU, or NPU — created
with IWinMLRuntime. Reuse a target to build stages and create tensors:
reusing one target expresses same-target affinity, while using distinct
targets makes cross-target transfer boundaries explicit. When the pipeline
builder's Build method runs, the runtime places each stage on the execution
backend for its target and commits target and backend compatibility.
For a hardware class without enumerating devices yourself, call
CreateExecutionTarget with a WINML_EXECUTION_TARGET_KIND and a
WINML_EXECUTION_TARGET_PREFERENCE that ranks candidates of that class, such as
an integrated versus a discrete GPU. Treat ERROR_NOT_FOUND and
ERROR_NOT_SUPPORTED alike and fall back to another class:
ComPtr<IWinMLExecutionTarget> target;
HRESULT hr = runtime->CreateExecutionTarget(
WINML_EXECUTION_TARGET_KIND_NPU,
WINML_EXECUTION_TARGET_PREFERENCE_EFFICIENCY,
target.GetAddressOf());
if (hr == HRESULT_FROM_WIN32(ERROR_NOT_FOUND) ||
hr == HRESULT_FROM_WIN32(ERROR_NOT_SUPPORTED))
{
THROW_IF_FAILED(runtime->CreateExecutionTarget(
WINML_EXECUTION_TARGET_KIND_GPU,
WINML_EXECUTION_TARGET_PREFERENCE_PERFORMANCE,
target.GetAddressOf()));
}
else
{
THROW_IF_FAILED(hr);
}
CreateCpuExecutionTarget is equivalent to requesting the CPU kind.
CreateExecutionTargetFromAdapter and CreateExecutionTargetFromD3D12 bind a
target to a specific DXCore adapter or D3D12 device instead of letting the
Runtime pick one. For ONNX models that depend on an installed execution
provider (EP), an optional ONNX Runtime
compatibility surface lets you pin a provider you have already registered.
Each backend resolves these targets differently — for example, the llama.cpp
backend used by .gguf models requires at least partial GPU offload to honor a
GPU target, and doesn't support NPU targets or provider pinning. For the full
target-resolution matrix per backend, and for a walkthrough of running a GGUF
model, see Windows ML Runtime API
concepts.