ModelDeploymentStatus Class
Definition
Important
Some information relates to prerelease product that may be substantially modified before it’s released. Microsoft makes no warranties, express or implied, with respect to the information provided here.
The runtime status of a model deployment. All fields are read-only and populated once reconciliation has started.
public class ModelDeploymentStatus : System.ClientModel.Primitives.IJsonModel<Azure.ResourceManager.ContainerServiceAIManager.Models.ModelDeploymentStatus>, System.ClientModel.Primitives.IPersistableModel<Azure.ResourceManager.ContainerServiceAIManager.Models.ModelDeploymentStatus>
type ModelDeploymentStatus = class
interface IJsonModel<ModelDeploymentStatus>
interface IPersistableModel<ModelDeploymentStatus>
Public Class ModelDeploymentStatus
Implements IJsonModel(Of ModelDeploymentStatus), IPersistableModel(Of ModelDeploymentStatus)
- Inheritance
-
ModelDeploymentStatus
- Implements
Properties
| Name | Description |
|---|---|
| CurrentReplicas |
The current number of ready replicas serving traffic. |
| DesiredReplicas |
The desired replica count reported by the controller. Equals |
| Endpoint |
The inference endpoint URL exposed by the deployment, once ready. |
| Engine |
The inference engine used to serve the model, e.g. "vllm". |
| EngineVersion |
The version of the inference engine, e.g. "0.17". |
| EstimatedProvisionTimeSeconds |
Estimated total time, in seconds, for the deployment to become ready end-to-end (GPU vm provisioning, image/weight pull, engine warm-up). |
| MaxModelLen |
The maximum model context length, in tokens, configured for this deployment. |
| PeakTokensPerMinute |
The peak tokens per minute measured by live stress test. |
| Quantization |
The quantization level applied to the model weights, e.g. "fp16", "awq-int4". |