Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
In this quickstart, you run the .NET create-index sample for Azure Cosmos DB for NoSQL. The sample creates two containers with different vector index types—DiskANN and QuantizedFlat—loads a pre-vectorized hotel dataset, and compares the scores and rankings produced by Cosine, DotProduct, and Euclidean distance functions.
DiskANN and QuantizedFlat are two of the vector index types available in Azure Cosmos DB for NoSQL. For a comparison of when to use each, see Vector indexing policies.
The sample uses a JSON dataset with hotel names, regions, descriptions, and 1536-dimension vectors generated by the text-embedding-3-small model. The sample partitions the hotel documents by geographic region.
Prerequisites
- An Azure subscription. If you don't have one, create a free account.
- Azure CLI installed.
- Git installed.
- Azure Developer CLI (azd) installed.
- .NET 10.0 SDK installed.
Azure resources are created in this quickstart with the Azure Developer CLI.
Tip
Agent Kit helps coding agents work with Azure Cosmos DB quickly and efficiently using recommended best practices. To get started, run:
npx skills add AzureCosmosDB/cosmosdb-agent-kit
To learn more, see Azure Cosmos DB Agent Kit.
App dependencies
The sample uses the following NuGet package dependencies:
Azure.Identity: Authenticates with Microsoft Entra ID by usingDefaultAzureCredential.Azure.ResourceManager.CosmosDB: Creates and deletes the vector-indexed containers through Azure Resource Manager.Microsoft.Azure.Cosmos: Ingests documents and runs vector queries.Azure.AI.OpenAI: Generates vector embeddings with Azure OpenAI.Microsoft.Extensions.Configuration.EnvironmentVariables: Loads configuration from environment variables.Microsoft.Extensions.Configuration.Json: Loads configuration fromappsettings.json.Newtonsoft.Json: Serializes and deserializes JSON data.
Authenticate to Azure
The sample uses passwordless authentication through DefaultAzureCredential and Microsoft Entra ID. Sign in to Azure before you run the sample so it can access your Azure resources securely.
azd auth login
Provision the resources
Clone the sample repository.
git clone https://github.com/Azure-Samples/cosmos-db-vector-samples.git cd cosmos-db-vector-samplesCreate an Azure Developer CLI environment.
azd env new cosmos-nosqlSet the database name for the create-index samples.
azd env set AZURE_COSMOSDB_CREATE_INDEX_DATABASENAME "HotelsCreateIndex"Provision the Azure resources and role assignments.
azd upChange to the .NET sample directory.
cd nosql-create-index-dotnet
Set up environment variables
This sample uses appsettings.json to store the Azure resource and project settings. Generate the file from the active Azure Developer CLI environment.
.\scripts\generate-appsettings.ps1
The checked-in example shows the default configuration shape:
{
"CosmosDbSettings": {
"Endpoint": "https://YOUR-COSMOS-ACCOUNT.documents.azure.com:443/",
"DatabaseName": "HotelsCreateIndex",
"SubscriptionId": "",
"ResourceGroup": "",
"AccountName": "",
"Location": "eastus2",
"DiskANNContainerName": "hotels_diskann",
"QuantizedFlatContainerName": "hotels_quantizedflat",
"AllowCustomContainerDeletion": false
},
"OpenAiSettings": {
"Endpoint": "https://YOUR-OPENAI-RESOURCE.openai.azure.com/",
"Deployment": "text-embedding-3-small"
}
}
The source defaults to the hotels_diskann and hotels_quantizedflat containers. The sample deletes and recreates these containers, and then deletes them again during cleanup.
Build and run the project
Restore the project dependencies.
dotnet restoreConfirm that the generated
appsettings.jsonfile is in the sample directory.Run the sample.
dotnet run --project nosql-create-index-dotnet.csprojReview an excerpt from the recorded output.
| Index Type | Distance Function | Top 1 Result | Score | Top 2 Result | Score | Diff | |----------------|-------------------|----------------------------|--------|----------------------------|--------|--------| | DiskANN | Cosine | City Center Summer Wind... | 0.4025 | Red Tide Hotel | 0.4000 | 0.0025 | | DiskANN | DotProduct | City Center Summer Wind... | 0.4027 | Red Tide Hotel | 0.4001 | 0.0025 | | DiskANN | Euclidean | City Center Summer Wind... | 1.0934 | Red Tide Hotel | 1.0957 | -0.0023 | | QuantizedFlat | Cosine | City Center Summer Wind... | 0.4025 | Red Tide Hotel | 0.4000 | 0.0025 | | QuantizedFlat | DotProduct | City Center Summer Wind... | 0.4027 | Red Tide Hotel | 0.4001 | 0.0025 | | QuantizedFlat | Euclidean | City Center Summer Wind... | 1.0934 | Red Tide Hotel | 1.0957 | -0.0023 |
The sample demonstrates two goals:
- Control plane: Authenticate, use the existing database, and delete and recreate the two configured vector-indexed containers.
- Data plane: Load the same 50-document dataset, validate 1536-dimension embeddings, ingest documents by using
Regionas the partition key, generate a query embedding, run Cosine, DotProduct, and Euclidean queries, display ranked results and a comparison summary, and clean up the two containers.
Understand the vector distance functions
The sample's core purpose is comparing all three distance functions against the same containers and query embedding. Each function measures "closeness" differently, and that difference affects both the numeric scores and, for near ties, the ranking.
What each function measures
| Function | What it measures | Score range |
|---|---|---|
| Cosine | The angle between two vectors — magnitude-independent | -1 to 1; higher values indicate greater similarity |
| DotProduct | The inner product — accounts for both direction and magnitude | Any real number; higher values indicate greater similarity |
| Euclidean | Straight-line (L2) distance between vector endpoints — magnitude-sensitive | 0 to 2 for unit-normalized vectors; lower values indicate greater similarity |
How the example output reflects these differences
Across the sample's recorded output for the 50-hotel dataset, Cosine and DotProduct produce nearly identical scores and return the hotels in the same order. Euclidean produces scores in a different magnitude range for the same hotels.
Why Cosine and DotProduct can align closely here: This sample uses embeddings from the same model and compares the same documents across each distance function. As a result, Cosine and DotProduct often produce similar rankings. Depending on your data distribution and whether embeddings are normalized, the exact scores and ordering can still differ.
Cosine and DotProduct are similarity measures, not mathematical distance metrics. Cosine measures how closely the vectors point in the same direction. DotProduct measures that alignment while also accounting for vector magnitude. For both measures, a larger score means greater similarity.
Why Euclidean differs: Euclidean is a distance metric that measures the straight-line distance between vector endpoints. A smaller score means the vectors are closer together. The exact score ranges depend on the embeddings in your dataset.
How the distance function override works
Each query in this sample passes the function through the distanceFunction option in VectorDistance():
VectorDistance(c.embedding, @embedding, false, {'distanceFunction': 'Cosine'})
In this query, false tells Azure Cosmos DB to use the vector index instead of performing a brute-force search. The distanceFunction option tells it to calculate similarity by using cosine distance. This setting applies only to the current query—it doesn't change the distance function configured for the container. For the full signature, see the VectorDistance reference.
Tip
Match the query distance function to how your embedding model was trained. cosine is the standard default for OpenAI text embedding models, including text-embedding-3-small. Querying with a different function than the one stored in the container's vector policy still computes correctly, but might not use the vector index as efficiently.
Explore the app code
The following sections describe the main code paths in the .NET sample.
nosql-create-index-dotnet/
├── appsettings.example.json
├── nosql-create-index-dotnet.csproj
├── output/
│ └── sample-output.txt
├── scripts/
│ ├── generate-appsettings.ps1
│ └── generate-appsettings.sh
└── src/
├── Config.cs
├── ControlPlane.cs
├── DataPlane.cs
├── HotelDocument.cs
└── Program.cs
| File | Purpose |
|---|---|
src/Program.cs |
Loads configuration and orchestrates control-plane and data-plane operations. |
src/Config.cs |
Loads and validates app settings and environment variables. |
src/ControlPlane.cs |
Creates, recreates, and deletes the vector-indexed containers. |
src/DataPlane.cs |
Loads documents, generates embeddings, ingests data, and runs vector queries. |
Explore the credential and client setup
The orchestration code loads the configuration, creates one DefaultAzureCredential, and uses that credential to create the Azure Cosmos DB and Azure OpenAI clients.
var config = Config.Load();
Config.Validate(config);
var credential = new DefaultAzureCredential();
using var cosmosClient = DataPlane.CreateCosmosClient(config, credential);
var azureOpenAIClient = DataPlane.CreateAzureOpenAIClient(config, credential);
var database = cosmosClient.GetDatabase(config.DatabaseName);
You also pass the same credential to the control-plane method that creates the two containers.
Explore the control-plane container creation
The CreateContainersAsync method gets the existing database, iterates over the DiskANN and QuantizedFlat container definitions, and calls BuildContainerDefinition for each one.
public static async Task CreateContainersAsync(
SampleConfig config,
TokenCredential credential,
CancellationToken cancellationToken = default)
{
ArgumentNullException.ThrowIfNull(config);
ArgumentNullException.ThrowIfNull(credential);
var armClient = new ArmClient(credential);
var accountIdentifier = CosmosDBAccountResource.CreateResourceIdentifier(
config.SubscriptionId,
config.ResourceGroup,
config.AccountName);
var accountResource = armClient.GetCosmosDBAccountResource(accountIdentifier);
var account = await accountResource.GetAsync(cancellationToken);
var database = await account.Value.GetCosmosDBSqlDatabaseAsync(config.DatabaseName, cancellationToken);
var containers = database.Value.GetCosmosDBSqlContainers();
var embeddingPath = $"/{config.EmbeddingFieldName}";
// Create separate containers for each index type
var indexConfigs = new[]
{
(Name: config.DiskANNContainerName, IndexType: CosmosDBVectorIndexType.DiskAnn),
(Name: config.QuantizedFlatContainerName, IndexType: CosmosDBVectorIndexType.QuantizedFlat)
};
foreach (var indexConfig in indexConfigs)
{
Console.WriteLine("\n=== Step 1: Create Container with Vector Index ===");
Console.WriteLine($" Container: {indexConfig.Name}");
Console.WriteLine($" Index type: {indexConfig.IndexType}");
Console.WriteLine($" Dimensions: {config.ExpectedDimensions}");
Console.WriteLine($" Distance func: cosine (queried with all 3 metrics)");
await DeleteContainerIfExistsAsync(containers, indexConfig.Name, cancellationToken);
var start = DateTime.UtcNow;
var containerDefinition = BuildContainerDefinition(
indexConfig.Name,
indexConfig.IndexType,
embeddingPath,
config.ExpectedDimensions,
account.Value.Data.Location);
await containers.CreateOrUpdateAsync(WaitUntil.Completed, indexConfig.Name, containerDefinition, cancellationToken);
var elapsed = (DateTime.UtcNow - start).TotalSeconds;
Console.WriteLine($" Created in {elapsed:F1}s");
Console.WriteLine(" Vector index is IMMUTABLE — cannot be changed after creation");
}
}
The BuildContainerDefinition method combines the selected index type with the shared partition key, vector path, data type, dimensions, and stored distance function.
private static CosmosDBSqlContainerCreateOrUpdateContent BuildContainerDefinition(
string containerName,
CosmosDBVectorIndexType indexType,
string embeddingPath,
int dimensions,
AzureLocation location)
{
var indexingPolicy = new CosmosDBIndexingPolicy
{
IsAutomatic = true,
IndexingMode = CosmosDBIndexingMode.Consistent
};
indexingPolicy.IncludedPaths.Add(new CosmosDBIncludedPath { Path = "/*" });
indexingPolicy.ExcludedPaths.Add(new CosmosDBExcludedPath { Path = "/_etag/?" });
indexingPolicy.ExcludedPaths.Add(new CosmosDBExcludedPath { Path = $"{embeddingPath}/*" });
indexingPolicy.VectorIndexes.Add(new CosmosDBVectorIndex(embeddingPath, indexType));
var resource = new CosmosDBSqlContainerResourceInfo(containerName)
{
PartitionKey = new CosmosDBContainerPartitionKey
{
Kind = CosmosDBPartitionKind.MultiHash,
Version = 2
},
IndexingPolicy = indexingPolicy
};
resource.PartitionKey.Paths.Add(PartitionKeyPath);
resource.VectorEmbeddings.Add(
new CosmosDBVectorEmbedding(
embeddingPath,
CosmosDBVectorDataType.Float32,
VectorDistanceFunction.Cosine,
dimensions));
return new CosmosDBSqlContainerCreateOrUpdateContent(location, resource);
}
Together, these methods complete the control plane goal: they recreate hotels_diskann and hotels_quantizedflat in the database created by azd up. The containers share /Region as the partition key path and /embedding as the 1536-dimension float32 vector path. The vector policy and index are immutable, so the sample recreates both containers on every run.
Explore the document ingestion
The sample loads 50 hotel documents, associates each document with its Region partition key value, and writes the documents to both containers.
public static async Task<IngestionSummary> IngestDocumentsAsync(Container container, string containerName, IReadOnlyList<HotelDocument> documents, CancellationToken cancellationToken)
{
var existingCount = await GetExistingDocumentCountAsync(container, cancellationToken);
if (existingCount > 0)
{
return new IngestionSummary(containerName, documents.Count, 0, existingCount, true, 0);
}
var documentsByRegion = GroupDocumentsByRegion(documents);
var failures = new List<string>();
var insertedDocuments = 0;
var skippedDocuments = 0;
var totalRequestCharge = 0d;
foreach (var regionGroup in documentsByRegion)
{
var transactionalBatch = container.CreateTransactionalBatch(BuildRegionPartitionKey(regionGroup.Key));
foreach (var document in regionGroup.Value)
{
if (string.IsNullOrWhiteSpace(document.Region))
{
throw new InvalidOperationException($"Document {document.HotelId} missing Region property");
}
transactionalBatch.UpsertItem(document);
}
using TransactionalBatchResponse response = await transactionalBatch.ExecuteAsync(cancellationToken);
totalRequestCharge += response.RequestCharge;
if (response.IsSuccessStatusCode)
{
insertedDocuments += regionGroup.Value.Count;
Console.WriteLine($" Region '{regionGroup.Key}': {regionGroup.Value.Count} documents");
continue;
}
for (var operationIndex = 0; operationIndex < response.Count; operationIndex++)
{
var operationResult = response[operationIndex];
if ((int)operationResult.StatusCode is >= 200 and < 300)
{
continue;
}
failures.Add(
$"Region '{regionGroup.Key}', document '{regionGroup.Value[operationIndex].HotelId}': {operationResult.StatusCode}");
}
}
if (failures.Count > 0)
{
throw new InvalidOperationException(
$"Failed to ingest one or more region batches into {containerName}: {string.Join("; ", failures.Take(5))}");
}
return new IngestionSummary(containerName, documents.Count, insertedDocuments, skippedDocuments, false, totalRequestCharge);
}
The preceding code:
- Groups the documents by
Region. - Creates one transactional batch per region and upserts every document in that group.
- Skips ingestion when the target container already has documents.
- Reports the first failed operations and throws an exception. Returns the ingestion summary only when every batch succeeds.
Explore the vector similarity queries
The VerifyEmbeddingDimensionsAsync method calls GenerateEmbeddingAsync with probe text and confirms that Azure OpenAI returns the 1536 dimensions required by the container policy.
public static async Task<float[]> VerifyEmbeddingDimensionsAsync(AzureOpenAIClient azureOpenAIClient, SampleConfig config, CancellationToken cancellationToken)
{
var embedding = await GenerateEmbeddingAsync(azureOpenAIClient, config, "dimension check", cancellationToken);
if (embedding.Length != config.ExpectedDimensions)
{
throw new InvalidOperationException(
$"Embedding dimensions do not match the container definition. Expected {config.ExpectedDimensions}, received {embedding.Length}.");
}
return embedding;
}
public static async Task<float[]> GenerateEmbeddingAsync(AzureOpenAIClient azureOpenAIClient, SampleConfig config, string text, CancellationToken cancellationToken)
{
EmbeddingClient embeddingClient = azureOpenAIClient.GetEmbeddingClient(config.OpenAIEmbeddingDeployment);
var response = await embeddingClient.GenerateEmbeddingAsync(text, cancellationToken: cancellationToken);
return response.Value.ToFloats().ToArray();
}
The orchestration code then generates the query embedding and calls the query method for every combination of the two containers and three comparison functions.
// --- Step 3: Query with all distance functions ---
var queryEmbedding = await DataPlane.GenerateEmbeddingAsync(azureOpenAIClient, config, config.QueryText, cancellationToken);
Console.WriteLine($"\nQuery: \"{config.QueryText}\"");
Console.WriteLine($"Embedding generated ({queryEmbedding.Length} dimensions)");
Console.WriteLine($"\nRunning searches (top {config.TopCount} results for each distance function)...");
var distanceFunctions = new[] { "Cosine", "DotProduct", "Euclidean" };
var allResults = new List<(string Label, string DistanceFunction, DataPlane.QuerySummary Summary)>();
foreach (var containerName in Config.TargetContainers(config))
{
var container = database.GetContainer(containerName);
foreach (var distanceFunction in distanceFunctions)
{
var summary = await DataPlane.QueryTopMatchesAsync(container, containerName, config, queryEmbedding, distanceFunction, cancellationToken);
var label = Config.AlgorithmLabel(containerName, config);
Console.WriteLine($" ✓ {containerName} queried ({summary.RequestCharge:F2} RUs)");
allResults.Add((label, distanceFunction, summary));
}
}
The QueryTopMatchesAsync method binds the result count and query embedding as parameters, scopes the request to one partition, and reads the ranked results.
public static async Task<QuerySummary> QueryTopMatchesAsync(
Container container,
string containerName,
SampleConfig config,
IReadOnlyList<float> queryEmbedding,
string distanceFunction,
CancellationToken cancellationToken)
{
var queryDefinition = new QueryDefinition(BuildVectorDistanceQueryText(config.EmbeddingFieldName, distanceFunction))
.WithParameter("@topK", config.TopCount)
.WithParameter("@embedding", queryEmbedding);
var queryRequestOptions = new QueryRequestOptions
{
PartitionKey = BuildRegionPartitionKey(config.PartitionKeyValue)
};
using var iterator = container.GetItemQueryIterator<VectorSearchRow>(
queryDefinition,
requestOptions: queryRequestOptions);
var results = new List<QueryResult>();
var totalRequestCharge = 0d;
string? activityId = null;
while (iterator.HasMoreResults)
{
FeedResponse<VectorSearchRow> page = await iterator.ReadNextAsync(cancellationToken);
totalRequestCharge += page.RequestCharge;
activityId = page.ActivityId;
foreach (var row in page)
{
results.Add(new QueryResult(row.HotelId, row.HotelName, row.Description, row.SimilarityScore));
}
}
return new QuerySummary(containerName, totalRequestCharge, activityId, results);
}
The BuildVectorDistanceQueryText method calls ValidateEmbeddingFieldName before interpolating the field name, then builds the VectorDistance() expression used in both SELECT and ORDER BY.
public static string ValidateEmbeddingFieldName(string fieldName)
{
if (!EmbeddingFieldNamePattern().IsMatch(fieldName))
{
throw new InvalidOperationException($"Invalid embedding field name: {fieldName}");
}
return fieldName;
}
public static string BuildVectorDistanceQueryText(string embeddingFieldName, string distanceFunction)
{
var validatedFieldName = ValidateEmbeddingFieldName(embeddingFieldName);
if (!SupportedDistanceFunctions.Contains(distanceFunction, StringComparer.Ordinal))
{
throw new InvalidOperationException(
$"Unsupported distance function: {distanceFunction}. Expected one of: {string.Join(", ", SupportedDistanceFunctions)}.");
}
return
$"SELECT TOP @topK c.HotelId, c.HotelName, c.Description, " +
$"VectorDistance(c.{validatedFieldName}, @embedding, false, {{'distanceFunction': '{distanceFunction}'}}) AS SimilarityScore " +
$"FROM c " +
$"ORDER BY VectorDistance(c.{validatedFieldName}, @embedding, false, {{'distanceFunction': '{distanceFunction}'}})";
}
These steps complete the data-plane comparison goal by running the same query embedding with Cosine, DotProduct, and Euclidean against both vector index types. For guidance on choosing a distance function for your own data, see the vector distance function overview earlier in this article.
Explore the single-partition query pattern
The sample scopes each query to the configured Region value, which defaults to Northeast, through the SDK partition key option.
| Mechanism | How it works | Sample code |
|---|---|---|
| SDK partition key | Sets QueryRequestOptions.PartitionKey so the request targets the configured partition key value. |
PartitionKey = BuildRegionPartitionKey(config.PartitionKeyValue) |
| Comparison | Single-partition query | Cross-partition query |
|---|---|---|
| Query scope | One configured region. | All regions. |
| Documents in the supplied dataset | 10 documents in the default Northeast region. |
50 documents across all regions. |
| Routing | Uses the SDK partition key value. | Requires cross-partition fan-out. |
You can alternatively add a WHERE c.Region = @region predicate to the SQL statement. This sample uses the SDK partition key option instead.
An ORDER BY VectorDistance(...) clause is required for nearest-neighbor ranking. The expression repeats the same VectorDistance(...) call used in the SELECT clause.
View and manage data in Visual Studio Code
Use the Azure Databases extension for Visual Studio Code to connect to your Azure Cosmos DB account and browse the hotels_diskann and hotels_quantizedflat containers.
- In Visual Studio Code, select the Azure icon in the Activity Bar.
- Under Resources, expand Azure Cosmos DB, and locate your account.
- Expand your account > HotelsCreateIndex > hotels_diskann or hotels_quantizedflat.
- Select a document to view its hotel fields and the
embeddingvector array.
Note
The sample deletes both containers at the end of each run. To retain the containers for inspection, comment out the ControlPlane.CleanupContainersAsync call in src/Program.cs before you run the sample.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
InvalidOperationException: Missing required environment variables |
The generated configuration is missing or incomplete. | Regenerate appsettings.json and verify that it contains all required Azure resource values. |
DefaultAzureCredential can't authenticate |
The current shell doesn't have an authenticated Azure identity. | Run azd auth login, and then rerun the sample. |
| 403 response during container creation | The identity lacks control-plane access. | Verify that the identity has permission to create and delete Azure Cosmos DB containers through Azure Resource Manager. |
| 403 response during document ingestion or query | The identity lacks data-plane access. | Verify that the identity has Cosmos DB Built-in Data Contributor. Role assignments can take several minutes to propagate. |
| Azure OpenAI authorization error | The identity lacks access to the embedding deployment. | Verify that the identity has Cognitive Services OpenAI User on the Azure OpenAI resource. |
| Embedding dimension mismatch | The embedding deployment returns a vector size other than 1536. | Verify that AZURE_OPENAI_EMBEDDING_DEPLOYMENT identifies the expected text-embedding-3-small deployment. |
| Database or container configuration error | The configured database doesn't exist or the two container names are identical. | Verify the database and container settings, and then rerun the sample. |
Clean up resources
The sample automatically deletes its vector-indexed containers at the end of each run. To remove the remaining Azure resources that the sample provisions, run azd down.
azd down