Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
In this quickstart, you run the Python create-index sample for Azure Cosmos DB for NoSQL. The sample creates two containers with different vector index types—DiskANN and QuantizedFlat—loads a pre-vectorized hotel dataset, and compares the scores and rankings produced by Cosine, DotProduct, and Euclidean distance functions.
DiskANN and QuantizedFlat are two of the vector index types available in Azure Cosmos DB for NoSQL. For a comparison of when to use each, see Vector indexing policies.
The sample uses a JSON dataset with hotel names, regions, descriptions, and 1536-dimension vectors generated by the text-embedding-3-small model. The sample partitions the hotel documents by geographic region.
Prerequisites
- An Azure subscription. If you don't have one, create a free account.
- Azure CLI installed.
- Git installed.
- Azure Developer CLI (azd) installed.
- Python 3.10 or later installed.
Azure resources are created in this quickstart with the Azure Developer CLI.
Tip
Agent Kit helps coding agents work with Azure Cosmos DB quickly and efficiently using recommended best practices. To get started, run:
npx skills add AzureCosmosDB/cosmosdb-agent-kit
To learn more, see Azure Cosmos DB Agent Kit.
App dependencies
The sample uses the following Python package dependencies:
azure-identity: Authenticates with Microsoft Entra ID by usingDefaultAzureCredential.azure-mgmt-cosmosdb: Creates and deletes the vector-indexed containers through Azure Resource Manager.azure-cosmos: Ingests documents and runs vector queries.openai: Generates vector embeddings with Azure OpenAI.
Authenticate to Azure
The sample uses passwordless authentication through DefaultAzureCredential and Microsoft Entra ID. Sign in to Azure before you run the sample so it can access your Azure resources securely.
azd auth login
Provision the resources
Clone the sample repository.
git clone https://github.com/Azure-Samples/cosmos-db-vector-samples.git cd cosmos-db-vector-samplesCreate an Azure Developer CLI environment.
azd env new cosmos-nosqlSet the database name for the create-index samples.
azd env set AZURE_COSMOSDB_CREATE_INDEX_DATABASENAME "HotelsCreateIndex"Provision the Azure resources and role assignments.
azd upChange to the Python sample directory.
cd nosql-create-index-python
Set up environment variables
This sample uses a local .env file to store the Azure resource and project settings. Generate the file from the active Azure Developer CLI environment.
azd env get-values > .env
Load the values into the current shell because the Python app reads process environment variables and doesn't load .env automatically.
Get-Content .env | Where-Object { $_ -match '^[^#].*=' } | ForEach-Object { $k,$v = $_ -split '=',2; [Environment]::SetEnvironmentVariable($k.Trim(), $v.Trim().Trim('"').Trim("'")) }
The checked-in example shows the default configuration shape:
AZURE_COSMOSDB_ENDPOINT=https://YOUR-COSMOS-ACCOUNT.documents.azure.com:443/
AZURE_COSMOSDB_CREATE_INDEX_DATABASENAME=HotelsCreateIndex
AZURE_OPENAI_EMBEDDING_ENDPOINT=https://YOUR-OPENAI-RESOURCE.openai.azure.com/
AZURE_OPENAI_EMBEDDING_DEPLOYMENT=text-embedding-3-small
AZURE_SUBSCRIPTION_ID=
AZURE_RESOURCE_GROUP=
AZURE_COSMOSDB_ACCOUNT_NAME=
AZURE_LOCATION=eastus2
DATA_FILE_WITH_VECTORS_AND_REGIONS=./data/HotelsData_toCosmosDB_Vector_byRegion.json
The source defaults to the hotels_diskann and hotels_quantizedflat containers. The sample deletes and recreates these containers, and then deletes them again during cleanup.
Build and run the project
Create a virtual environment and install the project dependencies.
python -m venv .venv .\.venv\Scripts\Activate.ps1 pip install -r requirements.txtIf you opened a new terminal after setup, activate the virtual environment and load the
.envvalues into that shell again.Run the sample.
python -m src.indexReview an excerpt from the recorded output.
| Index Type | Distance Function | Top 1 Result | Score | Top 2 Result | Score | Diff | |----------------|-------------------|----------------------------|--------|----------------------------|--------|--------| | DiskANN | Cosine | City Center Summer Wind Re | 0.4025 | Red Tide Hotel | 0.4000 | 0.0025 | | DiskANN | DotProduct | City Center Summer Wind Re | 0.4027 | Red Tide Hotel | 0.4001 | 0.0025 | | DiskANN | Euclidean | City Center Summer Wind Re | 1.0934 | Red Tide Hotel | 1.0957 | -0.0023 | | QuantizedFlat | Cosine | City Center Summer Wind Re | 0.4025 | Red Tide Hotel | 0.4000 | 0.0025 | | QuantizedFlat | DotProduct | City Center Summer Wind Re | 0.4027 | Red Tide Hotel | 0.4001 | 0.0025 | | QuantizedFlat | Euclidean | City Center Summer Wind Re | 1.0934 | Red Tide Hotel | 1.0957 | -0.0023 |
The sample demonstrates two goals:
- Control plane: Authenticate, use the existing database, and delete and recreate the two configured vector-indexed containers.
- Data plane: Load the same 50-document dataset, validate 1536-dimension embeddings, ingest documents by using
Regionas the partition key, generate a query embedding, run Cosine, DotProduct, and Euclidean queries, display ranked results and a comparison summary, and clean up the two containers.
Understand the vector distance functions
The sample's core purpose is comparing all three distance functions against the same containers and query embedding. Each function measures "closeness" differently, and that difference affects both the numeric scores and, for near ties, the ranking.
What each function measures
| Function | What it measures | Score range |
|---|---|---|
| Cosine | The angle between two vectors — magnitude-independent | -1 to 1; higher values indicate greater similarity |
| DotProduct | The inner product — accounts for both direction and magnitude | Any real number; higher values indicate greater similarity |
| Euclidean | Straight-line (L2) distance between vector endpoints — magnitude-sensitive | 0 to 2 for unit-normalized vectors; lower values indicate greater similarity |
How the example output reflects these differences
Across the sample's recorded output for the 50-hotel dataset, Cosine and DotProduct produce nearly identical scores and return the hotels in the same order. Euclidean produces scores in a different magnitude range for the same hotels.
Why Cosine and DotProduct can align closely here: This sample uses embeddings from the same model and compares the same documents across each distance function. As a result, Cosine and DotProduct often produce similar rankings. Depending on your data distribution and whether embeddings are normalized, the exact scores and ordering can still differ.
Cosine and DotProduct are similarity measures, not mathematical distance metrics. Cosine measures how closely the vectors point in the same direction. DotProduct measures that alignment while also accounting for vector magnitude. For both measures, a larger score means greater similarity.
Why Euclidean differs: Euclidean is a distance metric that measures the straight-line distance between vector endpoints. A smaller score means the vectors are closer together. The exact score ranges depend on the embeddings in your dataset.
How the distance function override works
Each query in this sample passes the function through the distanceFunction option in VectorDistance():
VectorDistance(c.embedding, @embedding, false, {'distanceFunction': 'Cosine'})
In this query, false tells Azure Cosmos DB to use the vector index instead of performing a brute-force search. The distanceFunction option tells it to calculate similarity by using cosine distance. This setting applies only to the current query—it doesn't change the distance function configured for the container. For the full signature, see the VectorDistance reference.
Tip
Match the query distance function to how your embedding model was trained. cosine is the standard default for OpenAI text embedding models, including text-embedding-3-small. Querying with a different function than the one stored in the container's vector policy still computes correctly, but might not use the vector index as efficiently.
Explore the app code
The following sections describe the main code paths in the Python sample.
nosql-create-index-python/
├── .env.example
├── output/
│ └── sample-output.txt
├── requirements.txt
└── src/
├── config.py
├── control_plane.py
├── data_plane.py
└── index.py
| File | Purpose |
|---|---|
src/index.py |
Loads configuration and orchestrates control-plane and data-plane operations. |
src/config.py |
Loads and validates environment variables. |
src/control_plane.py |
Creates, recreates, and deletes the vector-indexed containers. |
src/data_plane.py |
Loads documents, generates embeddings, ingests data, and runs vector queries. |
Explore the credential and client setup
The orchestration code loads the configuration, creates one DefaultAzureCredential, and passes it to the data-plane client factories.
config = load_config()
validate_config(config)
credential = DefaultAzureCredential()
cosmos_client = create_cosmos_client(config, credential)
openai_client = create_azure_openai_client(config, credential)
database = cosmos_client.get_database_client(config.database_name)
The create_cosmos_client and create_azure_openai_client functions use that shared credential to create the two clients needed for document operations, embedding generation, and vector queries.
def create_cosmos_client(config: SampleConfig, credential: DefaultAzureCredential) -> CosmosClient:
return CosmosClient(url=config.cosmos_endpoint, credential=credential)
def create_azure_openai_client(
config: SampleConfig, credential: DefaultAzureCredential
) -> AzureOpenAI:
token_provider = get_bearer_token_provider(
credential, "https://cognitiveservices.azure.com/.default"
)
return AzureOpenAI(
azure_endpoint=config.openai_embedding_endpoint,
azure_ad_token_provider=token_provider,
api_version=config.openai_embedding_api_version,
)
The same credential is also passed to the control-plane function that creates the two containers.
Explore the control-plane container creation
The create_containers function iterates over the DiskANN and QuantizedFlat definitions and calls _build_container_create_parameters for each container.
def create_containers(credential: DefaultAzureCredential, config: SampleConfig) -> None:
"""Create SQL containers with vector indexes using control-plane (ARM) SDK.
Args:
credential: Azure credential for authentication
config: Sample configuration with resource details
"""
from .config import validate_container_deletion_targets
validate_container_deletion_targets(config)
client = CosmosDBManagementClient(
credential=credential,
subscription_id=config.subscription_id
)
embedding_path = f"/{config.embedding_field_name}"
containers_config = [
{"type": "diskANN", "container_name": config.diskann_container_name},
{"type": "quantizedFlat", "container_name": config.quantizedflat_container_name},
]
for container_config in containers_config:
print(f"\n=== Phase 1: Create Container with Vector Index ===")
print(f" Container: {container_config['container_name']}")
print(f" Index type: {container_config['type']}")
print(f" Dimensions: {config.expected_dimensions}")
print(f" Distance func: cosine (queried with all 3 metrics)")
# Delete existing container to ensure clean state (idempotent)
try:
client.sql_resources.begin_delete_sql_container(
resource_group_name=config.resource_group,
account_name=config.account_name,
database_name=config.database_name,
container_name=container_config["container_name"]
).result()
print(f" Deleted existing container")
except HttpResponseError as e:
if e.status_code == 404:
print(f" Container does not exist (OK)")
else:
raise
params = _build_container_create_parameters(
container_name=container_config["container_name"],
partition_key_path="/Region",
embedding_field=embedding_path,
dimensions=config.expected_dimensions,
index_type=container_config["type"],
location=config.location,
)
client.sql_resources.begin_create_update_sql_container(
resource_group_name=config.resource_group,
account_name=config.account_name,
database_name=config.database_name,
container_name=container_config["container_name"],
create_update_sql_container_parameters=params
).result()
print(f" Created in ~1s")
print(f" Vector index is IMMUTABLE — cannot be changed after creation")
The _build_container_create_parameters helper combines the selected index type with the shared partition key, vector path, data type, dimensions, and stored distance function.
def _build_container_create_parameters(
container_name: str,
partition_key_path: str,
embedding_field: str,
dimensions: int,
index_type: str,
location: str,
) -> SqlContainerCreateUpdateParameters:
resource = SqlContainerResource(
id=container_name,
partition_key=ContainerPartitionKey(
paths=[partition_key_path],
kind="Hash",
),
indexing_policy=IndexingPolicy(
indexing_mode="Consistent",
automatic=True,
included_paths=[IncludedPath(path="/*")],
excluded_paths=[
ExcludedPath(path="/_etag/?"),
ExcludedPath(path=f"{embedding_field}/*"),
],
vector_indexes=[
VectorIndex(path=embedding_field, type=index_type),
],
),
vector_embedding_policy=VectorEmbeddingPolicy(
vector_embeddings=[
VectorEmbedding(
path=embedding_field,
data_type="float32",
dimensions=dimensions,
distance_function="cosine",
),
],
),
)
return SqlContainerCreateUpdateParameters(
resource=resource,
location=location,
)
Together, these functions complete the control-plane goal: they recreate hotels_diskann and hotels_quantizedflat in the database created by azd up. The containers share /Region as the partition key path and /embedding as the 1536-dimension float32 vector path. The vector policy and index are immutable, so the sample recreates both containers on every run.
Explore the document ingestion
The ingest_documents function associates each of the 50 hotel documents with its Region partition key and writes the same data to both containers.
def ingest_documents(container: Any, container_name: str, documents: Sequence[Dict[str, Any]]) -> IngestionSummary:
existing_count = _document_count(container)
if existing_count > 0:
return IngestionSummary(
container_name=container_name,
total_documents=len(documents),
inserted_documents=0,
skipped=True,
request_charge=0.0,
)
# Validate Region property and group by region
_validate_region_property(documents)
docs_by_region = _group_by_region(documents)
inserted_documents = 0
failed_documents = 0
total_request_charge = 0.0
# Batch ingest by region (one batch per region)
for region, region_docs in docs_by_region.items():
operations = [("upsert", (document,)) for document in region_docs]
results = container.execute_item_batch(
batch_operations=operations,
partition_key=region,
)
for result in results:
if int(result.get("statusCode", 0)) < 300:
inserted_documents += 1
else:
failed_documents += 1
total_request_charge += _request_charge(container)
if failed_documents > 0:
print(
"WARNING: {0} of {1} documents failed to insert.".format(
failed_documents, len(documents)
)
)
raise RuntimeError(
"Batch ingestion incomplete: {0} documents failed in container '{1}'.".format(
failed_documents, container_name
)
)
return IngestionSummary(
container_name=container_name,
total_documents=len(documents),
inserted_documents=inserted_documents,
skipped=False,
request_charge=total_request_charge,
)
The preceding code:
- Groups the documents by
Region. - Creates one transactional batch per region and upserts every document in that group.
- Skips ingestion when the target container already has documents.
- Reports the first failed operations and throws an exception. Returns the ingestion summary only when every batch succeeds.
Explore the vector similarity queries
The generate_embedding function creates an embedding for supplied text. The verify_embedding_dimensions function calls it with probe text and confirms that Azure OpenAI returns the 1536 dimensions required by the container policy.
def generate_embedding(
openai_client: AzureOpenAI, config: SampleConfig, text: str
) -> Sequence[float]:
response = openai_client.embeddings.create(
model=config.openai_embedding_deployment,
input=[text],
)
return response.data[0].embedding
def verify_embedding_dimensions(
openai_client: AzureOpenAI, config: SampleConfig
) -> Sequence[float]:
embedding = generate_embedding(openai_client, config, "dimension check")
actual_dimensions = len(embedding)
if actual_dimensions != config.expected_dimensions:
raise ValueError(
"Embedding dimensions do not match the container definition. "
"Expected {0}, received {1}.".format(
config.expected_dimensions, actual_dimensions
)
)
return embedding
The orchestration code then generates the query embedding and calls query_top_matches for every combination of the two containers and three comparison functions.
# --- Query ---
query_embedding = generate_embedding(
openai_client, config, config.query_text
)
print("\nQuery: \"{0}\"".format(config.query_text))
print("Embedding generated ({0} dimensions)".format(len(query_embedding)))
print("\nRunning searches (top {0} results for each distance function)...".format(config.top_count))
distance_functions = ["Cosine", "DotProduct", "Euclidean"]
all_results = []
for container_name in target_containers(config):
container = database.get_container_client(container_name)
for distance_function in distance_functions:
summary = query_top_matches(
container=container,
container_name=container_name,
config=config,
query_embedding=query_embedding,
distance_function=distance_function,
)
label = algorithm_label(container_name, config)
print(" ✓ {0} queried ({1:.2f} RUs)".format(container_name, summary.request_charge))
all_results.append((label, distance_function, summary))
The query_top_matches function validates the embedding field name, binds the result count and embedding as parameters, scopes the request to one partition, and reads the ranked results.
def query_top_matches(
container: Any,
container_name: str,
config: SampleConfig,
query_embedding: Sequence[float],
distance_function: str = "Cosine",
) -> QuerySummary:
embedding_field = validate_field_name(config.embedding_field_name)
# ORDER BY VectorDistance(...) is REQUIRED for nearest-neighbor ranking.
# Without it, SELECT TOP N returns N arbitrary documents, not the closest ones.
# Field names cannot be query parameters in Cosmos DB SQL; interpolate them inline.
query_text = (
"SELECT TOP @topK c.HotelId, c.HotelName, c.Description, "
"VectorDistance(c.{0}, @embedding, false, {{'distanceFunction': '{1}'}}) AS SimilarityScore "
"FROM c "
"ORDER BY VectorDistance(c.{0}, @embedding, false, {{'distanceFunction': '{1}'}})"
).format(embedding_field, distance_function)
raw_results = list(
container.query_items(
query=query_text,
parameters=[
{"name": "@topK", "value": config.top_count},
{"name": "@embedding", "value": list(query_embedding)},
],
partition_key=config.partition_key_value,
)
)
results = [
QueryResult(
hotel_id=str(item["HotelId"]),
hotel_name=str(item["HotelName"]),
description=str(item["Description"]),
score=float(item["SimilarityScore"]),
)
for item in raw_results
]
headers = getattr(container.client_connection, "last_response_headers", {}) or {}
return QuerySummary(
container_name=container_name,
request_charge=_request_charge(container),
activity_id=str(headers.get("x-ms-activity-id", "")),
results=results,
)
These steps complete the data-plane comparison goal by running the same query embedding with Cosine, DotProduct, and Euclidean against both vector index types. For guidance on choosing a distance function for your own data, see the vector distance function overview earlier in this article.
Explore the single-partition query pattern
The sample scopes each query to the configured Region value, which defaults to Northeast, through the SDK partition key option.
| Mechanism | How it works | Sample code |
|---|---|---|
| SDK partition key | Sets partition_key on query_items so the request targets the configured partition key value. |
container.query_items(..., partition_key=config.partition_key_value) |
| Comparison | Single-partition query | Cross-partition query |
|---|---|---|
| Query scope | One configured region. | All regions. |
| Documents in the supplied dataset | 10 documents in the default Northeast region. |
50 documents across all regions. |
| Routing | Uses the SDK partition key value. | Requires cross-partition fan-out. |
You can alternatively add a WHERE c.Region = @region predicate to the SQL statement. This sample uses the SDK partition key option instead.
An ORDER BY VectorDistance(...) clause is required for nearest-neighbor ranking. The expression repeats the same VectorDistance(...) call used in the SELECT clause.
View and manage data in Visual Studio Code
Use the Azure Databases extension for Visual Studio Code to connect to your Azure Cosmos DB account and browse the hotels_diskann and hotels_quantizedflat containers.
- In Visual Studio Code, select the Azure icon in the Activity Bar.
- Under Resources, expand Azure Cosmos DB, and locate your account.
- Expand your account > HotelsCreateIndex > hotels_diskann or hotels_quantizedflat.
- Select a document to view its hotel fields and the
embeddingvector array.
Note
The sample deletes both containers at the end of each run. To retain the containers for inspection, comment out the delete_containers call in src/index.py before you run the sample.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
ConfigError: Missing required environment variables |
The generated configuration is missing or isn't loaded into the current shell. | Verify the .env values, load them into the current shell, and rerun the sample. |
DefaultAzureCredential can't authenticate |
The current shell doesn't have an authenticated Azure identity. | Run azd auth login, and then rerun the sample. |
| 403 response during container creation | The identity lacks control-plane access. | Verify that the identity has permission to create and delete Azure Cosmos DB containers through Azure Resource Manager. |
| 403 response during document ingestion or query | The identity lacks data-plane access. | Verify that the identity has Cosmos DB Built-in Data Contributor. Role assignments can take several minutes to propagate. |
| Azure OpenAI authorization error | The identity lacks access to the embedding deployment. | Verify that the identity has Cognitive Services OpenAI User on the Azure OpenAI resource. |
| Embedding dimension mismatch | The embedding deployment returns a vector size other than 1536. | Verify that AZURE_OPENAI_EMBEDDING_DEPLOYMENT identifies the expected text-embedding-3-small deployment. |
| Database or container configuration error | The configured database doesn't exist or the two container names are identical. | Verify the database and container settings, and then rerun the sample. |
Clean up resources
The sample automatically deletes its vector-indexed containers at the end of each run. To remove the remaining Azure resources that the sample provisions, run azd down.
azd down