Quickstart: Create and query vector indexes in Azure Cosmos DB for NoSQL using Java

In this quickstart, you run the Java create-index sample for Azure Cosmos DB for NoSQL. The sample creates two containers with different vector index types—DiskANN and QuantizedFlat—loads a pre-vectorized hotel dataset, and compares the scores and rankings produced by Cosine, DotProduct, and Euclidean distance functions.

DiskANN and QuantizedFlat are two of the vector index types available in Azure Cosmos DB for NoSQL. For a comparison of when to use each, see Vector indexing policies.

The sample uses a JSON dataset with hotel names, regions, descriptions, and 1536-dimension vectors generated by the text-embedding-3-small model. The sample partitions the hotel documents by geographic region.

Prerequisites

Azure resources are created in this quickstart with the Azure Developer CLI.

Tip

Agent Kit helps coding agents work with Azure Cosmos DB quickly and efficiently using recommended best practices. To get started, run:

npx skills add AzureCosmosDB/cosmosdb-agent-kit

To learn more, see Azure Cosmos DB Agent Kit.

App dependencies

The sample uses the following Maven dependencies:

Authenticate to Azure

The sample uses passwordless authentication through DefaultAzureCredential and Microsoft Entra ID. Sign in to Azure before you run the sample so it can access your Azure resources securely.

azd auth login

Provision the resources

  1. Clone the sample repository.

    git clone https://github.com/Azure-Samples/cosmos-db-vector-samples.git
    cd cosmos-db-vector-samples
    
  2. Create an Azure Developer CLI environment.

    azd env new cosmos-nosql
    
  3. Set the database name for the create-index samples.

    azd env set AZURE_COSMOSDB_CREATE_INDEX_DATABASENAME "HotelsCreateIndex"
    
  4. Provision the Azure resources and role assignments.

    azd up
    
  5. Change to the Java sample directory.

    cd nosql-create-index-java
    

Set up environment variables

This sample uses a local .env file to store the Azure resource and project settings. Generate the file from the active Azure Developer CLI environment.

azd env get-values > .env

Load the values into the current shell because the Java app reads process environment variables and doesn't load .env automatically.

Get-Content .env | Where-Object { $_ -match '^[^#].*=' } | ForEach-Object { $k,$v = $_ -split '=',2; [Environment]::SetEnvironmentVariable($k.Trim(), $v.Trim().Trim('"').Trim("'")) }

The checked-in example shows the default configuration shape:

AZURE_COSMOSDB_ENDPOINT=https://YOUR-COSMOS-ACCOUNT.documents.azure.com:443/
AZURE_COSMOSDB_CREATE_INDEX_DATABASENAME=HotelsCreateIndex
AZURE_OPENAI_EMBEDDING_ENDPOINT=https://YOUR-OPENAI-RESOURCE.openai.azure.com/
AZURE_OPENAI_EMBEDDING_DEPLOYMENT=text-embedding-3-small
AZURE_SUBSCRIPTION_ID=
AZURE_RESOURCE_GROUP=
AZURE_COSMOSDB_ACCOUNT_NAME=
AZURE_LOCATION=eastus2
AZURE_COSMOSDB_CREATE_INDEX_DISKANN_CONTAINER_NAME=hotels_diskann
AZURE_COSMOSDB_CREATE_INDEX_QUANTIZEDFLAT_CONTAINER_NAME=hotels_quantizedflat
# Set to true only when intentionally allowing the sample to overwrite and delete custom container names.
AZURE_COSMOSDB_CREATE_INDEX_ALLOW_DESTRUCTIVE_OPERATIONS=false

The source defaults to the hotels_diskann and hotels_quantizedflat containers. The sample deletes and recreates these containers, and then deletes them again during cleanup.

Build and run the project

  1. Compile the project and download its dependencies.

    mvn compile
    
  2. If you opened a new terminal after setup, load the .env values into that shell again.

  3. Run the sample.

    mvn exec:java
    
  4. Review an excerpt from the recorded output.

    
    | Index Type     | Distance Function | Top 1 Result               | Score  | Top 2 Result               | Score  | Diff   |
    |----------------|-------------------|----------------------------|--------|----------------------------|--------|--------|
    | DiskANN        | Cosine            | City Center Summer Wind... | 0.4025 | Red Tide Hotel             | 0.4000 | 0.0025 |
    | DiskANN        | DotProduct        | City Center Summer Wind... | 0.4027 | Red Tide Hotel             | 0.4001 | 0.0025 |
    | DiskANN        | Euclidean         | City Center Summer Wind... | 1.0934 | Red Tide Hotel             | 1.0957 | -0.0023 |
    | QuantizedFlat  | Cosine            | City Center Summer Wind... | 0.4025 | Red Tide Hotel             | 0.4000 | 0.0025 |
    | QuantizedFlat  | DotProduct        | City Center Summer Wind... | 0.4027 | Red Tide Hotel             | 0.4001 | 0.0025 |
    | QuantizedFlat  | Euclidean         | City Center Summer Wind... | 1.0934 | Red Tide Hotel             | 1.0957 | -0.0023 |
    

The sample demonstrates two goals:

  • Control plane: Authenticate, use the existing database, and delete and recreate the two configured vector-indexed containers.
  • Data plane: Load the same 50-document dataset, validate 1536-dimension embeddings, ingest documents by using Region as the partition key, generate a query embedding, run Cosine, DotProduct, and Euclidean queries, display ranked results and a comparison summary, and clean up the two containers.

Understand the vector distance functions

The sample's core purpose is comparing all three distance functions against the same containers and query embedding. Each function measures "closeness" differently, and that difference affects both the numeric scores and, for near ties, the ranking.

What each function measures

Function What it measures Score range
Cosine The angle between two vectors — magnitude-independent -1 to 1; higher values indicate greater similarity
DotProduct The inner product — accounts for both direction and magnitude Any real number; higher values indicate greater similarity
Euclidean Straight-line (L2) distance between vector endpoints — magnitude-sensitive 0 to 2 for unit-normalized vectors; lower values indicate greater similarity

How the example output reflects these differences

Across the sample's recorded output for the 50-hotel dataset, Cosine and DotProduct produce nearly identical scores and return the hotels in the same order. Euclidean produces scores in a different magnitude range for the same hotels.

Why Cosine and DotProduct can align closely here: This sample uses embeddings from the same model and compares the same documents across each distance function. As a result, Cosine and DotProduct often produce similar rankings. Depending on your data distribution and whether embeddings are normalized, the exact scores and ordering can still differ.

Cosine and DotProduct are similarity measures, not mathematical distance metrics. Cosine measures how closely the vectors point in the same direction. DotProduct measures that alignment while also accounting for vector magnitude. For both measures, a larger score means greater similarity.

Why Euclidean differs: Euclidean is a distance metric that measures the straight-line distance between vector endpoints. A smaller score means the vectors are closer together. The exact score ranges depend on the embeddings in your dataset.

How the distance function override works

Each query in this sample passes the function through the distanceFunction option in VectorDistance():

VectorDistance(c.embedding, @embedding, false, {'distanceFunction': 'Cosine'})

In this query, false tells Azure Cosmos DB to use the vector index instead of performing a brute-force search. The distanceFunction option tells it to calculate similarity by using cosine distance. This setting applies only to the current query—it doesn't change the distance function configured for the container. For the full signature, see the VectorDistance reference.

Tip

Match the query distance function to how your embedding model was trained. cosine is the standard default for OpenAI text embedding models, including text-embedding-3-small. Querying with a different function than the one stored in the container's vector policy still computes correctly, but might not use the vector index as efficiently.

Explore the app code

The following sections describe the main code paths in the Java sample.

nosql-create-index-java/
├── .env.example
├── output/
│   └── sample-output.txt
├── pom.xml
└── src/main/java/com/azure/cosmos/createindex/
    ├── App.java
    ├── Config.java
    ├── ControlPlane.java
    └── DataPlane.java
File Purpose
App.java Loads configuration and orchestrates control-plane and data-plane operations.
Config.java Loads and validates environment variables.
ControlPlane.java Creates, recreates, and deletes the vector-indexed containers.
DataPlane.java Loads documents, generates embeddings, ingests data, and runs vector queries.

Explore the credential and client setup

The orchestration code loads the configuration, creates one DefaultAzureCredential, and passes it to the data-plane client factories.

SampleConfig config = Config.load();
Config.validate(config);

var credential = new DefaultAzureCredentialBuilder().build();

try (var cosmosClient = DataPlane.createCosmosClient(credential, config)) {
    var openAiClient = DataPlane.createAzureOpenAIClient(credential, config);
    var database = cosmosClient.getDatabase(config.databaseName());

The createCosmosClient and createAzureOpenAIClient methods use that shared credential to create the two clients needed for document operations, embedding generation, and vector queries.

public static CosmosClient createCosmosClient(TokenCredential credential, SampleConfig config) {
    return new CosmosClientBuilder()
            .endpoint(config.cosmosEndpoint())
            .credential(credential)
            .contentResponseOnWriteEnabled(false)
            .buildClient();
}

public static OpenAIClient createAzureOpenAIClient(TokenCredential credential, SampleConfig config) {
    return new OpenAIClientBuilder()
            .endpoint(config.openAiEmbeddingEndpoint())
            .credential(credential)
            .buildClient();
}

You also pass the same credential to the control-plane method that creates the two containers.

Explore the control-plane container creation

The createContainersWithVectorIndexes method calls createContainer once for DiskANN and once for QuantizedFlat. The helper deletes the previous container, builds the selected vector index and shared vector policy, and creates the replacement.

public static void createContainersWithVectorIndexes(
        TokenCredential credential,
        String subscriptionId,
        String resourceGroup,
        String accountName,
        String location,
        String databaseName,
        String embeddingFieldName,
        String diskannContainerName,
        String quantizedflatContainerName) throws Exception {

    AzureProfile profile = new AzureProfile(AzureEnvironment.AZURE);
    AzureResourceManager azure = AzureResourceManager
            .authenticate(credential, profile)
            .withSubscription(subscriptionId);

    SqlResourcesClient sqlResourcesClient = azure.cosmosDBAccounts()
            .manager()
            .serviceClient()
            .getSqlResources();

    String embeddingPath = "/" + embeddingFieldName;

    createContainer(sqlResourcesClient, resourceGroup, accountName, location,
            databaseName, embeddingPath, diskannContainerName, VectorIndexType.DISK_ANN);
    createContainer(sqlResourcesClient, resourceGroup, accountName, location,
            databaseName, embeddingPath, quantizedflatContainerName, VectorIndexType.QUANTIZED_FLAT);
}

private static void createContainer(
        SqlResourcesClient sqlResourcesClient,
        String resourceGroup,
        String accountName,
        String location,
        String databaseName,
        String embeddingPath,
        String containerName,
        VectorIndexType indexType) throws Exception {

    String indexLabel = indexType == VectorIndexType.DISK_ANN ? "diskANN" : "quantizedFlat";

    System.out.println("\n=== Step 1: Create Container with Vector Index ===");
    System.out.printf("  Container:      %s%n", containerName);
    System.out.printf("  Index type:     %s%n", indexLabel);
    System.out.printf("  Dimensions:     %d%n", EMBEDDING_DIMENSIONS);
    System.out.println("  Distance func:  cosine (queried with all 3 metrics)");

    System.out.println("  Deleting existing container if present...");
    try {
        sqlResourcesClient.getSqlContainer(resourceGroup, accountName, databaseName, containerName);
        sqlResourcesClient.deleteSqlContainer(resourceGroup, accountName, databaseName, containerName);
        System.out.println("  Deleted existing container");
    } catch (ManagementException e) {
        if (e.getResponse() != null && e.getResponse().getStatusCode() == 404) {
            System.out.println("  Container does not exist (OK)");
        } else {
            throw e;
        }
    }

    long startMs = System.currentTimeMillis();

    VectorEmbedding vectorEmbedding = new VectorEmbedding()
            .withPath(embeddingPath)
            .withDataType(VectorDataType.FLOAT32)
            .withDimensions(EMBEDDING_DIMENSIONS)
            .withDistanceFunction(DistanceFunction.COSINE);

    VectorEmbeddingPolicy vectorEmbeddingPolicy = new VectorEmbeddingPolicy()
            .withVectorEmbeddings(Arrays.asList(vectorEmbedding));

    VectorIndex vectorIndex = new VectorIndex()
            .withPath(embeddingPath)
            .withType(indexType);

    IndexingPolicy indexingPolicy = new IndexingPolicy()
            .withIndexingMode(IndexingMode.CONSISTENT)
            .withAutomatic(true)
            .withIncludedPaths(Arrays.asList(new IncludedPath().withPath("/*")))
            .withExcludedPaths(Arrays.asList(
                    new ExcludedPath().withPath("/_etag/?"),
                    new ExcludedPath().withPath(embeddingPath + "/*")))
            .withVectorIndexes(Arrays.asList(vectorIndex));

    SqlContainerResource containerResource = new SqlContainerResource()
            .withId(containerName)
            .withPartitionKey(new ContainerPartitionKey()
                    .withPaths(Arrays.asList(REGION_PARTITION_KEY))
                    .withKind(PartitionKind.HASH))
            .withVectorEmbeddingPolicy(vectorEmbeddingPolicy)
            .withIndexingPolicy(indexingPolicy);

    SqlContainerCreateUpdateParameters containerParams = new SqlContainerCreateUpdateParameters()
            .withLocation(location)
            .withResource(containerResource)
            .withOptions(new CreateUpdateOptions());

    sqlResourcesClient.createUpdateSqlContainer(
            resourceGroup, accountName, databaseName, containerName, containerParams);

    double elapsed = (System.currentTimeMillis() - startMs) / 1000.0;
    System.out.printf("  Created in %.1fs%n", elapsed);
    System.out.println("  Vector index is IMMUTABLE \u2014 cannot be changed after creation");
}

Together, these methods complete the control plane goal: they recreate hotels_diskann and hotels_quantizedflat in the database created by azd up. The containers share /Region as the partition key path and /embedding as the 1536-dimension float32 vector path. The vector policy and index are immutable, so the sample recreates both containers on every run.

Explore the document ingestion

The ingestDocuments method associates each of the 50 hotel documents with its Region partition key and writes the same data to both containers.

public static IngestionSummary ingestDocuments(
        CosmosContainer container,
        String containerName,
        List<Map<String, Object>> documents) {

    // Group by region and print per-region counts as batch progress (matches .NET output)
    Map<String, List<Map<String, Object>>> byRegion = new java.util.TreeMap<>();
    for (Map<String, Object> doc : documents) {
        String region = String.valueOf(doc.get("Region"));
        byRegion.computeIfAbsent(region, k -> new ArrayList<>()).add(doc);
    }
    for (Map.Entry<String, List<Map<String, Object>>> entry : byRegion.entrySet()) {
        System.out.printf("  Region '%s': %d documents%n", entry.getKey(), entry.getValue().size());
    }

    List<CosmosItemOperation> operations = new ArrayList<>(documents.size());
    for (Map<String, Object> document : documents) {
        // Extract Region from document for partition key
        Object region = document.get("Region");
        if (region == null) {
            throw new IllegalStateException("Document missing Region property");
        }
        operations.add(CosmosBulkOperations.getUpsertItemOperation(
                document,
                new PartitionKeyBuilder().add(String.valueOf(region)).build()));
    }

    int upsertedDocuments = 0;
    int failedDocuments = 0;
    double requestCharge = 0.0;

    for (var response : container.executeBulkOperations(operations)) {
        var itemResponse = response.getResponse();
        if (itemResponse == null) {
            failedDocuments++;
            continue;
        }

        requestCharge += itemResponse.getRequestCharge();
        int statusCode = itemResponse.getStatusCode();
        if (statusCode >= 200 && statusCode < 300) {
            upsertedDocuments++;
        } else {
            failedDocuments++;
        }
    }

    return new IngestionSummary(containerName, documents.size(), upsertedDocuments, failedDocuments, requestCharge);
}

The method creates one upsert operation per document, submits the operations through executeBulkOperations, and reports the successful and failed operations.

Explore the vector similarity queries

The generateEmbedding method creates an embedding for supplied text. The verifyEmbeddingDimensions method calls it with probe text and confirms that Azure OpenAI returns the 1536 dimensions required by the container policy.

public static List<Float> generateEmbedding(OpenAIClient client, SampleConfig config, String text) {
    EmbeddingsOptions options = new EmbeddingsOptions(List.of(text));
    return client.getEmbeddings(config.openAiEmbeddingDeployment(), options)
            .getData()
            .get(0)
            .getEmbedding();
}

public static void verifyEmbeddingDimensions(OpenAIClient client, SampleConfig config) {
    List<Float> embedding = generateEmbedding(client, config, "dimension check");
    int actualDimensions = embedding.size();

    if (actualDimensions != config.expectedDimensions()) {
        throw new IllegalStateException(
                "Embedding dimensions do not match the container definition. Expected "
                        + config.expectedDimensions() + ", received " + actualDimensions + ".");
    }
}

The orchestration code then generates the query embedding and calls queryTopMatches for every combination of the two containers and three comparison functions.

// --- Step 3: Query with all distance functions ---
List<Float> queryEmbedding = DataPlane.generateEmbedding(openAiClient, config, config.queryText());
System.out.println("\nQuery: \"" + config.queryText() + "\"");
System.out.println("Embedding generated (" + queryEmbedding.size() + " dimensions)");
System.out.println("\nRunning searches (top " + config.topCount() + " results for each distance function)...");

List<String> distanceFunctions = List.of("Cosine", "DotProduct", "Euclidean");
List<Object[]> allResults = new java.util.ArrayList<>();
for (String containerName : Config.targetContainers(config)) {
    var container = database.getContainer(containerName);
    for (String distanceFunction : distanceFunctions) {
        QuerySummary summary = DataPlane.queryTopMatches(container, containerName, config, queryEmbedding, distanceFunction);
        String label = Config.algorithmLabel(containerName, config);
        System.out.printf("  \u2713 %s queried (%.2f RUs)%n", containerName, summary.requestCharge());
        allResults.add(new Object[]{label, distanceFunction, summary});
    }
}

The queryTopMatches method validates the embedding field name, binds the result count and embedding as parameters, scopes the request to one partition, and reads the ranked results.

public static QuerySummary queryTopMatches(
        CosmosContainer container,
        String containerName,
        SampleConfig config,
        List<Float> queryEmbedding,
        String distanceFunction) {
    String embeddingField = validateFieldName(config.embeddingFieldName());
    // Scope the vector query to one partition through SDK request options so the SQL focuses on ranking.
    String queryText = "SELECT TOP @topK c.HotelId, c.HotelName, c.Description, "
            + "VectorDistance(c." + embeddingField + ", @embedding, false, {'distanceFunction': '" + distanceFunction + "'}) AS SimilarityScore "
            + "FROM c "
            + "ORDER BY VectorDistance(c." + embeddingField + ", @embedding, false, {'distanceFunction': '" + distanceFunction + "'})";

    List<QueryResult> results = new ArrayList<>();
    double requestCharge = 0.0;

    // Query single partition (config.partitionKeyValue) for efficiency
    String partitionKeyValue = config.partitionKeyValue();
    SqlQuerySpec querySpec = new SqlQuerySpec(
            queryText,
            List.of(
                    new SqlParameter("@topK", config.topCount()),
                    new SqlParameter("@embedding", toDoubleList(queryEmbedding))));

    CosmosQueryRequestOptions options = new CosmosQueryRequestOptions();
    options.setPartitionKey(new PartitionKeyBuilder().add(partitionKeyValue).build());

    for (var page : container.queryItems(querySpec, options, Map.class).iterableByPage()) {
        requestCharge += page.getRequestCharge();
        for (Object item : page.getResults()) {
            @SuppressWarnings("unchecked")
            Map<String, Object> result = (Map<String, Object>) item;
            results.add(new QueryResult(
                    String.valueOf(result.get("HotelId")),
                    String.valueOf(result.get("HotelName")),
                    String.valueOf(result.get("Description")),
                    ((Number) result.get("SimilarityScore")).doubleValue()));
        }
    }

    return new QuerySummary(containerName, requestCharge, results);
}

These steps complete the data-plane comparison goal by running the same query embedding with Cosine, DotProduct, and Euclidean against both vector index types. For guidance on choosing a distance function for your own data, see the vector distance function overview earlier in this article.

Explore the single-partition query pattern

The sample scopes each query to the configured Region value, which defaults to Northeast, through the SDK partition key option.

Mechanism How it works Sample code
SDK partition key Sets CosmosQueryRequestOptions.setPartitionKey so the request targets the configured partition key value. options.setPartitionKey(new PartitionKeyBuilder().add(partitionKeyValue).build())
Comparison Single-partition query Cross-partition query
Query scope One configured region. All regions.
Documents in the supplied dataset 10 documents in the default Northeast region. 50 documents across all regions.
Routing Uses the SDK partition key value. Requires cross-partition fan-out.

You can alternatively add a WHERE c.Region = @region predicate to the SQL statement. This sample uses the SDK partition key option instead.

An ORDER BY VectorDistance(...) clause is required for nearest-neighbor ranking. The expression repeats the same VectorDistance(...) call used in the SELECT clause.

View and manage data in Visual Studio Code

Use the Azure Databases extension for Visual Studio Code to connect to your Azure Cosmos DB account and browse the hotels_diskann and hotels_quantizedflat containers.

  1. In Visual Studio Code, select the Azure icon in the Activity Bar.
  2. Under Resources, expand Azure Cosmos DB, and locate your account.
  3. Expand your account > HotelsCreateIndex > hotels_diskann or hotels_quantizedflat.
  4. Select a document to view its hotel fields and the embedding vector array.

Note

The sample deletes both containers at the end of each run. To retain the containers for inspection, comment out the ControlPlane.cleanupContainers call in App.java before you run the sample.

Troubleshooting

Symptom Cause Fix
IllegalArgumentException: Missing required environment variables The generated configuration is missing or isn't loaded into the current shell. Verify the .env values, load them into the current shell, and rerun the sample.
DefaultAzureCredential can't authenticate The current shell doesn't have an authenticated Azure identity. Run azd auth login, and then rerun the sample.
403 response during container creation The identity lacks control-plane access. Verify that the identity has permission to create and delete Azure Cosmos DB containers through Azure Resource Manager.
403 response during document ingestion or query The identity lacks data-plane access. Verify that the identity has Cosmos DB Built-in Data Contributor. Role assignments can take several minutes to propagate.
Azure OpenAI authorization error The identity lacks access to the embedding deployment. Verify that the identity has Cognitive Services OpenAI User on the Azure OpenAI resource.
Embedding dimension mismatch The embedding deployment returns a vector size other than 1536. Verify that AZURE_OPENAI_EMBEDDING_DEPLOYMENT identifies the expected text-embedding-3-small deployment.
Database or container configuration error The configured database doesn't exist or the two container names are identical. Verify the database and container settings, and then rerun the sample.

Clean up resources

The sample automatically deletes its vector-indexed containers at the end of each run. To remove the remaining Azure resources that the sample provisions, run azd down.

azd down