Tutorial: Build agent memory with Azure Managed Redis

In this tutorial, you use Azure Managed Redis to implement memory for an AI agent. Agent memory isn't a separate Azure Managed Redis managed feature. It's an application pattern that combines Redis data structures, expiration policies, and vector search.

The pattern has two layers:

  • Short-term or session memory: recent chat turns, session state, tool outputs, current plans, and checkpoints that expire or trim automatically.
  • Long-term semantic memory: durable facts, preferences, summaries, and observations extracted from interactions, stored with embeddings and metadata, and retrieved by vector similarity and filters.

Azure Managed Redis supports vector search when your cache has the required RediSearch/vector search capability. For more information, see vector embeddings and vector search in Azure Managed Redis. RediSearch and RedisJSON are Redis modules, and modules must be enabled when you create the Azure Managed Redis instance. For more information, see Use Redis modules with Azure Managed Redis.

Prerequisites

  • An Azure Managed Redis instance. If you need to create one, see Quickstart: Create an Azure Managed Redis instance.
  • Short-term memory (Lists, Hashes, Streams) works without any modules. Long-term semantic memory requires RediSearch, which stores and searches vectors in hashes or JSON documents, as described in Azure Managed Redis vector search and Redis vector search concepts. This tutorial builds both layers together, so you need RediSearch to follow it end to end. If you store session or memory records as JSON documents instead of hashes, also enable RedisJSON. Modules can only be enabled when you create the instance and can't be added later, so decide up front which ones you'll need. See Redis modules with Azure Managed Redis.
  • Python 3.10 or later with packages such as redis, redis-entraid, azure-identity, and your embedding or agent framework SDK.

Architecture

The following conceptual architecture separates volatile session context from durable semantic memory.

User request
    |
    v
Agent runtime
    |
    +--> Short-term memory in Azure Managed Redis
    |       - Lists: bounded chat history
    |       - Hashes or JSON: session state, current plan, checkpoints
    |       - Streams: tool events and intermediate outputs
    |       - TTL and trimming on every write
    |
    +--> Long-term semantic memory in Azure Managed Redis
            - Extracted facts, preferences, summaries
            - Embedding vector + tenant/user/session metadata
            - RediSearch vector index
            - KNN search with metadata filters

Redis is well suited to agent state because agents need fast memory, search, and stateful workflows. Redis AI agent concepts describe short-term memory, long-term memory, planning, and tool execution as core parts of agent architecture. For more information, see Redis AI agent concepts.

Connect to Azure Managed Redis

Use Microsoft Entra ID authentication and TLS in production. The following example uses placeholders only.

import redis
from azure.identity import DefaultAzureCredential
from redis_entraid.cred_provider import create_from_default_azure_credential

credential_provider = create_from_default_azure_credential(
    ("https://redis.azure.com/.default",)
)

redis_client = redis.Redis(
    host="<azure-managed-redis-hostname>",
    port=10000,
    ssl=True,
    decode_responses=False,
    credential_provider=credential_provider,
)

Add short-term session memory

Short-term memory should be bounded. Store recent raw turns in a list, state and checkpoints in a hash or JSON document, and tool output references in a stream. Apply TTL and trimming on each write so inactive sessions age out.

Store bounded chat history

import json
import time

SESSION_TTL_SECONDS = 60 * 60
MAX_RECENT_TURNS = 40

def history_key(tenant_id: str, session_id: str) -> str:
    return f"agent:{tenant_id}:session:{session_id}:history"

def add_turn(tenant_id: str, session_id: str, role: str, content: str) -> None:
    turn = {
        "role": role,
        "content": content,
        "created_at": int(time.time()),
    }
    key = history_key(tenant_id, session_id)

    pipe = redis_client.pipeline()
    pipe.rpush(key, json.dumps(turn))
    pipe.ltrim(key, -MAX_RECENT_TURNS, -1)
    pipe.expire(key, SESSION_TTL_SECONDS)
    pipe.execute()

def load_recent_turns(tenant_id: str, session_id: str) -> list[dict]:
    return [
        json.loads(item)
        for item in redis_client.lrange(history_key(tenant_id, session_id), 0, -1)
    ]

Store session state, plans, and checkpoints

Use a hash for compact state that changes frequently. Use RedisJSON when you need nested plans or complex checkpoint structures. RedisJSON requires the RedisJSON module, which must be enabled when you create the cache. For module availability and creation-time requirements, see Redis modules with Azure Managed Redis.

def state_key(tenant_id: str, session_id: str) -> str:
    return f"agent:{tenant_id}:session:{session_id}:state"

def save_checkpoint(
    tenant_id: str,
    session_id: str,
    *,
    summary: str,
    current_plan: list[str],
    checkpoint_name: str,
) -> None:
    key = state_key(tenant_id, session_id)
    redis_client.hset(
        key,
        mapping={
            "summary": summary,
            "current_plan": json.dumps(current_plan),
            "checkpoint": checkpoint_name,
            "updated_at": str(int(time.time())),
        },
    )
    redis_client.expire(key, SESSION_TTL_SECONDS)

Store tool events and outputs

Store large tool outputs in application storage or a bounded Redis key, and put only references or summaries in the stream. Use MAXLEN so the stream doesn't grow without bound.

def tool_events_key(tenant_id: str, session_id: str) -> str:
    return f"agent:{tenant_id}:session:{session_id}:tool-events"

def record_tool_event(
    tenant_id: str,
    session_id: str,
    tool_name: str,
    status: str,
    output_ref: str,
) -> None:
    key = tool_events_key(tenant_id, session_id)
    redis_client.xadd(
        key,
        {
            "tool": tool_name,
            "status": status,
            "output_ref": output_ref,
            "created_at": str(int(time.time())),
        },
        maxlen=200,
        approximate=True,
    )
    redis_client.expire(key, SESSION_TTL_SECONDS)

Add long-term semantic memory

Long-term memory stores only selected durable information. Don't embed every raw turn. Instead, promote concise facts, preferences, summaries, or observations after an extraction step. Store each memory with metadata so retrieval can filter by tenant, user, source, memory type, and retention policy.

Redis vector search uses a secondary index over hashes or JSON documents, supports FLAT and HNSW vector indexes, and supports vector search with metadata filters. For more information, see Redis vector search concepts. On Azure Managed Redis, vector search requires RediSearch/vector search capability. For more information, see Azure Managed Redis vector search.

Create a memory index

Use one embedding model, dimension, and distance metric per index. Rebuild or create a new index if you change the embedding model.

from redis.commands.search.field import NumericField, TagField, TextField, VectorField
from redis.commands.search.index_definition import IndexDefinition, IndexType
from redis.exceptions import ResponseError

INDEX_NAME = "idx:agent_memories"
MEMORY_PREFIX = "agent:memory:"
VECTOR_DIMENSIONS = 1536

def create_memory_index() -> None:
    try:
        redis_client.ft(INDEX_NAME).create_index(
            fields=[
                TagField("tenant_id"),
                TagField("user_id"),
                TagField("memory_type"),
                TextField("text"),
                TextField("source"),
                NumericField("created_at"),
                VectorField(
                    "embedding",
                    "HNSW",
                    {
                        "TYPE": "FLOAT32",
                        "DIM": VECTOR_DIMENSIONS,
                        "DISTANCE_METRIC": "COSINE",
                    },
                ),
            ],
            definition=IndexDefinition(
                prefix=[MEMORY_PREFIX],
                index_type=IndexType.HASH,
            ),
        )
    except ResponseError as ex:
        if "Index already exists" not in str(ex):
            raise

Promote facts into long-term memory

The embed and extract_memories functions are application-specific placeholders. The extraction step should return only durable, consented memories. Include source attribution so a future response can explain why a memory was used.

import uuid
import numpy as np

def embed(text: str) -> list[float]:
    return [0.0] * VECTOR_DIMENSIONS  # Replace with your embedding model call.

def extract_memories(transcript: list[dict]) -> list[dict]:
    return []  # Replace with your extraction logic (LLM call or rules).

def vector_bytes(values: list[float]) -> bytes:
    return np.array(values, dtype=np.float32).tobytes()

def store_memory(
    tenant_id: str,
    user_id: str,
    memory_type: str,
    text: str,
    source: str,
    source_turn_ids: list[str],
) -> str:
    memory_id = f"{MEMORY_PREFIX}{tenant_id}:{user_id}:{uuid.uuid4()}"
    created_at = int(time.time())

    redis_client.hset(
        memory_id,
        mapping={
            "tenant_id": tenant_id,
            "user_id": user_id,
            "memory_type": memory_type,
            "text": text,
            "source": source,
            "source_turn_ids": json.dumps(source_turn_ids),
            "created_at": created_at,
            "embedding": vector_bytes(embed(text)),
        },
    )
    return memory_id

Retrieve memories for a request

Filter by tenant and user before running vector similarity. Return only the fields needed by the agent prompt.

from redis.commands.search.query import Query

def escape_tag(value: str) -> str:
    return (
        value.replace("\\", "\\\\")
        .replace("{", "\\{")
        .replace("}", "\\}")
        .replace("|", "\\|")
    )

def retrieve_memories(tenant_id: str, user_id: str, user_message: str) -> list[dict]:
    tenant_filter = escape_tag(tenant_id)
    user_filter = escape_tag(user_id)

    query = (
        Query(
            f"(@tenant_id:{{{tenant_filter}}} @user_id:{{{user_filter}}})"
            "=>[KNN 5 @embedding $vector AS score]"
        )
        .sort_by("score")
        .return_fields("text", "memory_type", "source", "created_at", "score")
        .dialect(2)
    )

    results = redis_client.ft(INDEX_NAME).search(
        query,
        query_params={"vector": vector_bytes(embed(user_message))},
    )

    return [
        {
            "text": doc.text,
            "type": doc.memory_type,
            "source": doc.source,
            "score": doc.score,
        }
        for doc in results.docs
    ]

Use memories in an agent loop

At each turn, combine the session summary, recent turns, and relevant long-term memories. Keep the injected memory compact and attributed.

async def run_agent_turn(tenant_id: str, user_id: str, session_id: str, message: str):
    recent_turns = load_recent_turns(tenant_id, session_id)
    state = redis_client.hgetall(state_key(tenant_id, session_id))
    memories = retrieve_memories(tenant_id, user_id, message)

    prompt_context = {
        "session_summary": state.get(b"summary", b"").decode("utf-8"),
        "recent_turns": recent_turns,
        "relevant_memories": memories,
    }

    response = await call_agent_model(
        user_message=message,
        context=prompt_context,
    )

    add_turn(tenant_id, session_id, "user", message)
    add_turn(tenant_id, session_id, "assistant", response.text)
    return response

Microsoft Agent Framework integration notes

Microsoft Agent Framework supports sessions, memory, context providers, and history providers. Use a session to share context across runs. For more information, see Agent Framework memory and persistence.

Use the following mapping when integrating Azure Managed Redis:

Agent Framework concept Azure Managed Redis pattern
Session Store session-scoped summary, plan, checkpoint, and recent-turn keys with TTL.
History provider Load and store recent messages from a bounded List or Stream.
Context provider Retrieve long-term memories from RediSearch before a run, and extract/promote new memories after a run.
Provider state Store per-session provider state in the AgentSession, with Redis keys referenced by tenant, user, and session ID.

Agent Framework context providers run before and after each invocation, can add context before execution, and can process data after execution. This pattern fits Redis memory retrieval and memory promotion. For more information and custom provider examples, see Agent Framework context providers.

For built-in integrations, Agent Framework lists a Python Redis History Provider and Python Redis Provider as preview memory providers. The integrations page also lists a .NET Redis vector store connector through vector store abstractions, but doesn't list a built-in .NET Redis chat history provider. For current provider status, see Agent Framework integrations and memory AI context providers.

Important

In Agent Framework Python, configure only one history provider with load_messages=True. Use additional history providers with load_messages=False for audit or evaluation logs so the same messages aren't loaded more than once. This guidance is documented in Agent Framework context providers.

Design guidance

Concern Guidance
TTL and bounded windows Use TTL on session keys and trim Lists and Streams on every write. Keep only the recent turns needed for response quality.
Summarization Summarize older turns into a session summary before trimming. Use summary + recent turns + retrieved memories, not the full transcript.
Promotion to long-term memory Promote only stable facts, explicit preferences, durable summaries, or high-value observations. Avoid embedding every assistant response.
Metadata and tenant separation Use tenant-aware key prefixes and filterable metadata fields such as tenant_id, user_id, memory_type, source, and created_at. Always filter retrieval by tenant and user.
Deletion and privacy Keep a key pattern or secondary lookup that lets you delete all session and memory keys for a user. Apply retention policies to memories that shouldn't be permanent.
Source attribution Store source, source turn IDs, timestamps, confidence, and extraction reason so retrieved memories can be audited and explained.
Feedback loops Extract new memories only from external user input and selected final assistant responses. Don't re-ingest context provider messages or retrieved memories as new facts.
Embedding consistency Use one embedding model, vector dimension, and distance metric per index. Create a new index when changing embedding models.
Payload size Return only fields needed by the agent, such as text, type, source, and score. Keep K small enough to fit the model context window.

Delete session and user memory

Build deletion into your application. Delete short-term session keys when a session ends or when a user requests deletion. Delete long-term keys by tenant and user key prefix or by maintaining a per-user Set of memory IDs.

def delete_session(tenant_id: str, session_id: str) -> None:
    redis_client.delete(
        history_key(tenant_id, session_id),
        state_key(tenant_id, session_id),
        tool_events_key(tenant_id, session_id),
    )

def delete_user_memory_ids(memory_ids: list[str]) -> None:
    if memory_ids:
        redis_client.delete(*memory_ids)