Observação
O acesso a essa página exige autorização. Você pode tentar entrar ou alterar diretórios.
O acesso a essa página exige autorização. Você pode tentar alterar os diretórios.
In this tutorial, you use Azure Managed Redis to implement memory for an AI agent. Agent memory isn't a separate Azure Managed Redis managed feature. It's an application pattern that combines Redis data structures, expiration policies, and vector search.
The pattern has two layers:
- Short-term or session memory: recent chat turns, session state, tool outputs, current plans, and checkpoints that expire or trim automatically.
- Long-term semantic memory: durable facts, preferences, summaries, and observations extracted from interactions, stored with embeddings and metadata, and retrieved by vector similarity and filters.
Azure Managed Redis supports vector search when your cache has the required RediSearch/vector search capability. For more information, see vector embeddings and vector search in Azure Managed Redis. RediSearch and RedisJSON are Redis modules, and modules must be enabled when you create the Azure Managed Redis instance. For more information, see Use Redis modules with Azure Managed Redis.
Prerequisites
- An Azure Managed Redis instance. If you need to create one, see Quickstart: Create an Azure Managed Redis instance.
- Short-term memory (Lists, Hashes, Streams) works without any modules. Long-term semantic memory requires RediSearch, which stores and searches vectors in hashes or JSON documents, as described in Azure Managed Redis vector search and Redis vector search concepts. This tutorial builds both layers together, so you need RediSearch to follow it end to end. If you store session or memory records as JSON documents instead of hashes, also enable RedisJSON. Modules can only be enabled when you create the instance and can't be added later, so decide up front which ones you'll need. See Redis modules with Azure Managed Redis.
- Python 3.10 or later with packages such as
redis,redis-entraid,azure-identity, and your embedding or agent framework SDK.
Architecture
The following conceptual architecture separates volatile session context from durable semantic memory.
User request
|
v
Agent runtime
|
+--> Short-term memory in Azure Managed Redis
| - Lists: bounded chat history
| - Hashes or JSON: session state, current plan, checkpoints
| - Streams: tool events and intermediate outputs
| - TTL and trimming on every write
|
+--> Long-term semantic memory in Azure Managed Redis
- Extracted facts, preferences, summaries
- Embedding vector + tenant/user/session metadata
- RediSearch vector index
- KNN search with metadata filters
Redis is well suited to agent state because agents need fast memory, search, and stateful workflows. Redis AI agent concepts describe short-term memory, long-term memory, planning, and tool execution as core parts of agent architecture. For more information, see Redis AI agent concepts.
Connect to Azure Managed Redis
Use Microsoft Entra ID authentication and TLS in production. The following example uses placeholders only.
import redis
from azure.identity import DefaultAzureCredential
from redis_entraid.cred_provider import create_from_default_azure_credential
credential_provider = create_from_default_azure_credential(
("https://redis.azure.com/.default",)
)
redis_client = redis.Redis(
host="<azure-managed-redis-hostname>",
port=10000,
ssl=True,
decode_responses=False,
credential_provider=credential_provider,
)
Add short-term session memory
Short-term memory should be bounded. Store recent raw turns in a list, state and checkpoints in a hash or JSON document, and tool output references in a stream. Apply TTL and trimming on each write so inactive sessions age out.
Store bounded chat history
import json
import time
SESSION_TTL_SECONDS = 60 * 60
MAX_RECENT_TURNS = 40
def history_key(tenant_id: str, session_id: str) -> str:
return f"agent:{tenant_id}:session:{session_id}:history"
def add_turn(tenant_id: str, session_id: str, role: str, content: str) -> None:
turn = {
"role": role,
"content": content,
"created_at": int(time.time()),
}
key = history_key(tenant_id, session_id)
pipe = redis_client.pipeline()
pipe.rpush(key, json.dumps(turn))
pipe.ltrim(key, -MAX_RECENT_TURNS, -1)
pipe.expire(key, SESSION_TTL_SECONDS)
pipe.execute()
def load_recent_turns(tenant_id: str, session_id: str) -> list[dict]:
return [
json.loads(item)
for item in redis_client.lrange(history_key(tenant_id, session_id), 0, -1)
]
Store session state, plans, and checkpoints
Use a hash for compact state that changes frequently. Use RedisJSON when you need nested plans or complex checkpoint structures. RedisJSON requires the RedisJSON module, which must be enabled when you create the cache. For module availability and creation-time requirements, see Redis modules with Azure Managed Redis.
def state_key(tenant_id: str, session_id: str) -> str:
return f"agent:{tenant_id}:session:{session_id}:state"
def save_checkpoint(
tenant_id: str,
session_id: str,
*,
summary: str,
current_plan: list[str],
checkpoint_name: str,
) -> None:
key = state_key(tenant_id, session_id)
redis_client.hset(
key,
mapping={
"summary": summary,
"current_plan": json.dumps(current_plan),
"checkpoint": checkpoint_name,
"updated_at": str(int(time.time())),
},
)
redis_client.expire(key, SESSION_TTL_SECONDS)
Store tool events and outputs
Store large tool outputs in application storage or a bounded Redis key, and put only references or summaries in the stream. Use MAXLEN so the stream doesn't grow without bound.
def tool_events_key(tenant_id: str, session_id: str) -> str:
return f"agent:{tenant_id}:session:{session_id}:tool-events"
def record_tool_event(
tenant_id: str,
session_id: str,
tool_name: str,
status: str,
output_ref: str,
) -> None:
key = tool_events_key(tenant_id, session_id)
redis_client.xadd(
key,
{
"tool": tool_name,
"status": status,
"output_ref": output_ref,
"created_at": str(int(time.time())),
},
maxlen=200,
approximate=True,
)
redis_client.expire(key, SESSION_TTL_SECONDS)
Add long-term semantic memory
Long-term memory stores only selected durable information. Don't embed every raw turn. Instead, promote concise facts, preferences, summaries, or observations after an extraction step. Store each memory with metadata so retrieval can filter by tenant, user, source, memory type, and retention policy.
Redis vector search uses a secondary index over hashes or JSON documents, supports FLAT and HNSW vector indexes, and supports vector search with metadata filters. For more information, see Redis vector search concepts. On Azure Managed Redis, vector search requires RediSearch/vector search capability. For more information, see Azure Managed Redis vector search.
Create a memory index
Use one embedding model, dimension, and distance metric per index. Rebuild or create a new index if you change the embedding model.
from redis.commands.search.field import NumericField, TagField, TextField, VectorField
from redis.commands.search.index_definition import IndexDefinition, IndexType
from redis.exceptions import ResponseError
INDEX_NAME = "idx:agent_memories"
MEMORY_PREFIX = "agent:memory:"
VECTOR_DIMENSIONS = 1536
def create_memory_index() -> None:
try:
redis_client.ft(INDEX_NAME).create_index(
fields=[
TagField("tenant_id"),
TagField("user_id"),
TagField("memory_type"),
TextField("text"),
TextField("source"),
NumericField("created_at"),
VectorField(
"embedding",
"HNSW",
{
"TYPE": "FLOAT32",
"DIM": VECTOR_DIMENSIONS,
"DISTANCE_METRIC": "COSINE",
},
),
],
definition=IndexDefinition(
prefix=[MEMORY_PREFIX],
index_type=IndexType.HASH,
),
)
except ResponseError as ex:
if "Index already exists" not in str(ex):
raise
Promote facts into long-term memory
The embed and extract_memories functions are application-specific placeholders. The extraction step should return only durable, consented memories. Include source attribution so a future response can explain why a memory was used.
import uuid
import numpy as np
def embed(text: str) -> list[float]:
return [0.0] * VECTOR_DIMENSIONS # Replace with your embedding model call.
def extract_memories(transcript: list[dict]) -> list[dict]:
return [] # Replace with your extraction logic (LLM call or rules).
def vector_bytes(values: list[float]) -> bytes:
return np.array(values, dtype=np.float32).tobytes()
def store_memory(
tenant_id: str,
user_id: str,
memory_type: str,
text: str,
source: str,
source_turn_ids: list[str],
) -> str:
memory_id = f"{MEMORY_PREFIX}{tenant_id}:{user_id}:{uuid.uuid4()}"
created_at = int(time.time())
redis_client.hset(
memory_id,
mapping={
"tenant_id": tenant_id,
"user_id": user_id,
"memory_type": memory_type,
"text": text,
"source": source,
"source_turn_ids": json.dumps(source_turn_ids),
"created_at": created_at,
"embedding": vector_bytes(embed(text)),
},
)
return memory_id
Retrieve memories for a request
Filter by tenant and user before running vector similarity. Return only the fields needed by the agent prompt.
from redis.commands.search.query import Query
def escape_tag(value: str) -> str:
return (
value.replace("\\", "\\\\")
.replace("{", "\\{")
.replace("}", "\\}")
.replace("|", "\\|")
)
def retrieve_memories(tenant_id: str, user_id: str, user_message: str) -> list[dict]:
tenant_filter = escape_tag(tenant_id)
user_filter = escape_tag(user_id)
query = (
Query(
f"(@tenant_id:{{{tenant_filter}}} @user_id:{{{user_filter}}})"
"=>[KNN 5 @embedding $vector AS score]"
)
.sort_by("score")
.return_fields("text", "memory_type", "source", "created_at", "score")
.dialect(2)
)
results = redis_client.ft(INDEX_NAME).search(
query,
query_params={"vector": vector_bytes(embed(user_message))},
)
return [
{
"text": doc.text,
"type": doc.memory_type,
"source": doc.source,
"score": doc.score,
}
for doc in results.docs
]
Use memories in an agent loop
At each turn, combine the session summary, recent turns, and relevant long-term memories. Keep the injected memory compact and attributed.
async def run_agent_turn(tenant_id: str, user_id: str, session_id: str, message: str):
recent_turns = load_recent_turns(tenant_id, session_id)
state = redis_client.hgetall(state_key(tenant_id, session_id))
memories = retrieve_memories(tenant_id, user_id, message)
prompt_context = {
"session_summary": state.get(b"summary", b"").decode("utf-8"),
"recent_turns": recent_turns,
"relevant_memories": memories,
}
response = await call_agent_model(
user_message=message,
context=prompt_context,
)
add_turn(tenant_id, session_id, "user", message)
add_turn(tenant_id, session_id, "assistant", response.text)
return response
Microsoft Agent Framework integration notes
Microsoft Agent Framework supports sessions, memory, context providers, and history providers. Use a session to share context across runs. For more information, see Agent Framework memory and persistence.
Use the following mapping when integrating Azure Managed Redis:
| Agent Framework concept | Azure Managed Redis pattern |
|---|---|
| Session | Store session-scoped summary, plan, checkpoint, and recent-turn keys with TTL. |
| History provider | Load and store recent messages from a bounded List or Stream. |
| Context provider | Retrieve long-term memories from RediSearch before a run, and extract/promote new memories after a run. |
| Provider state | Store per-session provider state in the AgentSession, with Redis keys referenced by tenant, user, and session ID. |
Agent Framework context providers run before and after each invocation, can add context before execution, and can process data after execution. This pattern fits Redis memory retrieval and memory promotion. For more information and custom provider examples, see Agent Framework context providers.
For built-in integrations, Agent Framework lists a Python Redis History Provider and Python Redis Provider as preview memory providers. The integrations page also lists a .NET Redis vector store connector through vector store abstractions, but doesn't list a built-in .NET Redis chat history provider. For current provider status, see Agent Framework integrations and memory AI context providers.
Important
In Agent Framework Python, configure only one history provider with load_messages=True. Use additional history providers with load_messages=False for audit or evaluation logs so the same messages aren't loaded more than once. This guidance is documented in Agent Framework context providers.
Design guidance
| Concern | Guidance |
|---|---|
| TTL and bounded windows | Use TTL on session keys and trim Lists and Streams on every write. Keep only the recent turns needed for response quality. |
| Summarization | Summarize older turns into a session summary before trimming. Use summary + recent turns + retrieved memories, not the full transcript. |
| Promotion to long-term memory | Promote only stable facts, explicit preferences, durable summaries, or high-value observations. Avoid embedding every assistant response. |
| Metadata and tenant separation | Use tenant-aware key prefixes and filterable metadata fields such as tenant_id, user_id, memory_type, source, and created_at. Always filter retrieval by tenant and user. |
| Deletion and privacy | Keep a key pattern or secondary lookup that lets you delete all session and memory keys for a user. Apply retention policies to memories that shouldn't be permanent. |
| Source attribution | Store source, source turn IDs, timestamps, confidence, and extraction reason so retrieved memories can be audited and explained. |
| Feedback loops | Extract new memories only from external user input and selected final assistant responses. Don't re-ingest context provider messages or retrieved memories as new facts. |
| Embedding consistency | Use one embedding model, vector dimension, and distance metric per index. Create a new index when changing embedding models. |
| Payload size | Return only fields needed by the agent, such as text, type, source, and score. Keep K small enough to fit the model context window. |
Delete session and user memory
Build deletion into your application. Delete short-term session keys when a session ends or when a user requests deletion. Delete long-term keys by tenant and user key prefix or by maintaining a per-user Set of memory IDs.
def delete_session(tenant_id: str, session_id: str) -> None:
redis_client.delete(
history_key(tenant_id, session_id),
state_key(tenant_id, session_id),
tool_events_key(tenant_id, session_id),
)
def delete_user_memory_ids(memory_ids: list[str]) -> None:
if memory_ids:
redis_client.delete(*memory_ids)