Note
Access to this page requires authorization. You can try signing in or changing directories.
Access to this page requires authorization. You can try changing directories.
Before you create an Azure Managed Redis instance, decide which tier fits your workload and estimate the capacity you need. This article helps you compare the four tiers, weigh memory versus compute tradeoffs, and plan for connections, reserved memory, growth, and scaling limits using only information you can gather in advance. It doesn't require a deployed instance or an Azure tenant.
For instructions on how to create or resize a cache, see Quickstart: Create an Azure Managed Redis instance and Scale an Azure Managed Redis instance.
Compare the four tiers
Azure Managed Redis offers four tiers. Three store data entirely in memory and differ by their memory-to-vCPU ratio. The fourth, Flash Optimized, stores data across memory and NVMe storage. Use the following table to narrow your choice before you look at specific SKU sizes.
| Tier | Memory-to-vCPU ratio | Good fit | Poor fit |
|---|---|---|---|
| Memory Optimized | 8:1 | Memory-intensive workloads that don't need the highest throughput, such as development, testing, and straightforward caching. | Workloads that need high throughput or heavy command processing relative to dataset size. |
| Balanced | 4:1 | Standard workloads that need an even mix of memory and compute, such as general-purpose caching and session stores. | Workloads with extreme throughput requirements, or very large datasets that don't need much compute. |
| Compute Optimized | 2:1 | Performance-intensive workloads that need maximum throughput, such as high-request-rate caching or heavy command pipelines. | Workloads that mainly need large memory capacity, because this tier has the smallest maximum size among the in-memory tiers. |
| Flash Optimized | Not applicable (memory + NVMe) | Very large datasets (hundreds of GB to multiple TB) with a clear hot/cold access pattern, where you can accept some latency variability between hot and cold data. | Write-heavy workloads, random or uniform access across the whole dataset, or workloads that need RediSearch, RedisBloom, RedisTimeSeries, or active geo-replication. |
For a full feature-by-feature comparison, including size ranges per tier, see Choosing the right tier. For pricing details across tiers and sizes, see the Azure Managed Redis pricing page. After you identify a stable region, tier, and baseline node quantity, see Plan reservation savings to evaluate a one-year or three-year compute commitment.
Memory versus compute tradeoffs
Each in-memory tier trades memory capacity for vCPUs and Redis shards. Memory Optimized allocates the fewest vCPUs and shards per GB, Compute Optimized allocates the most, and Balanced sits in between. More shards let Redis Enterprise process commands in parallel across more vCPUs, which increases throughput and can lower latency under load.
Use these general guidelines when you're unsure which dimension to scale:
- If CPU utilization or throughput is your bottleneck, move to a tier with a lower memory-to-vCPU ratio (Balanced or Compute Optimized) instead of only increasing the memory size.
- If memory capacity is your bottleneck but performance is acceptable, Memory Optimized offers the highest memory-to-vCPU ratio of the three in-memory tiers.
- Increasing the instance size within the same tier also adds vCPUs and shards, but switching to a more compute-oriented tier is typically more effective for compute-bound workloads than scaling size alone.
The exact number of vCPUs and shards per SKU can change over time as Azure Managed Redis optimizes performance, so don't treat any specific ratio as fixed. For illustrative ratios at sample sizes and more detail on shard behavior, see Sharding configuration. For guidance on measuring throughput and latency for a given tier and size, see Performance testing with Azure Managed Redis.
Flash Optimized: hot and cold data considerations
On Flash Optimized instances, approximately 20% of cache space is on RAM and the remaining 80% uses NVMe flash storage. All key names are always stored in RAM, while values move between RAM ("hot") and flash ("cold") based on access frequency, and data must move to RAM before it's read or written. Because the instance fills available RAM before using flash storage, a lightly loaded test instance can show better latency than a fully loaded production instance, since a smaller share of data fits in RAM as usage grows.
These characteristics affect tier selection: Flash Optimized doesn't support active geo-replication, non-clustered instances, RediSearch (including vector search), RedisBloom, or RedisTimeSeries — RedisJSON is the only supported module. If you're evaluating value sizes for this tier, its documented guidance recommends keeping individual values under 512 KB. For detailed workload guidance and best practices, see Best practices for the Flash Optimized tier.
Connection limits and headroom
Each SKU has a maximum number of client connections, and the limit increases with higher performance tiers and larger sizes. Estimate your expected peak connection count in advance, because this limit can force a move to a larger size or higher tier even when memory and CPU have headroom. For the current, authoritative connection-limit figures by tier and size, see Maximum number of client connections. Some of the larger SKUs listed there are in preview; see Preview SKU qualification later in this article.
Reserved memory
On every Azure Managed Redis instance, approximately 20% of available memory is reserved as a buffer for noncache operations, such as replication during failover and the active geo-replication buffer. This buffer helps improve cache performance and prevent memory starvation. Plan your usable capacity accordingly, and confirm the effective memory available for your dataset on the SKU you're considering rather than assuming the full advertised size is available for data. For more information, see Reserved memory.
This reserved-memory buffer is separate from how high availability (HA) affects the Used Memory metric. When HA is enabled, that metric reports memory used by both primary and replica shards, which can appear roughly twice the size of your actual dataset — but the SKU's memory limit already accounts for this replication, so choose a SKU based on your actual dataset size rather than doubling it to compensate for replicas. Monitor the Used Memory Percentage metric, which already normalizes for the SKU limit, and plan to scale before that percentage is consistently above 75%. For more information, see Estimate memory for capacity planning.
Dataset size and growth headroom
When you size an instance, account for more than the raw size of your values:
- Per-key overhead: Each key carries internal metadata (pointers, type information, expiration tracking), typically 50 to 100 bytes per key depending on key name length and value type. This overhead adds up for workloads with large numbers of small keys.
- Key names: Longer key names increase memory use at scale.
- Expiration tracking: Keys with a TTL consume extra memory for expiration bookkeeping.
- Growth over time: Plan initial capacity to include headroom for expected data growth, not just your current dataset size.
Growth headroom matters more for Azure Managed Redis than for some other data stores because scaling down has documented limitations: reducing memory requires your current usage to already be below the target size, and you can only scale down to a subset of compatible SKUs based on vCPU and shard configuration. Review Limitations of scaling Azure Managed Redis so your initial sizing choice doesn't depend on being able to shrink the instance later.
Modules and index overhead
Modules (RediSearch, RedisBloom, RedisTimeSeries, and RedisJSON) must be enabled when you create the instance. You can't add or update modules afterward. Plan module requirements before you provision, because changing them later means recreating the cache. RediSearch, RedisBloom, and RedisTimeSeries aren't available on Flash Optimized; only RedisJSON is. Only RediSearch and RedisJSON support active geo-replication. For the authoritative, current module support matrix across all tiers, see Scope of Redis modules.
If you use RediSearch, you must use the Enterprise clustering policy and the NoEviction eviction policy. With NoEviction, Redis returns errors on writes when memory is full instead of evicting keys, so search-indexed workloads need extra memory headroom compared to workloads that rely on eviction to manage capacity. Search indexes and other module data structures also consume memory beyond the raw key/value data they represent, so include that overhead in your capacity estimate even though the exact overhead depends on your schema and data. For clustering policy tradeoffs, see Cluster policies.
High availability and replica implications
High availability (HA) is recommended for all production scenarios and is required for coverage under the Azure Managed Redis SLA. You can select HA when you provision an instance, or enable it afterward on an instance that doesn't already have it — but you can't disable HA on an instance where it's already enabled. When HA is enabled, the instance deploys with primary and replica shards distributed across at least two nodes, and in regions that support availability zones, nodes are distributed across zones by default. An instance provisioned without HA doesn't have SLA coverage and can experience data loss and downtime during maintenance or an unexpected failure, so reserve that configuration for development and test scenarios.
If you plan to use active geo-replication, note that the Balanced B0 and B1 SKUs don't support it. For more information, see High availability and Active geo-replication.
Preview SKU qualification
Some sizes across all four tiers are in public preview: Memory Optimized, Balanced, and Compute Optimized sizes over 350 GB (480 GB and larger), and Flash Optimized sizes of 1,920 GB and 4,500 GB. Scaling geo-replicated caches is also currently in preview.
Before you plan production capacity around a preview size, confirm the size is available in your target region, and review the Supplemental Terms of Use for Microsoft Azure Previews, since preview functionality and availability can change. Validate any capacity assumptions for a preview SKU with your own testing rather than relying solely on documented ranges, and revisit your plan periodically in case preview status or size availability changes before you deploy.
Scaling limitations to factor into your plan
Because some scaling paths are restricted, choosing the right tier and size upfront reduces the risk of needing a disruptive migration later:
- You can't scale between the Flash Optimized tier and any in-memory tier (Memory Optimized, Balanced, or Compute Optimized), or the reverse. Choosing the wrong tier family means recreating the cache.
- Reducing memory size requires your current memory usage to already be below the target size, and you can only scale down to SKUs with a compatible vCPU and shard configuration.
- Scaling can change the instance's underlying IP address. If you hardcode IP addresses in firewall rules, network security groups, or client configuration instead of using the DNS name, plan for that possibility.
- If you use active geo-replication, all instances in the geo-replication group must use the same cache size, so scaling one instance typically requires coordinating scaling across the whole group.
- The clustering policy (OSS, Enterprise, or Non-clustered) is set at creation and generally can't be changed afterward, except that a Non-clustered cache can later move to a clustered policy. Non-clustered configurations are only available for caches sized 25 GB or smaller. Decide on a clustering policy before you provision, because it also affects which modules and features you can use.
For complete details, see Limitations of scaling Azure Managed Redis.
Pre-provisioning worksheet
Answer these questions before you provision an Azure Managed Redis instance. Each row points to the section or article with more detail.
| Planning question | Why it matters | Where to learn more |
|---|---|---|
| What's your current dataset size, including per-key and key-name overhead? | Determines your minimum viable SKU size. | Dataset size and growth headroom |
| How much will your dataset grow over the life of this instance? | Because scaling down is limited, undersizing is riskier than modest oversizing. | Dataset size and growth headroom |
| Is your workload read-heavy with a clear hot/cold access pattern, or write-heavy and uniformly accessed? | Determines whether Flash Optimized is a good fit or a poor fit. | Flash Optimized: hot and cold data considerations |
| Is throughput/CPU or memory capacity your primary constraint? | Determines whether to prioritize a higher-vCPU tier or a higher-memory tier. | Memory versus compute tradeoffs |
| What's your expected peak client connection count? | Connection limits scale with tier and size and can force a larger SKU independent of memory needs. | Connection limits and headroom |
| Do you need high availability and SLA coverage? | HA is required for the SLA and is recommended outside of dev/test. | High availability and replica implications |
| Do you need active geo-replication? | Not supported on Flash Optimized or on Balanced B0/B1; requires matching cache sizes across the replication group. | High availability and replica implications |
| Which modules, if any, do you need (RediSearch, RedisBloom, RedisTimeSeries, RedisJSON)? | Modules must be selected at creation and can't be added later; availability varies by tier. | Modules and index overhead |
| Which clustering policy does your client library support (OSS, Enterprise, or Non-clustered)? | Set at creation and generally fixed afterward; affects module support and non-clustered size limits. | Modules and index overhead |
| Does your capacity plan rely on a SKU size that's currently in preview? | Preview sizes can change and need extra validation before you rely on them for production. | Preview SKU qualification |
| Will the region, tier, and baseline node quantity remain stable for one or three years? | Predictable usage might qualify for reservation savings after you finalize the capacity plan. | Plan reservation savings |