Guide · Redis

Redis for AI: Vector Search, Agent Memory and MCP in Production

Redis is now three things in an AI stack at once: a vector database, a cache for model responses, and the place agents keep their working state. Each of those wants different persistence, eviction and security settings, and most production problems come from running all three on one instance configured for only one of them. This guide takes the decisions in order; it applies to Valkey too, with the differences called out.

Tyler Eastridge

By Tyler Eastridge, Head of Operations

LinkedIn · Updated

9 min read11 sections
On this page
Redis for AI in one paragraph

Redis for AI in one paragraph

Redis 8 ships the Redis Query Engine and vector sets in Redis Open Source, so vector search no longer needs Redis Stack; on Valkey the equivalent is the valkey-search module. Pick FLAT for small or exact workloads and HNSW for large ones, and size memory from the arithmetic before you trust a benchmark. Treat a semantic cache as disposable and agent memory as data: they need different eviction policies and different persistence, and ideally different instances. Give any MCP server or agent a dedicated ACL user, and treat prompts in the cache as personal data.

Pick the vector engine: Query Engine, vector sets or valkey-search

Redis Open Source 8.0 folded Redis Search, JSON, time series and the probabilistic types into the core server, so the separate RediSearch and RedisJSON modules are no longer needed. For vectors it offers two paths. The Redis Query Engine (FT.CREATE, FT.SEARCH, FT.HYBRID) indexes vectors stored in hashes or JSON documents alongside text, tag, numeric and geo fields, which is what most retrieval workloads need. Vector sets (VADD, VSIM) are a separate data type modeled on sorted sets, introduced as a preview in 8.0, with simple attribute filtering and int8 quantization on by default. Choose the Query Engine when you filter on metadata or combine keyword and vector search; choose vector sets when a key holding similar items is all you need.

On Valkey, vector search comes from the BSD-licensed valkey-search module, which supports HNSW and exact KNN, standalone and cluster mode, and from version 1.2 full-text, tag and numeric queries with aggregations. The valkey-bundle container image ships it preloaded. Command names overlap with Redis, but test your client and query syntax against the module rather than assuming parity.

FLAT or HNSW: exact answers or fast ones

A FLAT index compares the query against every vector. Recall is perfect, and Redis suggests it for small datasets, under about a million vectors, or where exact results matter more than latency. HNSW builds a layered graph and searches it approximately; Redis recommends it above about a million vectors. Its knobs are M (links per node, default 16, doubled on the bottom layer), EF_CONSTRUCTION (default 200) and EF_RUNTIME (default 10). Higher M buys recall with memory and build time; higher EF_RUNTIME buys it with latency, and can be set per query without a rebuild.

Redis 8.2 added a third type, SVS-VAMANA, a graph index designed to pair with compression. Note the caveat in the docs: Intel's LVQ and LeanVec compression are not available in Redis Open Source, which falls back to plain 8-bit scalar quantization. Measure recall on your own embeddings before you commit to any compressed index.

Size memory from the arithmetic, then measure

Raw vector memory is the number of vectors times the dimension times the bytes per element: 8 for FLOAT64, 4 for FLOAT32, 2 for FLOAT16 and BFLOAT16, 1 for INT8 and UINT8. As a worked example, one million 1,536-dimension FLOAT32 embeddings are 1,000,000 × 1,536 × 4 = 6,144,000,000 bytes, about 5.7 GiB, before anything else. Storing the same vectors as FLOAT16 halves that line, and vector sets' default int8 quantization quarters it.

Then add the graph. The vector sets documentation gives a method you can reuse as an estimate: each node holds M × 2 links on the bottom layer at 8 bytes each, plus roughly 0.33 × M × 8 bytes for the upper layers. At the default M of 16 that is 256 + about 42, near 300 bytes per vector, or about 0.3 GB per million. Add the key and the metadata fields you filter on, then the overheads every Redis node carries: replication and client output buffers, which are not counted against maxmemory, and copy-on-write headroom for the fork during an RDB save or AOF rewrite. Load a representative sample, read the index size from FT.INFO and INFO memory, and extrapolate from that rather than from a vendor benchmark.

Semantic caching for LLM responses

A semantic cache stores model responses keyed by the embedding of the prompt and returns a stored answer when a new prompt is close enough. RedisVL, Redis's Python library, implements this as SemanticCache: the match threshold is a cosine distance from 0 to 2 with a default of 0.1, entries can carry a TTL, and filterable_fields let you scope lookups by user, tenant or model. Redis also sells LangCache, a managed semantic cache on Redis Cloud behind a REST API, for teams that do not want to run it.

The threshold is the decision that matters. Too loose and the cache answers a different question with confidence; too tight and the hit rate does not pay for the embedding call. Tune it on logged real prompts, put the model and system prompt version in the filter, and never let one tenant's cached response answer another tenant. A cache that leaks across users is a data breach.

Agent memory: short-term state and long-term recall

Agents need two kinds of memory. Short-term memory is the state of one conversation or run: messages so far, tool results, the step the graph is on. Long-term memory is what survives the session: facts, preferences and past outcomes, usually stored as text plus an embedding so the agent can retrieve them by meaning. Redis's Agent Memory Server repository models exactly this split, with session memory under a configurable TTL and long-term memory promoted in the background; Redis now describes that open-source version as a research foundation for its managed memory product.

For LangGraph, the langgraph-checkpoint-redis package provides RedisSaver, AsyncRedisSaver and ShallowRedisSaver checkpointers for thread state, and RedisStore for long-term memory with optional vector search. It needs the JSON and search capabilities, built into Redis 8 and available through Redis Stack on earlier versions, and it supports TTLs on checkpoints. Decide the TTLs deliberately: a checkpoint that expires mid-conversation loses the thread, and one that never expires is unbounded growth.

The Redis MCP server, and what an agent should be allowed to do

Redis maintains an official MCP server, redis/mcp-redis, under the MIT license. It exposes tools for strings, hashes, lists, sets, sorted sets, pub/sub, streams, JSON, the query engine and server information, runs over the stdio transport, and connects with a username and password, TLS, cluster mode, or Entra ID for Azure Managed Redis. Its README recommends restricting the connecting user with Redis ACLs, for example a read-only user.

Follow that advice literally. Create a dedicated ACL user per agent, limit it to the key patterns it needs, and deny @dangerous and @admin so no prompt can produce FLUSHALL, CONFIG SET or KEYS against production. Note that Redis 8 widened the existing ACL categories: a user granted +@read can now run FT.SEARCH, and +@write now includes JSON.SET, so re-check rules written for Redis 7. Point exploratory agents at a replica or a copy, not the primary that serves users.

Where RAG retrieval latency actually goes

In a retrieval step the Redis query is often not the slowest part; the embedding call that turns the question into a vector usually is, so measure each stage separately. Inside Redis, latency moves with EF_RUNTIME, with how selective your metadata filter is (the engine switches between paging through vector batches and brute-forcing the filtered set), and with cluster fan-out, where SHARD_K_RATIO trades recall for fewer results per shard. Two traps catch most teams: FT.SEARCH returns 10 results by default unless you add LIMIT 0 <k>, and results sort by document score unless you SORTBY the distance field. Watch the slow log: a long query delays the commands queued behind it.

Persistence and high availability for AI state

A pure semantic cache can run without persistence; losing it costs model calls, not data. Agent memory and checkpoints are state, and they need the settings you would give any primary store. AOF with the default appendfsync everysec bounds loss to about one second of writes; RDB snapshots give compact backups and faster restarts; Redis's own recommendation for data you care about is both. Both fork the process, so the copy-on-write headroom from step 3 is not optional.

Replication is asynchronous, and the Redis docs are explicit that even WAIT does not stop acknowledged writes being lost in a failover, so design agents to tolerate losing their last step. Run replicas under Sentinel for availability, or Redis Cluster when the dataset or query load outgrows one node. Before production, restart a node with the full vector index loaded and time how long it takes to serve queries again; that number belongs in your recovery objective.

Eviction: right for a cache, wrong for a memory store

A cache should evict; a memory store should not. On a shared instance, allkeys-lru will happily evict an agent's long-term memory or a checkpoint to make room for cached responses, and the agent forgets without an error. The volatile-* policies only evict keys that carry a TTL and behave like noeviction when none do, which is a common way to get surprise write errors. The Redis docs suggest separate instances when one would otherwise hold both cache and persistent keys, and that is the cleanest fix: a cache instance with allkeys-lru or allkeys-lfu and no persistence, and a memory instance with noeviction, AOF and alerts well below maxmemory.

Licensing: Redis 8 or Valkey

Redis Open Source 8.0 and later is offered under your choice of RSALv2, SSPLv1 or AGPLv3, and the built-in search, JSON and other components share that license. Redis 7.4 was RSALv2 or SSPLv1; 7.2 and earlier remain BSD-3-Clause. Valkey and valkey-search are BSD-3-Clause. If you modify Redis and offer it over a network, read the AGPLv3 terms with counsel first. Teams that stay on BSD-licensed Redis 7.2 or earlier for license reasons get no built-in vector search and fewer upstream fixes; OSSeva publishes patched builds for Redis 6.2, 7.0 and 7.2.

Prompts, PII and what sits in the cache

Prompts contain whatever users type, so a semantic cache or memory store will hold personal data. Treat it as a regulated store: TLS on every connection, ACL users per application, TTLs no longer than your retention policy, and tenant or user IDs as filter fields on every cache and memory lookup. Redis 8 also enforces that a user can only create or query an index over key prefixes it has access to, so align index prefixes with your ACL key patterns. Plan for deletion requests: a user's data may sit in the cache, in long-term memory, in checkpoints, in AOF and RDB files and in backups, and an erasure process has to reach all of them.

Frequently asked questions

Is Redis a good vector database?

For workloads that fit in memory and need low-latency retrieval with metadata filters, yes. Redis 8 includes the Redis Query Engine with FLAT, HNSW and SVS-VAMANA indexes and the vector set data type in Redis Open Source. The cost is memory: every vector, its graph links and its metadata live in RAM, so size it before you commit.

Does Valkey support vector search?

Yes, through the valkey-search module, BSD-3-Clause licensed. It supports HNSW and exact KNN in standalone and cluster mode, and version 1.2 added full-text, tag and numeric queries with aggregations. The valkey-bundle image includes it.

How much memory do embeddings need in Redis?

Multiply vectors by dimensions by bytes per element (4 for FLOAT32), then add HNSW graph links, keys, metadata, replication buffers and fork headroom. One million 1,536-dimension FLOAT32 vectors are about 5.7 GiB of raw vector data alone. Measure a sample load with FT.INFO and INFO memory before sizing production.

Is there an official Redis MCP server?

Yes, redis/mcp-redis, maintained by Redis under the MIT license. Connect it with a dedicated ACL user limited to what the agent needs.

Should a semantic cache and agent memory share a Redis instance?

Preferably not. A cache wants an eviction policy such as allkeys-lru and little or no persistence; agent memory wants noeviction and AOF. On one instance, the cache's eviction policy can silently remove an agent's memory.

Redis services

Where this gets done

The work behind this page, run by the same engineers who wrote it.

More resources

Other Redis guides, comparisons and research

From the blog

Recent Redis articles

Next step

Need this done on your Redis or Valkey estate?

Named senior engineers for latency, memory, replication and failover, 24/7, with a 15-minute emergency SLA.