Skip to content

feat: add EnvectorSemanticCache, a semantic LLM cache on an enVector index - #24

Open
minseokpark-CL wants to merge 3 commits into
mainfrom
feat/semantic-cache
Open

minseokpark-CL wants to merge 3 commits into
mainfrom
feat/semantic-cache

Conversation

@minseokpark-CL

@minseokpark-CL minseokpark-CL commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

한국어

배경: LangChain 의 LLM cache 란

LangChain 은 LLM(OpenAI, Anthropic, 로컬 model 등) 위에 app 을 만드는 Python 프레임워크입니다. app 은 model 에 텍스트를 보내고 답 텍스트를 받는데, LangChain 에서는 그것이 llm.invoke(...) 한 번입니다. 이 글에서는 이를 model 요청이라고 부릅니다. 사용자의 질문에 답하는 것이 대표적이지만 질문만은 아닙니다. 문서 요약, 분류, 정보 추출, 번역, agent 가 다음에 쓸 tool 을 고르는 단계가 전부 같은 model 요청입니다. 이 패키지의 주 용도인 RAG 에서는 system 지시문과 검색된 문서 조각과 질문을 하나로 합친 텍스트가 model 요청이 됩니다.

model 요청은 돈과 시간이 들고, app 은 같은 요청을 여러 번 보내기 쉽습니다. 그래서 LangChain 에는 cache 자리가 있습니다. set_llm_cache(cache) 로 cache 객체를 등록하면 그 뒤의 모든 model 요청이 이렇게 흐릅니다.

  1. 보낼 텍스트를 만든다.
  2. cache.lookup(prompt, llm_string) 을 부른다. 저장된 답이 있으면(hit) 그것을 돌려주고 끝. 없으면(miss) None.
  3. miss 면 model 에 요청을 보내 답을 받는다.
  4. cache.update(prompt, llm_string, return_val) 로 그 답을 저장한다.

cache 는 BaseCache 인터페이스를 구현한 객체이고, 메서드는 lookup, update, clear()(저장된 것을 지움) 셋입니다. 주고받는 세 값은 실제로 이렇게 생겼습니다. FakeListChatModel 에 "What is the capital of France?" 를 보내고 InMemoryCache 에 남은 것을 그대로 옮긴 것입니다.

  • prompt: model 에 보낸 텍스트 전체를 문자열로 만든 것. chat model 은 메시지 목록을 JSON 으로 직렬화합니다. 질문 한 줄도 이런 껍데기에 싸여 옵니다.
    [{"lc": 1, "type": "constructor", "id": ["langchain", "schema", "messages", "HumanMessage"], "kwargs": {"content": "What is the capital of France?", "type": "human"}}]
  • llm_string: "어떤 model 에 어떤 설정으로 보냈나"를 LangChain 이 문자열로 만든 것. model 이름과 temperature, stop 토큰 같은 요청 파라미터가 들어갑니다. gpt-4o 가 준 답을 gpt-4o-mini 요청에 돌려주면 안 되므로, cache 는 prompt 와 llm_string 이 둘 다 맞는 답만 돌려줘야 합니다.
    [('_type', 'fake-list-chat-model'), ('responses', ['Paris.']), ('stop', None)]
  • return_val: model 의 답. Generation 객체의 목록이고, chat model 이면 그 하위 클래스 ChatGeneration(안에 AIMessage 를 담음)입니다. LangChain 의 dumps / loads 로 JSON 문자열과 상호 변환합니다.
    [ChatGeneration(text='Paris.', message=AIMessage(content='Paris.', ...))]

LangChain 에 들어 있는 cache(InMemoryCache, SQLiteCache)는 prompt 가 글자 단위로 똑같을 때만 hit 입니다. semantic cache 는 prompt 의 embedding 을 저장해 두고, 새 prompt 의 embedding 이 저장된 것과 충분히 가까우면 hit 으로 칩니다. embedding 은 뜻이 비슷한 텍스트가 가까운 vector 가 되도록 텍스트를 숫자 목록으로 바꾸는 것이고, LLM 과는 별개의 embedding model 이 합니다. 그래서 "What is the capital of France?" 와 "Which city is France's capital?" 가 같은 답을 받습니다. langchain-community 의 RedisSemanticCache 가 같은 종류입니다.

semantic cache 는 결국 vector 검색이고, enVector 는 암호화된 vector 를 검색합니다. 그 위에 cache 를 만들면 프롬프트와 답을 서버가 읽을 수 없는 LLM cache 가 됩니다. 이 패키지의 기존 Envector 는 RAG 의 문서 검색에 enVector 를 쓰고, 이 PR 은 같은 enVector 를 답 cache 에 씁니다.

이 PR 이 하는 것

EnvectorSemanticCache 를 추가합니다. 위의 semantic cache 를 enVector index 위에 만든 것입니다. 프롬프트 embedding 은 암호화된 vector 로 저장되고 검색되며, IndexSettings.metadata_encryption=True 로 index 를 만들면 프롬프트 원문과 답(metadata 문자열로 저장됨)도 AES 로 암호화됩니다. 암호화 index 에서만 가능한 기능이라 이 패키지의 showcase 로 넣습니다.

사용법은 README 의 "Semantic cache" 절에 있습니다. Envector 와 같은 EnvectorConfig 와 embedding model 을 받고, set_llm_cache(...) 로 등록하면 그 다음 model 요청부터 적용됩니다.

설계: 무엇을 정했고 왜

1. model 별 분리는 enVector 의 partition 으로 합니다.
partition 은 enVector index 안의 이름 붙은 칸입니다. 넣을 때 칸을 지정하고, 검색할 때 "이 칸 안에서만"이라고 지정할 수 있습니다. 이 cache 는 llm_string 마다 칸을 하나씩 둡니다. llm_string 은 수백 자가 될 수 있어 이름으로 못 쓰므로 sha256 해시 앞 16자로 lc_cache_<hash> 라고 이름 붙입니다.
다른 방법은 모든 행을 한 칸에 넣고 각 행의 metadata 에 llm_string 을 적어 두었다가 검색 결과를 받은 뒤 client 쪽에서 거르는 것(metadata filter)인데, 두 가지가 안 됩니다. 서버는 "전체에서 가장 가까운 k 개"를 돌려주므로, 다른 model 의 행이 그 k 개를 채우면 내 model 의 정답 행은 보지도 못하고 miss 가 됩니다. 그리고 한 model 의 행만 지우려면 그 행들의 ID 가 필요한데 enVector 에는 행 목록 API 가 없습니다. partition 이면 검색이 서버 쪽에서 칸 안으로 한정되고, 한 model 의 cache 를 지우는 것은 drop_partition 한 번입니다.
보험으로 각 행의 metadata 에 llm_string 원문도 저장하고, lookup 이 hit 한 행의 것을 요청값과 대조합니다. 해시가 겹치는 일이 생겨도 다른 model 의 답이 나가지 않고 miss 가 될 뿐입니다.

2. "얼마나 비슷해야 hit 인가"는 similarity_threshold 로, cosine similarity 기준, 기본값 0.9.
cosine similarity 는 두 vector 의 방향이 얼마나 같은지를 나타내는 값으로, 같으면 1.0, 무관하면 0 근처입니다. 길이가 1 인 vector(unit norm)끼리는 inner product 가 곧 cosine similarity 이고, enVector 의 검색 점수가 바로 그 inner product 입니다. 그래서 cache 는 저장 전과 검색 전에 embedding vector 를 직접 길이 1 로 맞춥니다. embedding model 이 길이 1 이 아닌 vector 를 돌려줘도 threshold 의 뜻이 변하지 않습니다.
이름을 score_threshold 로 하지 않은 이유: RedisSemanticCache 의 score_threshold=0.2 는 "거리 0.2 이하면 hit", 즉 숫자가 작을수록 엄격합니다. 우리는 "비슷함 0.9 이상이면 hit", 숫자가 클수록 엄격합니다. 같은 이름이면 Redis 에서 옮겨오며 0.2 를 넣는 순간 거의 모든 것이 hit 이 됩니다. 이름에 방향이 들어간 similarity_threshold 를 씁니다. 기본값 0.9 는 Redis 의 거리 0.2(cosine 0.8 에 해당)보다 엄격합니다. 잘못된 hit 은 다른 요청의 답을 조용히 돌려주는 것이라 miss 보다 나쁘기 때문입니다. 점수가 정확히 threshold 와 같으면 hit(>=). 생성자는 [-1, 1] 밖의 값, bool, NaN 을 ValueError 로 거부합니다.

3. chat 프롬프트는 메시지의 내용만 embed 합니다.
위 prompt 예시처럼 chat model 의 메시지에는 {"lc": 1, "type": "constructor", "id": [...], "kwargs": {...}} 껍데기가 붙습니다. 이 JSON 을 그대로 embed 하면 서로 다른 내용도 껍데기 때문에 비슷해 보여서 threshold 를 쓸 수 있는 폭이 좁아집니다. embedding_text() 가 메시지마다 human: What is the capital of France? 처럼 역할: 내용 한 줄로 줄여서 embed 합니다. 텍스트만 있는 메시지 목록이 아니면(이미지 block 이 섞였거나 JSON 이 아니면) 받은 prompt 를 그대로 embed 합니다. 행에 저장하는 것은 어느 경우든 받은 prompt 원문입니다.

4. update 는 행을 추가만 합니다. 같은 프롬프트의 옛 행을 찾아 덮어쓰지 않습니다.
LangChain 은 lookup 이 None 을 돌려준 뒤에만 update 를 부릅니다. None 이었다는 것은 threshold 이상으로 비슷한 행이 없었다는 뜻이므로 덮어쓸 행이 애초에 없습니다. 두 프로세스가 같은 프롬프트를 동시에 처리하면 행이 둘 생기지만, lookup 은 가장 가까운 한 행을 돌려주므로 답은 틀리지 않습니다. 덮어쓰기를 넣으면 model 요청마다 검색이 한 번 더 들어갑니다.

5. clear() 는 이 cache 의 partition 만 지웁니다.
clear(llm_string=...) 은 그 model 의 partition 하나를 drop 하고, 없으면 아무것도 하지 않습니다. clear() 는 list_partitions() 에서 이름이 lc_cache_ 로 시작하는 partition 을 전부 drop 합니다. 다른 partition 은 건드리지 않으므로 RAG 문서와 같은 index 에 cache 를 두어도 됩니다. clear 에 다른 keyword 를 주면 조용히 무시하지 않고 TypeError 입니다.

6. 입력 검증과 그 밖의 동작.
update 는 Generation 이 아닌 값(문자열, AIMessage 등)과 str 이 아닌 prompt 를 서버에 닿기 전에 거부합니다. 저장된 답을 loads 할 수 없는 행(다른 client 가 쓴 행 등)은 UserWarning 을 내고 miss 로 처리합니다. loads 에는 이 cache 가 쓰는 클래스(Generation, ChatGeneration, AIMessage 등)만 허용 목록으로 넘깁니다. alookup / aupdate / aclear 는 BaseCache 의 기본 구현(동기 메서드를 thread 에서 실행)을 그대로 씁니다. pyenvector 가 동기 SDK 라 그 이상 할 것이 없습니다.

변경 파일

  • libs/envector/langchain_envector/cache.py: EnvectorSemanticCache, embedding_text, PARTITION_PREFIX.
  • libs/envector/langchain_envector/__init__.py: EnvectorSemanticCache export.
  • notebooks/01-semantic-cache.ipynb: 로컬 Ollama(embedding all-minilm:l6-v2, chat gemma3:270m)로 cache 를 돌려 보는 notebook. envector-msa 의 deployment/notebooks/ 형식(제목, 소개 bullet, Prerequisites, 번호 절, Clean Up, 출력 저장 안 함)을 따릅니다. 질문 → 바꿔 말한 질문(cache) → 다른 질문(model), 요청 설정이 다르면 partition 이 갈리는 것, clear() 뒤 miss 를 보여 줍니다. 이 저장소에는 그동안 notebook 이 없었고 README 의 notebooks/ 문장만 있었습니다.
  • README.md: Features 한 줄, "Semantic cache" 예제 절(notebook 링크 포함), Limitations 두 줄, notebooks/ 문장을 실제 파일로.
  • pyproject.toml: 0.3.0 → 0.4.0 (새 public class).
  • tests/conftest.py: StoringFakeIndex. 기존 fake 는 고정된 다섯 문장만 채점하고 partition 을 무시해서, insert 된 vector 를 partition 별로 저장하고 그것을 채점하는 fake 를 추가했습니다.
  • tests/test_cache.py: 단위 테스트 56개.
  • tests/integration_tests/test_semantic_cache.py: 실서버 테스트 5개 × (plain metadata, encrypted metadata).

검증

  • 단위 테스트: python -m pytest -q -m "not integration" → 224 passed (기존 168 + 신규 56). 새로 넣은 거부 조건을 하나씩 약화시켜(경계를 <= 로, llm_string 대조 제거, unit norm 정규화 제거, threshold 범위 검사 제거, chat 프롬프트 축약 제거) 각각 테스트가 실패하는 것을 확인했습니다.
  • 통합 테스트: enVector 1.6.2 docker compose 스택(이 PR 을 위해 새로 띄운 것)에서 pytest -q -m integration tests/integration_tests/test_semantic_cache.py → 10 passed (0:01:58). 항목: partition 이 아직 없을 때 첫 lookup 이 miss, update 직후 같은 프롬프트와 cosine 0.95 인 프롬프트가 hit 이고 0.85 / 0.6 인 프롬프트가 miss, 행이 없는 partition 의 lookup, ChatGeneration 저장 후 복원, llm_string 이 다르면 서로 안 보임, clear(llm_string=) 과 clear() 뒤 다시 사용. 서버는 빈 partition 검색에 오류가 아니라 빈 결과를 돌려주었습니다. index 전체가 비었을 때 나는 "shard list is empty" 오류의 처리는 방어용으로 남겨 두었습니다.
  • 포맷: git ls-files '*.py' | xargs black --check 와 ruff check libs tests scripts 통과.
  • notebook: 같은 스택과 로컬 Ollama 에서 nbclient 로 처음부터 끝까지 실행했습니다. 첫 질문 9.2 s, 바꿔 말한 질문 0.09 s(같은 답), 다른 질문 7.0 s. options={"temperature": 0.8} 요청은 7.1 s 가 걸리고 partition 이 둘이 됐으며, clear() 뒤 바꿔 말한 질문이 다시 7.2 s 걸렸습니다. LangChain 의 RedisSemanticCache 예제처럼 hit 은 걸린 시간으로만 보입니다. 저장소의 notebook 은 출력 없이 둡니다.

알아 둘 것

  • encrypted metadata 로 통합 테스트를 돌리려면 key 묶음에 MetadataKey.json 이 있어야 합니다. pyenvector 의 auto key setup 은 metadata_encryption=True 로 생성할 때만 이 파일을 만듭니다(테스트 파일 docstring 에 적음).
  • 긴 system 지시문을 공유하는 chat 프롬프트들은 내용이 달라도 서로 점수가 높게 나옵니다. RAG 처럼 지시문이 길고 고정된 경우 similarity_threshold 를 올리라고 README Limitations 에 적었습니다.
  • ChatOllama(langchain-ollama 1.1.0)의 llm_string 은 [('_type', 'chat-ollama'), ('stop', None)] 뿐이라 model 이름도 생성자 설정도 들어가지 않습니다(_identifying_params 가 비어 있음). 그래서 Ollama 에서는 model 이 달라도 같은 partition 을 씁니다. 요청에 넘긴 값은 들어가므로 notebook 은 invoke(..., options={"temperature": 0.8}) 로 분리를 보여 주고, README Limitations 에 "요청 단위로 설정을 넘기거나 model 마다 index 를 따로 쓰라"고 적었습니다. 이 cache 가 아니라 langchain-ollama 쪽 동작입니다.

English

Background: what an LLM cache is in LangChain

LangChain is a Python framework for building applications on top of LLMs (OpenAI, Anthropic, local models, ...). The application sends text to a model and gets text back; in LangChain that is one llm.invoke(...). This description calls it a model request. Answering a user's question is the typical case, but not the only one: summarizing a document, classifying, extracting information, translating, and an agent choosing its next tool are all the same kind of model request. In RAG, this package's main use, the model request is the system instructions, the retrieved document chunks and the question combined into one text.

Model requests cost money and time, and an application easily sends the same one many times. So LangChain has a slot for a cache. Once set_llm_cache(cache) registers a cache object, every model request after that flows like this:

  1. Build the text to send.
  2. Call cache.lookup(prompt, llm_string). If a stored answer exists (hit), return it and stop. Otherwise (miss) it returns None.
  3. On a miss, send the request to the model and get the answer.
  4. Store the answer with cache.update(prompt, llm_string, return_val).

The cache is an object implementing the BaseCache interface, whose methods are lookup, update and clear() (remove what was stored). The three values exchanged look like this in practice; they were copied from what InMemoryCache held after sending "What is the capital of France?" to a FakeListChatModel.

  • prompt: the whole text sent to the model, as a string. A chat model serializes its message list as JSON, so even a one-line question arrives in this wrapper.
    [{"lc": 1, "type": "constructor", "id": ["langchain", "schema", "messages", "HumanMessage"], "kwargs": {"content": "What is the capital of France?", "type": "human"}}]
  • llm_string: LangChain's string for "which model, with which settings": the model name and request parameters such as temperature and stop tokens. An answer from gpt-4o must not be served to a gpt-4o-mini request, so a cache may only return an answer whose prompt and llm_string both match.
    [('_type', 'fake-list-chat-model'), ('responses', ['Paris.']), ('stop', None)]
  • return_val: the model's answer, a list of Generation objects; for a chat model, its subclass ChatGeneration (holding an AIMessage). LangChain's dumps / loads convert them to and from a JSON string.
    [ChatGeneration(text='Paris.', message=AIMessage(content='Paris.', ...))]

The caches that ship with LangChain (InMemoryCache, SQLiteCache) hit only when the prompt is identical character for character. A semantic cache stores the prompt's embedding and counts a new prompt as a hit when its embedding is close enough to a stored one. An embedding turns text into a list of numbers such that texts with similar meaning become nearby vectors; a separate embedding model, not the LLM, computes it. That is how "What is the capital of France?" and "Which city is France's capital?" get the same answer. langchain-community's RedisSemanticCache is one of these.

A semantic cache is a vector search in the end, and enVector searches encrypted vectors. A cache built on it is an LLM cache whose server can read neither the prompts nor the answers. This package's existing Envector uses enVector for RAG document retrieval; this PR uses the same enVector for the answer cache.

What this PR does

Adds EnvectorSemanticCache, the semantic cache above built on an enVector index. Prompt embeddings are stored and searched as encrypted vectors, and when the index is created with IndexSettings.metadata_encryption=True, the prompt text and the answer (stored as the metadata string) are AES-encrypted as well. It is only possible on an encrypted index, so it goes in as this package's showcase.

Usage is in the README's "Semantic cache" section. It takes the same EnvectorConfig and embedding model as Envector; register it with set_llm_cache(...) and every model request from then on uses it.

Design: what was decided and why

1. Models are kept apart with enVector partitions.
A partition is a named compartment inside an enVector index: rows are inserted into one, and a search can be limited to one. This cache keeps one partition per llm_string. An llm_string can be hundreds of characters and cannot be a name, so the partition is named lc_cache_<hash> after the first 16 hex characters of its sha256.
The alternative is to put every row in one place, write the llm_string into each row's metadata, and filter the search results on the client (a metadata filter). That fails in two ways. The server returns "the k nearest rows of everything", so when rows of another model fill those k, the right row for this model is never seen and the lookup misses. And removing one model's rows needs their IDs, while enVector has no API to list rows. With a partition the search is limited server-side, and clearing one model's cache is a single drop_partition.
As a safeguard, each row also stores the full llm_string in its metadata and lookup compares it with the requested one. Should two hashes ever collide, the result is a miss, not another model's answer.

2. "How similar counts as a hit" is similarity_threshold, a cosine similarity, default 0.9.
Cosine similarity measures how much two vectors point the same way: 1.0 for identical, around 0 for unrelated. For vectors of length 1 (unit norm) the inner product is the cosine similarity, and enVector's search score is exactly that inner product. So the cache scales embedding vectors to length 1 itself, before storing and before searching; the threshold keeps its meaning even when the embedding model returns vectors of another length.
Why not the name score_threshold: RedisSemanticCache's score_threshold=0.2 means "hit when the distance is at most 0.2", stricter as the number gets smaller. Ours means "hit when the similarity is at least 0.9", stricter as the number gets larger. With the same name, someone moving from Redis who passes 0.2 would get a hit on almost everything. similarity_threshold carries the direction in its name. The default 0.9 is stricter than Redis's 0.2 distance (a cosine of 0.8) because a wrong hit silently serves the answer to a different request, which is worse than a miss. A score exactly equal to the threshold is a hit (>=). The constructor rejects values outside [-1, 1], bool and NaN with ValueError.

3. Chat prompts are embedded by their message contents only.
As the prompt example above shows, each chat message comes wrapped in {"lc": 1, "type": "constructor", "id": [...], "kwargs": {...}}. Embedding that JSON makes different contents look alike because of the wrapper, which narrows the usable threshold range. embedding_text() reduces each message to one role: content line, such as human: What is the capital of France?, and embeds that. Anything that is not a text-only message list (an image block mixed in, or not JSON at all) is embedded as the received prompt. What the row stores is the received prompt in every case.

4. update only adds a row. It does not look for an older row for the same prompt and overwrite it.
LangChain calls update only after lookup returned None. None means no row within the threshold existed, so there is nothing to overwrite. Two processes handling the same prompt at once leave two rows, but lookup returns the single nearest row, so the answer is not wrong. Overwriting would add one more search to every model request.

5. clear() removes only this cache's partitions.
clear(llm_string=...) drops that model's one partition and does nothing when it is missing. clear() drops every partition in list_partitions() whose name starts with lc_cache_. Other partitions are left alone, so the cache can live in the same index as RAG documents. Any other keyword to clear raises TypeError instead of being silently ignored.

6. Input checks and the rest.
update refuses values that are not Generation (a string, an AIMessage, ...) and a prompt that is not a str before the server is contacted. A row whose stored answer cannot be loads-ed (written by another client, say) is reported with a UserWarning and counts as a miss. loads is given an allow list of only the classes this cache writes (Generation, ChatGeneration, AIMessage, ...). alookup / aupdate / aclear keep BaseCache's default implementation (the sync method run in a thread); pyenvector is a synchronous SDK, so there is nothing more to do.

Files changed

  • libs/envector/langchain_envector/cache.py: EnvectorSemanticCache, embedding_text, PARTITION_PREFIX.
  • libs/envector/langchain_envector/__init__.py: exports EnvectorSemanticCache.
  • notebooks/01-semantic-cache.ipynb: a notebook that runs the cache with a local Ollama (embedding all-minilm:l6-v2, chat gemma3:270m). It follows the layout of envector-msa's deployment/notebooks/ (title, intro bullets, Prerequisites, numbered sections, Clean Up, no saved outputs). It shows a question → the same question reworded (cache) → another question (model), a request with other settings landing in its own partition, and a miss after clear(). This repository had no notebook before, only the README sentence pointing at notebooks/.
  • README.md: one Features line, a "Semantic cache" example section (with a link to the notebook), two Limitations lines, and the notebooks/ sentence now naming a real file.
  • pyproject.toml: 0.3.0 → 0.4.0 (new public class).
  • tests/conftest.py: StoringFakeIndex. The existing fake scores five fixed sentences and ignores partitions, so a fake that stores inserted vectors per partition and scores those was added.
  • tests/test_cache.py: 56 unit tests.
  • tests/integration_tests/test_semantic_cache.py: 5 live-server tests × (plain metadata, encrypted metadata).

Verification

  • Unit tests: python -m pytest -q -m "not integration" → 224 passed (168 existing + 56 new). Each new rejection was weakened in turn (boundary changed to <=, llm_string comparison removed, unit-norm scaling removed, threshold range check removed, chat prompt reduction removed) and the tests failed each time.
  • Integration tests: on an enVector 1.6.2 docker compose stack started for this PR, pytest -q -m integration tests/integration_tests/test_semantic_cache.py → 10 passed (0:01:58). Covered: a first lookup missing while no partition exists yet, a hit right after update for the same prompt and for a prompt at cosine 0.95 with misses at 0.85 / 0.6, a lookup in a partition with no rows, a ChatGeneration stored and restored, different llm_string values not seeing each other, and reuse after clear(llm_string=) and clear(). The server answered the empty-partition search with an empty result, not an error. The handling of the "shard list is empty" error raised for a fully empty index is kept as a guard.
  • Formatting: git ls-files '*.py' | xargs black --check and ruff check libs tests scripts pass.
  • Notebook: executed top to bottom with nbclient against the same stack and the local Ollama. The first question took 9.2 s, the reworded one 0.09 s (same text), the other question 7.0 s. The request with options={"temperature": 0.8} took 7.1 s and a second partition appeared, and after clear() the reworded question took 7.2 s again. As in LangChain's RedisSemanticCache example, a hit shows only as elapsed time. The notebook in the repository is kept without outputs.

Worth knowing

  • The encrypted-metadata integration run needs a key bundle that contains MetadataKey.json; pyenvector's auto key setup writes that file only when generating with metadata_encryption=True (noted in the test file's docstring).
  • Chat prompts that share long system instructions score alike even when their contents differ. README Limitations says to raise similarity_threshold when the instructions are long and fixed, as in RAG.
  • ChatOllama (langchain-ollama 1.1.0) builds the llm_string [('_type', 'chat-ollama'), ('stop', None)] and nothing more: neither the model name nor constructor settings are in it (_identifying_params is empty). With Ollama, different models therefore share one partition. Values passed with the request do go in, so the notebook shows the separation with invoke(..., options={"temperature": 0.8}), and README Limitations says to pass settings per request or use one index per model. This is langchain-ollama's behaviour, not the cache's.

🤖 Generated with Claude Code

minseokpark-CL and others added 3 commits October 6, 2026 17:03
…index

A LangChain BaseCache that serves a prompt's cached generations when a
similar prompt was answered before. Prompt embeddings are searched under
encryption, and the prompt text and generations are AES-encrypted when the
index has metadata_encryption on, so the cache holds nothing the server can
read. This is the first feature of the package that only an encrypted index
makes worthwhile.

Design choices, and why:

- One partition per llm_string (named after its hash) instead of a metadata
  filter. The filter runs client-side on the server's top-k, so rows of a
  busier model could push the right row out of reach and turn a hit into a
  miss; a partition scopes the search server-side. It also makes
  clear(llm_string=...) a single drop_partition, which a filter could not do
  without listing rows. Each row still stores its llm_string and lookup
  checks it, so a hash collision is a miss, never another model's answer.
- similarity_threshold is a cosine similarity (the raw inner product of
  unit-norm vectors; 1.0 is an identical prompt), not the (1 + cos) / 2
  relevance score, and not RedisSemanticCache's distance-based
  score_threshold whose direction is the opposite. The cache scales the
  vectors to unit norm itself so the threshold means what it says whatever
  the embedding model returns. Default 0.9, stricter than Redis's 0.2
  distance, because a false hit answers a different question.
- Chat prompts arrive as LangChain's serialized message list; the JSON
  envelopes and a shared system message would make every prompt look alike,
  so the text embedded is one "role: content" line per message, falling back
  to the raw prompt for anything that is not a text-only message list. The
  stored row keeps the prompt as given.
- update appends. LangChain calls it only after a lookup missed, so no row
  within the threshold exists to replace; a concurrent duplicate is harmless
  because lookup returns the nearest row.
- clear() drops the cache's own partitions and nothing else, so the index
  can be shared with other data.

Tests: 56 unit tests on a new StoringFakeIndex that scores inserted vectors
per partition, and an integration file run against enVector 1.6.2 with plain
and encrypted metadata. Version bumped to 0.4.0 for the new public class.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Tests prove the cache works but do not show what it is for. The notebook
asks a question, asks it again in other words, and asks something else,
printing where each answer came from and how long it took, then shows a
request with other settings landing in its own partition and a miss after
clear(). It runs on a local Ollama so no API key is needed, and follows the
layout of envector-msa's deployment/notebooks.

The README has pointed at notebooks/ since e5b4be9 without the directory
existing; it now names this file.

While writing it: ChatOllama's llm_string carries neither the model name nor
constructor settings, so two ChatOllama instances share one cache partition.
The notebook passes the temperature per request, which does reach
llm_string, and README Limitations says so.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The notebook subclassed the cache to print whether each lookup hit, and
re-implemented `ollama pull` over the HTTP API. Neither is something a
user needs, and both made the example look harder than using the cache
is. LangChain's own RedisSemanticCache example shows a hit by timing the
call alone, so this does the same: the cache is registered as is, the
models are pulled with the ollama command, and each question prints its
elapsed time and answer.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant