Repository navigation
feat: add EnvectorSemanticCache, a semantic LLM cache on an enVector index - #24
Open
minseokpark-CL wants to merge 3 commits into
Open
minseokpark-CL wants to merge 3 commits into
minseokpark-CL wants to merge 3 commits into
Conversation
…index A LangChain BaseCache that serves a prompt's cached generations when a similar prompt was answered before. Prompt embeddings are searched under encryption, and the prompt text and generations are AES-encrypted when the index has metadata_encryption on, so the cache holds nothing the server can read. This is the first feature of the package that only an encrypted index makes worthwhile. Design choices, and why: - One partition per llm_string (named after its hash) instead of a metadata filter. The filter runs client-side on the server's top-k, so rows of a busier model could push the right row out of reach and turn a hit into a miss; a partition scopes the search server-side. It also makes clear(llm_string=...) a single drop_partition, which a filter could not do without listing rows. Each row still stores its llm_string and lookup checks it, so a hash collision is a miss, never another model's answer. - similarity_threshold is a cosine similarity (the raw inner product of unit-norm vectors; 1.0 is an identical prompt), not the (1 + cos) / 2 relevance score, and not RedisSemanticCache's distance-based score_threshold whose direction is the opposite. The cache scales the vectors to unit norm itself so the threshold means what it says whatever the embedding model returns. Default 0.9, stricter than Redis's 0.2 distance, because a false hit answers a different question. - Chat prompts arrive as LangChain's serialized message list; the JSON envelopes and a shared system message would make every prompt look alike, so the text embedded is one "role: content" line per message, falling back to the raw prompt for anything that is not a text-only message list. The stored row keeps the prompt as given. - update appends. LangChain calls it only after a lookup missed, so no row within the threshold exists to replace; a concurrent duplicate is harmless because lookup returns the nearest row. - clear() drops the cache's own partitions and nothing else, so the index can be shared with other data. Tests: 56 unit tests on a new StoringFakeIndex that scores inserted vectors per partition, and an integration file run against enVector 1.6.2 with plain and encrypted metadata. Version bumped to 0.4.0 for the new public class. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Tests prove the cache works but do not show what it is for. The notebook asks a question, asks it again in other words, and asks something else, printing where each answer came from and how long it took, then shows a request with other settings landing in its own partition and a miss after clear(). It runs on a local Ollama so no API key is needed, and follows the layout of envector-msa's deployment/notebooks. The README has pointed at notebooks/ since e5b4be9 without the directory existing; it now names this file. While writing it: ChatOllama's llm_string carries neither the model name nor constructor settings, so two ChatOllama instances share one cache partition. The notebook passes the temperature per request, which does reach llm_string, and README Limitations says so. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The notebook subclassed the cache to print whether each lookup hit, and re-implemented `ollama pull` over the HTTP API. Neither is something a user needs, and both made the example look harder than using the cache is. LangChain's own RedisSemanticCache example shows a hit by timing the call alone, so this does the same: the cache is registered as is, the models are pulled with the ollama command, and each question prints its elapsed time and answer. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
한국어
배경: LangChain 의 LLM cache 란
LangChain 은 LLM(OpenAI, Anthropic, 로컬 model 등) 위에 app 을 만드는 Python 프레임워크입니다. app 은 model 에 텍스트를 보내고 답 텍스트를 받는데, LangChain 에서는 그것이
llm.invoke(...)한 번입니다. 이 글에서는 이를 model 요청이라고 부릅니다. 사용자의 질문에 답하는 것이 대표적이지만 질문만은 아닙니다. 문서 요약, 분류, 정보 추출, 번역, agent 가 다음에 쓸 tool 을 고르는 단계가 전부 같은 model 요청입니다. 이 패키지의 주 용도인 RAG 에서는 system 지시문과 검색된 문서 조각과 질문을 하나로 합친 텍스트가 model 요청이 됩니다.model 요청은 돈과 시간이 들고, app 은 같은 요청을 여러 번 보내기 쉽습니다. 그래서 LangChain 에는 cache 자리가 있습니다.
set_llm_cache(cache)로 cache 객체를 등록하면 그 뒤의 모든 model 요청이 이렇게 흐릅니다.cache.lookup(prompt, llm_string)을 부른다. 저장된 답이 있으면(hit) 그것을 돌려주고 끝. 없으면(miss)None.cache.update(prompt, llm_string, return_val)로 그 답을 저장한다.cache 는
BaseCache인터페이스를 구현한 객체이고, 메서드는lookup,update,clear()(저장된 것을 지움) 셋입니다. 주고받는 세 값은 실제로 이렇게 생겼습니다.FakeListChatModel에 "What is the capital of France?" 를 보내고InMemoryCache에 남은 것을 그대로 옮긴 것입니다.prompt: model 에 보낸 텍스트 전체를 문자열로 만든 것. chat model 은 메시지 목록을 JSON 으로 직렬화합니다. 질문 한 줄도 이런 껍데기에 싸여 옵니다.[{"lc": 1, "type": "constructor", "id": ["langchain", "schema", "messages", "HumanMessage"], "kwargs": {"content": "What is the capital of France?", "type": "human"}}]llm_string: "어떤 model 에 어떤 설정으로 보냈나"를 LangChain 이 문자열로 만든 것. model 이름과 temperature, stop 토큰 같은 요청 파라미터가 들어갑니다. gpt-4o 가 준 답을 gpt-4o-mini 요청에 돌려주면 안 되므로, cache 는prompt와llm_string이 둘 다 맞는 답만 돌려줘야 합니다.[('_type', 'fake-list-chat-model'), ('responses', ['Paris.']), ('stop', None)]return_val: model 의 답.Generation객체의 목록이고, chat model 이면 그 하위 클래스ChatGeneration(안에AIMessage를 담음)입니다. LangChain 의dumps/loads로 JSON 문자열과 상호 변환합니다.[ChatGeneration(text='Paris.', message=AIMessage(content='Paris.', ...))]LangChain 에 들어 있는 cache(
InMemoryCache,SQLiteCache)는prompt가 글자 단위로 똑같을 때만 hit 입니다. semantic cache 는prompt의 embedding 을 저장해 두고, 새prompt의 embedding 이 저장된 것과 충분히 가까우면 hit 으로 칩니다. embedding 은 뜻이 비슷한 텍스트가 가까운 vector 가 되도록 텍스트를 숫자 목록으로 바꾸는 것이고, LLM 과는 별개의 embedding model 이 합니다. 그래서 "What is the capital of France?" 와 "Which city is France's capital?" 가 같은 답을 받습니다. langchain-community 의RedisSemanticCache가 같은 종류입니다.semantic cache 는 결국 vector 검색이고, enVector 는 암호화된 vector 를 검색합니다. 그 위에 cache 를 만들면 프롬프트와 답을 서버가 읽을 수 없는 LLM cache 가 됩니다. 이 패키지의 기존
Envector는 RAG 의 문서 검색에 enVector 를 쓰고, 이 PR 은 같은 enVector 를 답 cache 에 씁니다.이 PR 이 하는 것
EnvectorSemanticCache를 추가합니다. 위의 semantic cache 를 enVector index 위에 만든 것입니다. 프롬프트 embedding 은 암호화된 vector 로 저장되고 검색되며,IndexSettings.metadata_encryption=True로 index 를 만들면 프롬프트 원문과 답(metadata 문자열로 저장됨)도 AES 로 암호화됩니다. 암호화 index 에서만 가능한 기능이라 이 패키지의 showcase 로 넣습니다.사용법은 README 의 "Semantic cache" 절에 있습니다.
Envector와 같은EnvectorConfig와 embedding model 을 받고,set_llm_cache(...)로 등록하면 그 다음 model 요청부터 적용됩니다.설계: 무엇을 정했고 왜
1. model 별 분리는 enVector 의 partition 으로 합니다.
partition 은 enVector index 안의 이름 붙은 칸입니다. 넣을 때 칸을 지정하고, 검색할 때 "이 칸 안에서만"이라고 지정할 수 있습니다. 이 cache 는
llm_string마다 칸을 하나씩 둡니다.llm_string은 수백 자가 될 수 있어 이름으로 못 쓰므로 sha256 해시 앞 16자로lc_cache_<hash>라고 이름 붙입니다.다른 방법은 모든 행을 한 칸에 넣고 각 행의 metadata 에
llm_string을 적어 두었다가 검색 결과를 받은 뒤 client 쪽에서 거르는 것(metadata filter)인데, 두 가지가 안 됩니다. 서버는 "전체에서 가장 가까운 k 개"를 돌려주므로, 다른 model 의 행이 그 k 개를 채우면 내 model 의 정답 행은 보지도 못하고 miss 가 됩니다. 그리고 한 model 의 행만 지우려면 그 행들의 ID 가 필요한데 enVector 에는 행 목록 API 가 없습니다. partition 이면 검색이 서버 쪽에서 칸 안으로 한정되고, 한 model 의 cache 를 지우는 것은drop_partition한 번입니다.보험으로 각 행의 metadata 에
llm_string원문도 저장하고, lookup 이 hit 한 행의 것을 요청값과 대조합니다. 해시가 겹치는 일이 생겨도 다른 model 의 답이 나가지 않고 miss 가 될 뿐입니다.2. "얼마나 비슷해야 hit 인가"는
similarity_threshold로, cosine similarity 기준, 기본값0.9.cosine similarity 는 두 vector 의 방향이 얼마나 같은지를 나타내는 값으로, 같으면
1.0, 무관하면0근처입니다. 길이가 1 인 vector(unit norm)끼리는 inner product 가 곧 cosine similarity 이고, enVector 의 검색 점수가 바로 그 inner product 입니다. 그래서 cache 는 저장 전과 검색 전에 embedding vector 를 직접 길이 1 로 맞춥니다. embedding model 이 길이 1 이 아닌 vector 를 돌려줘도 threshold 의 뜻이 변하지 않습니다.이름을
score_threshold로 하지 않은 이유:RedisSemanticCache의score_threshold=0.2는 "거리 0.2 이하면 hit", 즉 숫자가 작을수록 엄격합니다. 우리는 "비슷함 0.9 이상이면 hit", 숫자가 클수록 엄격합니다. 같은 이름이면 Redis 에서 옮겨오며 0.2 를 넣는 순간 거의 모든 것이 hit 이 됩니다. 이름에 방향이 들어간similarity_threshold를 씁니다. 기본값 0.9 는 Redis 의 거리 0.2(cosine 0.8 에 해당)보다 엄격합니다. 잘못된 hit 은 다른 요청의 답을 조용히 돌려주는 것이라 miss 보다 나쁘기 때문입니다. 점수가 정확히 threshold 와 같으면 hit(>=). 생성자는[-1, 1]밖의 값,bool, NaN 을ValueError로 거부합니다.3. chat 프롬프트는 메시지의 내용만 embed 합니다.
위
prompt예시처럼 chat model 의 메시지에는{"lc": 1, "type": "constructor", "id": [...], "kwargs": {...}}껍데기가 붙습니다. 이 JSON 을 그대로 embed 하면 서로 다른 내용도 껍데기 때문에 비슷해 보여서 threshold 를 쓸 수 있는 폭이 좁아집니다.embedding_text()가 메시지마다human: What is the capital of France?처럼역할: 내용한 줄로 줄여서 embed 합니다. 텍스트만 있는 메시지 목록이 아니면(이미지 block 이 섞였거나 JSON 이 아니면) 받은prompt를 그대로 embed 합니다. 행에 저장하는 것은 어느 경우든 받은prompt원문입니다.4.
update는 행을 추가만 합니다. 같은 프롬프트의 옛 행을 찾아 덮어쓰지 않습니다.LangChain 은
lookup이None을 돌려준 뒤에만update를 부릅니다.None이었다는 것은 threshold 이상으로 비슷한 행이 없었다는 뜻이므로 덮어쓸 행이 애초에 없습니다. 두 프로세스가 같은 프롬프트를 동시에 처리하면 행이 둘 생기지만, lookup 은 가장 가까운 한 행을 돌려주므로 답은 틀리지 않습니다. 덮어쓰기를 넣으면 model 요청마다 검색이 한 번 더 들어갑니다.5.
clear()는 이 cache 의 partition 만 지웁니다.clear(llm_string=...)은 그 model 의 partition 하나를 drop 하고, 없으면 아무것도 하지 않습니다.clear()는list_partitions()에서 이름이lc_cache_로 시작하는 partition 을 전부 drop 합니다. 다른 partition 은 건드리지 않으므로 RAG 문서와 같은 index 에 cache 를 두어도 됩니다.clear에 다른 keyword 를 주면 조용히 무시하지 않고TypeError입니다.6. 입력 검증과 그 밖의 동작.
update는Generation이 아닌 값(문자열,AIMessage등)과str이 아닌prompt를 서버에 닿기 전에 거부합니다. 저장된 답을loads할 수 없는 행(다른 client 가 쓴 행 등)은UserWarning을 내고 miss 로 처리합니다.loads에는 이 cache 가 쓰는 클래스(Generation,ChatGeneration,AIMessage등)만 허용 목록으로 넘깁니다.alookup/aupdate/aclear는BaseCache의 기본 구현(동기 메서드를 thread 에서 실행)을 그대로 씁니다. pyenvector 가 동기 SDK 라 그 이상 할 것이 없습니다.변경 파일
libs/envector/langchain_envector/cache.py:EnvectorSemanticCache,embedding_text,PARTITION_PREFIX.libs/envector/langchain_envector/__init__.py:EnvectorSemanticCacheexport.notebooks/01-semantic-cache.ipynb: 로컬 Ollama(embeddingall-minilm:l6-v2, chatgemma3:270m)로 cache 를 돌려 보는 notebook. envector-msa 의deployment/notebooks/형식(제목, 소개 bullet, Prerequisites, 번호 절, Clean Up, 출력 저장 안 함)을 따릅니다. 질문 → 바꿔 말한 질문(cache) → 다른 질문(model), 요청 설정이 다르면 partition 이 갈리는 것,clear()뒤 miss 를 보여 줍니다. 이 저장소에는 그동안 notebook 이 없었고 README 의notebooks/문장만 있었습니다.README.md: Features 한 줄, "Semantic cache" 예제 절(notebook 링크 포함), Limitations 두 줄,notebooks/문장을 실제 파일로.pyproject.toml: 0.3.0 → 0.4.0 (새 public class).tests/conftest.py:StoringFakeIndex. 기존 fake 는 고정된 다섯 문장만 채점하고 partition 을 무시해서, insert 된 vector 를 partition 별로 저장하고 그것을 채점하는 fake 를 추가했습니다.tests/test_cache.py: 단위 테스트 56개.tests/integration_tests/test_semantic_cache.py: 실서버 테스트 5개 × (plain metadata, encrypted metadata).검증
python -m pytest -q -m "not integration"→ 224 passed (기존 168 + 신규 56). 새로 넣은 거부 조건을 하나씩 약화시켜(경계를<=로,llm_string대조 제거, unit norm 정규화 제거, threshold 범위 검사 제거, chat 프롬프트 축약 제거) 각각 테스트가 실패하는 것을 확인했습니다.pytest -q -m integration tests/integration_tests/test_semantic_cache.py→ 10 passed (0:01:58). 항목: partition 이 아직 없을 때 첫 lookup 이 miss,update직후 같은 프롬프트와 cosine 0.95 인 프롬프트가 hit 이고 0.85 / 0.6 인 프롬프트가 miss, 행이 없는 partition 의 lookup,ChatGeneration저장 후 복원,llm_string이 다르면 서로 안 보임,clear(llm_string=)과clear()뒤 다시 사용. 서버는 빈 partition 검색에 오류가 아니라 빈 결과를 돌려주었습니다. index 전체가 비었을 때 나는 "shard list is empty" 오류의 처리는 방어용으로 남겨 두었습니다.git ls-files '*.py' | xargs black --check와ruff check libs tests scripts통과.nbclient로 처음부터 끝까지 실행했습니다. 첫 질문 9.2 s, 바꿔 말한 질문 0.09 s(같은 답), 다른 질문 7.0 s.options={"temperature": 0.8}요청은 7.1 s 가 걸리고 partition 이 둘이 됐으며,clear()뒤 바꿔 말한 질문이 다시 7.2 s 걸렸습니다. LangChain 의 RedisSemanticCache 예제처럼 hit 은 걸린 시간으로만 보입니다. 저장소의 notebook 은 출력 없이 둡니다.알아 둘 것
MetadataKey.json이 있어야 합니다. pyenvector 의 auto key setup 은metadata_encryption=True로 생성할 때만 이 파일을 만듭니다(테스트 파일 docstring 에 적음).similarity_threshold를 올리라고 README Limitations 에 적었습니다.ChatOllama(langchain-ollama 1.1.0)의llm_string은[('_type', 'chat-ollama'), ('stop', None)]뿐이라 model 이름도 생성자 설정도 들어가지 않습니다(_identifying_params가 비어 있음). 그래서 Ollama 에서는 model 이 달라도 같은 partition 을 씁니다. 요청에 넘긴 값은 들어가므로 notebook 은invoke(..., options={"temperature": 0.8})로 분리를 보여 주고, README Limitations 에 "요청 단위로 설정을 넘기거나 model 마다 index 를 따로 쓰라"고 적었습니다. 이 cache 가 아니라 langchain-ollama 쪽 동작입니다.English
Background: what an LLM cache is in LangChain
LangChain is a Python framework for building applications on top of LLMs (OpenAI, Anthropic, local models, ...). The application sends text to a model and gets text back; in LangChain that is one
llm.invoke(...). This description calls it a model request. Answering a user's question is the typical case, but not the only one: summarizing a document, classifying, extracting information, translating, and an agent choosing its next tool are all the same kind of model request. In RAG, this package's main use, the model request is the system instructions, the retrieved document chunks and the question combined into one text.Model requests cost money and time, and an application easily sends the same one many times. So LangChain has a slot for a cache. Once
set_llm_cache(cache)registers a cache object, every model request after that flows like this:cache.lookup(prompt, llm_string). If a stored answer exists (hit), return it and stop. Otherwise (miss) it returnsNone.cache.update(prompt, llm_string, return_val).The cache is an object implementing the
BaseCacheinterface, whose methods arelookup,updateandclear()(remove what was stored). The three values exchanged look like this in practice; they were copied from whatInMemoryCacheheld after sending "What is the capital of France?" to aFakeListChatModel.prompt: the whole text sent to the model, as a string. A chat model serializes its message list as JSON, so even a one-line question arrives in this wrapper.[{"lc": 1, "type": "constructor", "id": ["langchain", "schema", "messages", "HumanMessage"], "kwargs": {"content": "What is the capital of France?", "type": "human"}}]llm_string: LangChain's string for "which model, with which settings": the model name and request parameters such as temperature and stop tokens. An answer from gpt-4o must not be served to a gpt-4o-mini request, so a cache may only return an answer whosepromptandllm_stringboth match.[('_type', 'fake-list-chat-model'), ('responses', ['Paris.']), ('stop', None)]return_val: the model's answer, a list ofGenerationobjects; for a chat model, its subclassChatGeneration(holding anAIMessage). LangChain'sdumps/loadsconvert them to and from a JSON string.[ChatGeneration(text='Paris.', message=AIMessage(content='Paris.', ...))]The caches that ship with LangChain (
InMemoryCache,SQLiteCache) hit only when thepromptis identical character for character. A semantic cache stores the prompt's embedding and counts a newpromptas a hit when its embedding is close enough to a stored one. An embedding turns text into a list of numbers such that texts with similar meaning become nearby vectors; a separate embedding model, not the LLM, computes it. That is how "What is the capital of France?" and "Which city is France's capital?" get the same answer. langchain-community'sRedisSemanticCacheis one of these.A semantic cache is a vector search in the end, and enVector searches encrypted vectors. A cache built on it is an LLM cache whose server can read neither the prompts nor the answers. This package's existing
Envectoruses enVector for RAG document retrieval; this PR uses the same enVector for the answer cache.What this PR does
Adds
EnvectorSemanticCache, the semantic cache above built on an enVector index. Prompt embeddings are stored and searched as encrypted vectors, and when the index is created withIndexSettings.metadata_encryption=True, the prompt text and the answer (stored as the metadata string) are AES-encrypted as well. It is only possible on an encrypted index, so it goes in as this package's showcase.Usage is in the README's "Semantic cache" section. It takes the same
EnvectorConfigand embedding model asEnvector; register it withset_llm_cache(...)and every model request from then on uses it.Design: what was decided and why
1. Models are kept apart with enVector partitions.
A partition is a named compartment inside an enVector index: rows are inserted into one, and a search can be limited to one. This cache keeps one partition per
llm_string. Anllm_stringcan be hundreds of characters and cannot be a name, so the partition is namedlc_cache_<hash>after the first 16 hex characters of its sha256.The alternative is to put every row in one place, write the
llm_stringinto each row's metadata, and filter the search results on the client (a metadata filter). That fails in two ways. The server returns "the k nearest rows of everything", so when rows of another model fill those k, the right row for this model is never seen and the lookup misses. And removing one model's rows needs their IDs, while enVector has no API to list rows. With a partition the search is limited server-side, and clearing one model's cache is a singledrop_partition.As a safeguard, each row also stores the full
llm_stringin its metadata and lookup compares it with the requested one. Should two hashes ever collide, the result is a miss, not another model's answer.2. "How similar counts as a hit" is
similarity_threshold, a cosine similarity, default0.9.Cosine similarity measures how much two vectors point the same way:
1.0for identical, around0for unrelated. For vectors of length 1 (unit norm) the inner product is the cosine similarity, and enVector's search score is exactly that inner product. So the cache scales embedding vectors to length 1 itself, before storing and before searching; the threshold keeps its meaning even when the embedding model returns vectors of another length.Why not the name
score_threshold:RedisSemanticCache'sscore_threshold=0.2means "hit when the distance is at most 0.2", stricter as the number gets smaller. Ours means "hit when the similarity is at least 0.9", stricter as the number gets larger. With the same name, someone moving from Redis who passes 0.2 would get a hit on almost everything.similarity_thresholdcarries the direction in its name. The default 0.9 is stricter than Redis's 0.2 distance (a cosine of 0.8) because a wrong hit silently serves the answer to a different request, which is worse than a miss. A score exactly equal to the threshold is a hit (>=). The constructor rejects values outside[-1, 1],booland NaN withValueError.3. Chat prompts are embedded by their message contents only.
As the
promptexample above shows, each chat message comes wrapped in{"lc": 1, "type": "constructor", "id": [...], "kwargs": {...}}. Embedding that JSON makes different contents look alike because of the wrapper, which narrows the usable threshold range.embedding_text()reduces each message to onerole: contentline, such ashuman: What is the capital of France?, and embeds that. Anything that is not a text-only message list (an image block mixed in, or not JSON at all) is embedded as the receivedprompt. What the row stores is the receivedpromptin every case.4.
updateonly adds a row. It does not look for an older row for the same prompt and overwrite it.LangChain calls
updateonly afterlookupreturnedNone.Nonemeans no row within the threshold existed, so there is nothing to overwrite. Two processes handling the same prompt at once leave two rows, but lookup returns the single nearest row, so the answer is not wrong. Overwriting would add one more search to every model request.5.
clear()removes only this cache's partitions.clear(llm_string=...)drops that model's one partition and does nothing when it is missing.clear()drops every partition inlist_partitions()whose name starts withlc_cache_. Other partitions are left alone, so the cache can live in the same index as RAG documents. Any other keyword toclearraisesTypeErrorinstead of being silently ignored.6. Input checks and the rest.
updaterefuses values that are notGeneration(a string, anAIMessage, ...) and apromptthat is not astrbefore the server is contacted. A row whose stored answer cannot beloads-ed (written by another client, say) is reported with aUserWarningand counts as a miss.loadsis given an allow list of only the classes this cache writes (Generation,ChatGeneration,AIMessage, ...).alookup/aupdate/aclearkeepBaseCache's default implementation (the sync method run in a thread); pyenvector is a synchronous SDK, so there is nothing more to do.Files changed
libs/envector/langchain_envector/cache.py:EnvectorSemanticCache,embedding_text,PARTITION_PREFIX.libs/envector/langchain_envector/__init__.py: exportsEnvectorSemanticCache.notebooks/01-semantic-cache.ipynb: a notebook that runs the cache with a local Ollama (embeddingall-minilm:l6-v2, chatgemma3:270m). It follows the layout of envector-msa'sdeployment/notebooks/(title, intro bullets, Prerequisites, numbered sections, Clean Up, no saved outputs). It shows a question → the same question reworded (cache) → another question (model), a request with other settings landing in its own partition, and a miss afterclear(). This repository had no notebook before, only the README sentence pointing atnotebooks/.README.md: one Features line, a "Semantic cache" example section (with a link to the notebook), two Limitations lines, and thenotebooks/sentence now naming a real file.pyproject.toml: 0.3.0 → 0.4.0 (new public class).tests/conftest.py:StoringFakeIndex. The existing fake scores five fixed sentences and ignores partitions, so a fake that stores inserted vectors per partition and scores those was added.tests/test_cache.py: 56 unit tests.tests/integration_tests/test_semantic_cache.py: 5 live-server tests × (plain metadata, encrypted metadata).Verification
python -m pytest -q -m "not integration"→ 224 passed (168 existing + 56 new). Each new rejection was weakened in turn (boundary changed to<=,llm_stringcomparison removed, unit-norm scaling removed, threshold range check removed, chat prompt reduction removed) and the tests failed each time.pytest -q -m integration tests/integration_tests/test_semantic_cache.py→ 10 passed (0:01:58). Covered: a first lookup missing while no partition exists yet, a hit right afterupdatefor the same prompt and for a prompt at cosine 0.95 with misses at 0.85 / 0.6, a lookup in a partition with no rows, aChatGenerationstored and restored, differentllm_stringvalues not seeing each other, and reuse afterclear(llm_string=)andclear(). The server answered the empty-partition search with an empty result, not an error. The handling of the "shard list is empty" error raised for a fully empty index is kept as a guard.git ls-files '*.py' | xargs black --checkandruff check libs tests scriptspass.nbclientagainst the same stack and the local Ollama. The first question took 9.2 s, the reworded one 0.09 s (same text), the other question 7.0 s. The request withoptions={"temperature": 0.8}took 7.1 s and a second partition appeared, and afterclear()the reworded question took 7.2 s again. As in LangChain's RedisSemanticCache example, a hit shows only as elapsed time. The notebook in the repository is kept without outputs.Worth knowing
MetadataKey.json; pyenvector's auto key setup writes that file only when generating withmetadata_encryption=True(noted in the test file's docstring).similarity_thresholdwhen the instructions are long and fixed, as in RAG.ChatOllama(langchain-ollama 1.1.0) builds thellm_string[('_type', 'chat-ollama'), ('stop', None)]and nothing more: neither the model name nor constructor settings are in it (_identifying_paramsis empty). With Ollama, different models therefore share one partition. Values passed with the request do go in, so the notebook shows the separation withinvoke(..., options={"temperature": 0.8}), and README Limitations says to pass settings per request or use one index per model. This is langchain-ollama's behaviour, not the cache's.🤖 Generated with Claude Code