Local memory with replayable evidence paths — not only top-k chunks. Structured graph retrieval for agents, with a local palace CLI, MCP closed loop, and public-benchmark KPIs.
Memory Path Engine models memory as typed nodes, edges, weights, and MemoryPath objects so a system can retrieve, traverse, and explain how it reached an answer. Product M1 adds an on-disk palace (SQLite) and the mpe CLI; Layer B fixtures remain the architecture proof surface.
Bundled markdown packs are ingested into a graph (MemoryNode / MemoryEdge). Retrievers return a MemoryPath. The CLI persists that graph under .mpe/ and prints answer + hops.
Memory Palace v1 adds a parallel domain (memory_engine.memory) mapped through palace_to_store. See docs/architecture.md and docs/ROADMAP.md.
examples/*_pack ──▶ mpe ingest ──▶ .mpe/store.sqlite (typed graph)
│
┌──────────────┼──────────────┐
▼ ▼ ▼
BaselineTopK other modes WeightedGraph
(flat answers) in `retrieve` (path + scores)
│
▼
mpe search/path → answer + hop list
Most RAG systems still look like this:
- Split documents into chunks.
- Embed chunks.
- Return
top-kmatches. - Ask the LLM to improvise the reasoning.
This project asks:
Can we retrieve a memory path instead of only retrieving similar chunks?
Three product bets:
structure: typed nodes and edges, not only a flat vector indexweight: importance, risk, novelty, reinforcement / forgettingpath: replayable evidence chains as the default search output
- compare multiple retrieval modes in one codebase
- inspect replayable evidence paths instead of only final answers
- test graph-aware retrieval on contract-like and operational documents
- run repository-owned structured benchmarks instead of toy snippets
Maintainers: configure the GitHub link-card image using docs/social-preview.md (docs/assets/open-graph-cover.png).
Install from PyPI (Product M8) or editable clone (see docs/install.md for pipx / uv / Docker):
pip install memory-path-engine
# pip install 'memory-path-engine[embed]' # optional dense backends
# or: python -m pip install --no-build-isolation -e .
mpe doctorPublish / first tag: docs/publish.md · Changelog: CHANGELOG.md.
mpe init --mode hybrid
mpe ingest examples/runbook_pack/runbooks --pack example_runbook_pack
mpe search "What if rollback does not recover the API?" --mode hybrid
mpe path "What if rollback does not recover the API?"
mpe status
mpe backupPalace files live in ./.mpe/ (or $MPE_PALACE). Search always prints an answer plus hop citations.
5-minute agent closed loop (MCP + hooks): see docs/getting-started.md.
mpe hooks install
mpe mcp # stdio MCP server for Cursor / Claude
# docker build -t mpe-mcp . && docker run -i --rm -v "$PWD/.mpe:/data/palace" -e MPE_PALACE=/data/palace mpe-mcpAcceptance checklist: docs/ACCEPTANCE.md.
Current progress vs vision stages: docs/progress.md.
mpe bench longmemeval --label tiny --granularity session
mpe bench longmemeval --label tiny --granularity turnChecked-in tiny reports live under benchmarks/external/longmemeval/baselines/.
Latest public recall (session, LongMemEval-S cleaned):
| Slice | Embedding | Mode | R@5 | R@10 | NDCG@10 |
|---|---|---|---|---|---|
| 50q | ngram | hybrid | 0.980 | 1.000 | 0.948 |
| 50q | ngram | lexical_baseline | 0.980 | 1.000 | 0.948 |
| 500q (full) | ngram | hybrid | 0.950 | 0.974 | 0.863 |
| 500q (full) | ngram | lexical_baseline | 0.950 | 0.974 | 0.863 |
| 500q (full) | fastembed | hybrid | 0.962 | 0.980 | 0.875 |
| 50q turn | ngram | hybrid | 0.840 | 0.920 | 0.714 |
| 50q turn | ngram | lexical_baseline | 0.840 | 0.920 | 0.714 |
Artifacts: benchmarks/external/longmemeval/baselines/longmemeval_kpi_{medium30,medium50,medium50_fastembed,medium50_turn,full_ngram,full_fastembed}.*.
M7 note: hybrid public ranking now reuses BM25/blend seed scores so NDCG no longer collapses after graph expansion. With fastembed, hybrid beats lexical on full R@5/R@10/NDCG. Turn mid-slice (Product M11) is harder than session (R@5 0.84 vs 0.98) as expected with finer units.
HotpotQA mid-slice (64q, evidence-hit, Product M10):
| Embedding | Mode | evidence_hit | comparison hit |
|---|---|---|---|
| ngram | lexical_baseline | 0.406 | 0.562 |
| ngram | hybrid | 0.844 | 0.938 |
| fastembed | embedding_baseline | 0.531 | 0.812 |
| fastembed | hybrid | 0.859 | 0.938 |
Artifacts: benchmarks/external/hotpotqa/baselines/hotpotqa_kpi_medium64*.md. Dense (fastembed) especially helps comparison/paraphrase-like questions; set MPE_EMBEDDING_CACHE_DIR to cache vectors on disk.
Optional dense embeddings (Product M5+): default remains dependency-free ngram. Install pip install 'memory-path-engine[embed]' then:
mpe bench longmemeval --label tiny --embedding fastembed
# or: export MPE_EMBEDDING=fastembed
# optional: export MPE_EMBEDDING_MAX_CHARS=4000Dense backends batch-encode when texts are short, and head+tail-truncate long session memories to fit model context.
Reproduce full LongMemEval-S (public KPI):
python scripts/download_longmemeval.py
mpe bench longmemeval \
--dataset benchmarks/external/longmemeval/data/longmemeval_s_cleaned.json \
--label full --granularity session --limit 0| Slice | Granularity | How to run | Notes |
|---|---|---|---|
| tiny (2q, checked in) | session / turn | mpe bench longmemeval --label tiny [--granularity turn] |
CI smoke + committed baselines |
| medium (50q) | session | nightly / local with --limit 50 |
positioning |
| full (LongMemEval-S) | session / turn | download + --label full --limit 0 |
headline public KPI |
See benchmarks/external/longmemeval/README.md and docs/ROADMAP.md.
Run the test suite:
python -m unittest discover -s tests -vRun the runbook demo:
python -m memory_engine.demo --scenario runbookTerminal-style capture of real stdout (refresh with python scripts/generate_runbook_demo_terminal_svg.py; latency_ms may differ run to run):
Run the research-notes demo:
python -m memory_engine.demo --scenario researchRun the HotpotQA tiny benchmark sanity check:
python scripts/run_hotpotqa_benchmark.pyDownload HotpotQA (CMU URL, with HuggingFace Hub fallback) and run the medium64 KPI slice:
python scripts/download_hotpotqa.py
python scripts/run_hotpotqa_benchmark.py \
--dataset benchmarks/external/hotpotqa/data/hotpot_dev_distractor_v1.json \
--limit 64 --top-k 10 \
--modes lexical_baseline,embedding_baseline,weighted_graph,hybrid,activation_spreading_v1 \
--summary-output benchmarks/external/hotpotqa/data/hotpotqa-medium64-summary.jsonOptional dense embedding disk cache (Product M10):
export MPE_EMBEDDING_CACHE_DIR="$PWD/.mpe/embed-cache"
export MPE_EMBEDDING=fastembedCommitted mid-slice KPI: benchmarks/external/hotpotqa/baselines/hotpotqa_kpi_medium64*.md.
Run the LongMemEval tiny benchmark sanity check:
python scripts/run_longmemeval_benchmark.pyPrint compact v1 palace metadata (spaces, routes, memory kinds) per case:
python scripts/run_longmemeval_benchmark.py --v1-recall-summaryGenerate a fixed-format Layer B report (path/route/space/lifecycle/activation snapshot):
python scripts/generate_layer_b_report.py --output "benchmarks/structured_memory/layer_b_report.json" --markdown-output "benchmarks/structured_memory/layer_b_report.md"Generate a fixed-format ablation matrix and latency summary (no-structure / no-weight / no-path-expansion):
python scripts/generate_ablation_report.py --output "benchmarks/structured_memory/ablation_report.json" --markdown-output "benchmarks/structured_memory/ablation_report.md"Generate a Layer A positioning report (external metrics only; keeps architecture claims in Layer B):
python scripts/generate_layer_a_report.py --slice-profile tiny --markdown-output "benchmarks/external/layer_a_report.md"Run Layer C transferable stand-in benchmarks:
python scripts/run_layer_c_benchmark.py --markdown-output "benchmarks/layer_c_minimal/layer_c_report.md"Download the official HotpotQA dev distractor file for local benchmark runs:
python scripts/download_hotpotqa.pyDownload the cleaned LongMemEval-S file for local benchmark runs:
python scripts/download_longmemeval.pypython -m memory_engine.demo prints a small banner, the query, then path-aware output: a BEST ANSWER line built from the winning walk, and a REPLAY PATH with one line per hop (node id, score, via=<edge type>) plus short scoring reasons on the following lines. With --scenario contract, a BASELINE block (flat top-k answers) appears above the path-aware section for the same query.
Representative runbook excerpt (answer line shortened; latency and hop scores can vary slightly between runs):
========================================================================
Memory Path Engine | demo
scenario: runbook
========================================================================
-------------------------------- QUERY ---------------------------------
What should we do if rollback does not recover the API after a
deployment incident?
----------------- PATH-AWARE weighted graph retrieval -----------------
BEST ANSWER
… stitched runbook units … [latency_ms=…]
REPLAY PATH
1. 01_api_incident_runbook:5 | score=0.500 | via=seed
seed hit semantic=0.501
2. 01_api_incident_runbook:4 | score=0.299 | via=next_unit
expanded at hop 1 total=0.299 exception=0.450 contradiction=0.000
========================================================================
The runbook demo loads incident and recovery procedures, then asks a multi-step operational question:
What should we do if rollback does not recover the API after a deployment incident?
The output includes:
- a BEST ANSWER line composed from the graph walk
- a REPLAY PATH with per-step scores,
viaedge types, and short reasons
For a representative stdout excerpt, see What you will see (under Quick start).
The contract demo runs the same query through a baseline retriever and the weighted graph retriever. Stdout shows flat top-k answers first, then the path-aware best answer and replay steps, so you can compare shapes of evidence without relying on a single aggregate metric.
| Retriever | What it emphasizes | Useful for |
|---|---|---|
| lexical baseline | keyword overlap | simple lookups and sanity checks |
| embedding baseline | semantic similarity | paraphrases and fuzzy matches |
| structure-only traversal | graph connectivity | linked evidence exploration |
| weighted graph retrieval | structure plus importance weighting | multi-hop retrieval with replayable evidence |
| activation spreading v1 | explicit propagation with decay | graph diffusion experiments |
The core is meant to stay domain-agnostic. The current examples use both contract-like documents and runbooks because together they stress:
- hierarchical structure
- exception and dependency chains
- critical risk-bearing units
- procedural and operational steps
- strong need for evidence-backed reasoning
If the retrieval and replay ideas cannot survive across these document types, they are unlikely to generalize well to other structured knowledge domains.
src/memory_engine: schema, storage, ingestion, retrieval, scoring, replayexamples/contract_pack: contract-like demo pack with dense dependencies and exceptionsexamples/runbook_pack: operational runbook pack for procedural retrievalbenchmarks/structured_memory: typed benchmark fixtures and evaluation assetsdocs: architecture, evaluation, hypotheses, and project visiontests: unit tests for schema, retrieval behavior, and benchmark support
docs/vision.md: why this project exists and where it is headingdocs/architecture.md: how the current system is structureddocs/api-tracks.md: legacy vs palace recall entry pointsdocs/evaluation.md: how retrieval modes are compareddocs/benchmark-strategy.md: how public, repo-owned, and private benchmarks should be useddocs/private-contract-dataset-guide.md: how to build and annotate a private contract golden setdocs/hypotheses.md: milestone hypotheses and success criteria
The first milestone tests three claims:
H1: graph-aware retrieval beats vanillatop-kretrieval on multi-hop questionsH2: anomaly and importance weighting improve recall of critical evidenceH3: replayable memory paths improve explainability without unacceptable latency
The retrieval stack separates:
- candidate generation
- semantic similarity backend
- scoring strategy
- path replay
That separation makes it possible to compare lexical baseline, embedding baseline, structure-only traversal, and weighted graph retrieval without rewriting the main search loop.
The evaluation layer can emit detailed per-question reports, which is useful for miss analysis and ablation debugging instead of relying only on a single aggregate score.
The repository also includes a dedicated structured benchmark bounded context with:
- strong pydantic dataset models
- a JSON repository for benchmark fixtures
- application services that load datasets, build stores, and run retrievers end to end
The benchmark story is intentionally split into three layers:
- External positioning: LongMemEval retrieval-only recall (
R@5,R@10,NDCG@10) at session or turn granularity - Public retrieval sanity: HotpotQA evidence retrieval on distractor-style multi-document questions
- Mechanism validation: repository-owned structured fixtures for path, semantic, contradiction, and dynamic-memory behavior
Current run matrix:
benchmarks/structured_memory/*.json: CIbenchmarks/structured_memory/spatial_recall_benchmark.json,route_replay_benchmark.json,consolidation_gain_benchmark.json,state_transition_benchmark.json,contradiction_tension_benchmark.json: Layer B checks for palace-oriented expectations (space, route shape, diffusion gain, lifecycle) and explicit contradiction / rule-tension pairsbenchmarks/external/hotpotqa/hotpot_tiny_fixture.json: CI sanitybenchmarks/external/hotpotqa/data/*.json: local / nightly (mediumdefault 64 samples,fulloptional)benchmarks/external/longmemeval/longmemeval_tiny_fixture.json: local sanitybenchmarks/external/longmemeval/data/*.json: local / nightly (mediumdefault 50 samples,fulloptional)benchmarks/layer_c_minimal/*: runnable Layer C transfer stand-ins + private annotation templates
- typed
MemoryNode/MemoryEdge/MemoryPathgraph - SQLite-backed local palace (
.mpe/) andmpeCLI - stdio MCP server + Cursor hook templates (
mpe hooks install) hybridretriever mode (lexical + embedding blend → graph expand)- LongMemEval session + turn KPI baselines and full-corpus reproduce recipe
- domain packs for contract / runbook / research documents
- Multi-backend vector zoo / hosted embedding services
- LLM-backed answer synthesis (path reasoning stays deterministic)
- multi-modal memory encoding
- full UI
See docs/ROADMAP.md. Product M1–M11 are done (memory-path-engine is on PyPI); next levers:
- Organization-side Layer C private gold labels
- Optional HotpotQA full distractor KPI / turn+fastembed tables
- Optional path explainability / agent UX polish
For suggested GitHub topic tags (About section), see docs/github-topics.md.
MIT. See LICENSE.