Skip to content

Repository files navigation

Memory Path Engine

CI PyPI Python 3.11+ License: MIT Status: Productizing

Local memory with replayable evidence paths — not only top-k chunks. Structured graph retrieval for agents, with a local palace CLI, MCP closed loop, and public-benchmark KPIs.

Memory Path Engine models memory as typed nodes, edges, weights, and MemoryPath objects so a system can retrieve, traverse, and explain how it reached an answer. Product M1 adds an on-disk palace (SQLite) and the mpe CLI; Layer B fixtures remain the architecture proof surface.

System shape (v0 + Memory Palace v1 + Product M1)

Bundled markdown packs are ingested into a graph (MemoryNode / MemoryEdge). Retrievers return a MemoryPath. The CLI persists that graph under .mpe/ and prints answer + hops.

Memory Palace v1 adds a parallel domain (memory_engine.memory) mapped through palace_to_store. See docs/architecture.md and docs/ROADMAP.md.

 examples/*_pack  ──▶  mpe ingest  ──▶  .mpe/store.sqlite (typed graph)
                                              │
                               ┌──────────────┼──────────────┐
                               ▼              ▼              ▼
                         BaselineTopK    other modes    WeightedGraph
                         (flat answers)  in `retrieve`  (path + scores)
                                              │
                                              ▼
                               mpe search/path → answer + hop list

Why this project is different

Most RAG systems still look like this:

  1. Split documents into chunks.
  2. Embed chunks.
  3. Return top-k matches.
  4. Ask the LLM to improvise the reasoning.

This project asks:

Can we retrieve a memory path instead of only retrieving similar chunks?

Three product bets:

  • structure: typed nodes and edges, not only a flat vector index
  • weight: importance, risk, novelty, reinforcement / forgetting
  • path: replayable evidence chains as the default search output

What you can do here

  • compare multiple retrieval modes in one codebase
  • inspect replayable evidence paths instead of only final answers
  • test graph-aware retrieval on contract-like and operational documents
  • run repository-owned structured benchmarks instead of toy snippets

Quick start

Maintainers: configure the GitHub link-card image using docs/social-preview.md (docs/assets/open-graph-cover.png).

Install from PyPI (Product M8) or editable clone (see docs/install.md for pipx / uv / Docker):

pip install memory-path-engine
# pip install 'memory-path-engine[embed]'   # optional dense backends
# or: python -m pip install --no-build-isolation -e .
mpe doctor

Publish / first tag: docs/publish.md · Changelog: CHANGELOG.md.

Product CLI (mpe) — local palace

mpe init --mode hybrid
mpe ingest examples/runbook_pack/runbooks --pack example_runbook_pack
mpe search "What if rollback does not recover the API?" --mode hybrid
mpe path "What if rollback does not recover the API?"
mpe status
mpe backup

Palace files live in ./.mpe/ (or $MPE_PALACE). Search always prints an answer plus hop citations.

5-minute agent closed loop (MCP + hooks): see docs/getting-started.md.

mpe hooks install
mpe mcp   # stdio MCP server for Cursor / Claude
# docker build -t mpe-mcp . && docker run -i --rm -v "$PWD/.mpe:/data/palace" -e MPE_PALACE=/data/palace mpe-mcp

Acceptance checklist: docs/ACCEPTANCE.md.
Current progress vs vision stages: docs/progress.md.

LongMemEval product KPI baseline

mpe bench longmemeval --label tiny --granularity session
mpe bench longmemeval --label tiny --granularity turn

Checked-in tiny reports live under benchmarks/external/longmemeval/baselines/.

Latest public recall (session, LongMemEval-S cleaned):

Slice Embedding Mode R@5 R@10 NDCG@10
50q ngram hybrid 0.980 1.000 0.948
50q ngram lexical_baseline 0.980 1.000 0.948
500q (full) ngram hybrid 0.950 0.974 0.863
500q (full) ngram lexical_baseline 0.950 0.974 0.863
500q (full) fastembed hybrid 0.962 0.980 0.875
50q turn ngram hybrid 0.840 0.920 0.714
50q turn ngram lexical_baseline 0.840 0.920 0.714

Artifacts: benchmarks/external/longmemeval/baselines/longmemeval_kpi_{medium30,medium50,medium50_fastembed,medium50_turn,full_ngram,full_fastembed}.*.

M7 note: hybrid public ranking now reuses BM25/blend seed scores so NDCG no longer collapses after graph expansion. With fastembed, hybrid beats lexical on full R@5/R@10/NDCG. Turn mid-slice (Product M11) is harder than session (R@5 0.84 vs 0.98) as expected with finer units.

HotpotQA mid-slice (64q, evidence-hit, Product M10):

Embedding Mode evidence_hit comparison hit
ngram lexical_baseline 0.406 0.562
ngram hybrid 0.844 0.938
fastembed embedding_baseline 0.531 0.812
fastembed hybrid 0.859 0.938

Artifacts: benchmarks/external/hotpotqa/baselines/hotpotqa_kpi_medium64*.md. Dense (fastembed) especially helps comparison/paraphrase-like questions; set MPE_EMBEDDING_CACHE_DIR to cache vectors on disk.

Optional dense embeddings (Product M5+): default remains dependency-free ngram. Install pip install 'memory-path-engine[embed]' then:

mpe bench longmemeval --label tiny --embedding fastembed
# or: export MPE_EMBEDDING=fastembed
# optional: export MPE_EMBEDDING_MAX_CHARS=4000

Dense backends batch-encode when texts are short, and head+tail-truncate long session memories to fit model context.

Reproduce full LongMemEval-S (public KPI):

python scripts/download_longmemeval.py
mpe bench longmemeval \
  --dataset benchmarks/external/longmemeval/data/longmemeval_s_cleaned.json \
  --label full --granularity session --limit 0
Slice Granularity How to run Notes
tiny (2q, checked in) session / turn mpe bench longmemeval --label tiny [--granularity turn] CI smoke + committed baselines
medium (50q) session nightly / local with --limit 50 positioning
full (LongMemEval-S) session / turn download + --label full --limit 0 headline public KPI

See benchmarks/external/longmemeval/README.md and docs/ROADMAP.md.

Run the test suite:

python -m unittest discover -s tests -v

Run the runbook demo:

python -m memory_engine.demo --scenario runbook

Terminal-style capture of real stdout (refresh with python scripts/generate_runbook_demo_terminal_svg.py; latency_ms may differ run to run):

Runbook demo terminal output

Run the research-notes demo:

python -m memory_engine.demo --scenario research

Run the HotpotQA tiny benchmark sanity check:

python scripts/run_hotpotqa_benchmark.py

Download HotpotQA (CMU URL, with HuggingFace Hub fallback) and run the medium64 KPI slice:

python scripts/download_hotpotqa.py
python scripts/run_hotpotqa_benchmark.py \
  --dataset benchmarks/external/hotpotqa/data/hotpot_dev_distractor_v1.json \
  --limit 64 --top-k 10 \
  --modes lexical_baseline,embedding_baseline,weighted_graph,hybrid,activation_spreading_v1 \
  --summary-output benchmarks/external/hotpotqa/data/hotpotqa-medium64-summary.json

Optional dense embedding disk cache (Product M10):

export MPE_EMBEDDING_CACHE_DIR="$PWD/.mpe/embed-cache"
export MPE_EMBEDDING=fastembed

Committed mid-slice KPI: benchmarks/external/hotpotqa/baselines/hotpotqa_kpi_medium64*.md.

Run the LongMemEval tiny benchmark sanity check:

python scripts/run_longmemeval_benchmark.py

Print compact v1 palace metadata (spaces, routes, memory kinds) per case:

python scripts/run_longmemeval_benchmark.py --v1-recall-summary

Generate a fixed-format Layer B report (path/route/space/lifecycle/activation snapshot):

python scripts/generate_layer_b_report.py --output "benchmarks/structured_memory/layer_b_report.json" --markdown-output "benchmarks/structured_memory/layer_b_report.md"

Generate a fixed-format ablation matrix and latency summary (no-structure / no-weight / no-path-expansion):

python scripts/generate_ablation_report.py --output "benchmarks/structured_memory/ablation_report.json" --markdown-output "benchmarks/structured_memory/ablation_report.md"

Generate a Layer A positioning report (external metrics only; keeps architecture claims in Layer B):

python scripts/generate_layer_a_report.py --slice-profile tiny --markdown-output "benchmarks/external/layer_a_report.md"

Run Layer C transferable stand-in benchmarks:

python scripts/run_layer_c_benchmark.py --markdown-output "benchmarks/layer_c_minimal/layer_c_report.md"

Download the official HotpotQA dev distractor file for local benchmark runs:

python scripts/download_hotpotqa.py

Download the cleaned LongMemEval-S file for local benchmark runs:

python scripts/download_longmemeval.py

What you will see

python -m memory_engine.demo prints a small banner, the query, then path-aware output: a BEST ANSWER line built from the winning walk, and a REPLAY PATH with one line per hop (node id, score, via=<edge type>) plus short scoring reasons on the following lines. With --scenario contract, a BASELINE block (flat top-k answers) appears above the path-aware section for the same query.

Representative runbook excerpt (answer line shortened; latency and hop scores can vary slightly between runs):

========================================================================
  Memory Path Engine  |  demo
  scenario: runbook
========================================================================
-------------------------------- QUERY ---------------------------------
  What should we do if rollback does not recover the API after a
  deployment incident?
----------------- PATH-AWARE  weighted graph retrieval -----------------
  BEST ANSWER
    … stitched runbook units … [latency_ms=…]

  REPLAY PATH
    1. 01_api_incident_runbook:5  |  score=0.500  |  via=seed
       seed hit semantic=0.501
    2. 01_api_incident_runbook:4  |  score=0.299  |  via=next_unit
       expanded at hop 1 total=0.299 exception=0.450 contradiction=0.000
========================================================================

What the demos show

Runbook demo

The runbook demo loads incident and recovery procedures, then asks a multi-step operational question:

What should we do if rollback does not recover the API after a deployment incident?

The output includes:

  • a BEST ANSWER line composed from the graph walk
  • a REPLAY PATH with per-step scores, via edge types, and short reasons

For a representative stdout excerpt, see What you will see (under Quick start).

Contract demo

The contract demo runs the same query through a baseline retriever and the weighted graph retriever. Stdout shows flat top-k answers first, then the path-aware best answer and replay steps, so you can compare shapes of evidence without relying on a single aggregate metric.

Retrieval modes in this repo

Retriever What it emphasizes Useful for
lexical baseline keyword overlap simple lookups and sanity checks
embedding baseline semantic similarity paraphrases and fuzzy matches
structure-only traversal graph connectivity linked evidence exploration
weighted graph retrieval structure plus importance weighting multi-hop retrieval with replayable evidence
activation spreading v1 explicit propagation with decay graph diffusion experiments

Why the examples span multiple document types

The core is meant to stay domain-agnostic. The current examples use both contract-like documents and runbooks because together they stress:

  • hierarchical structure
  • exception and dependency chains
  • critical risk-bearing units
  • procedural and operational steps
  • strong need for evidence-backed reasoning

If the retrieval and replay ideas cannot survive across these document types, they are unlikely to generalize well to other structured knowledge domains.

Repository layout

Read this first

Research hypotheses

The first milestone tests three claims:

  • H1: graph-aware retrieval beats vanilla top-k retrieval on multi-hop questions
  • H2: anomaly and importance weighting improve recall of critical evidence
  • H3: replayable memory paths improve explainability without unacceptable latency

Experimental framework

The retrieval stack separates:

  • candidate generation
  • semantic similarity backend
  • scoring strategy
  • path replay

That separation makes it possible to compare lexical baseline, embedding baseline, structure-only traversal, and weighted graph retrieval without rewriting the main search loop.

The evaluation layer can emit detailed per-question reports, which is useful for miss analysis and ablation debugging instead of relying only on a single aggregate score.

The repository also includes a dedicated structured benchmark bounded context with:

  • strong pydantic dataset models
  • a JSON repository for benchmark fixtures
  • application services that load datasets, build stores, and run retrievers end to end

Benchmarks

The benchmark story is intentionally split into three layers:

  • External positioning: LongMemEval retrieval-only recall (R@5, R@10, NDCG@10) at session or turn granularity
  • Public retrieval sanity: HotpotQA evidence retrieval on distractor-style multi-document questions
  • Mechanism validation: repository-owned structured fixtures for path, semantic, contradiction, and dynamic-memory behavior

Current run matrix:

  • benchmarks/structured_memory/*.json: CI
  • benchmarks/structured_memory/spatial_recall_benchmark.json, route_replay_benchmark.json, consolidation_gain_benchmark.json, state_transition_benchmark.json, contradiction_tension_benchmark.json: Layer B checks for palace-oriented expectations (space, route shape, diffusion gain, lifecycle) and explicit contradiction / rule-tension pairs
  • benchmarks/external/hotpotqa/hotpot_tiny_fixture.json: CI sanity
  • benchmarks/external/hotpotqa/data/*.json: local / nightly (medium default 64 samples, full optional)
  • benchmarks/external/longmemeval/longmemeval_tiny_fixture.json: local sanity
  • benchmarks/external/longmemeval/data/*.json: local / nightly (medium default 50 samples, full optional)
  • benchmarks/layer_c_minimal/*: runnable Layer C transfer stand-ins + private annotation templates

What is in scope for v0.4 (Product M3)

  • typed MemoryNode / MemoryEdge / MemoryPath graph
  • SQLite-backed local palace (.mpe/) and mpe CLI
  • stdio MCP server + Cursor hook templates (mpe hooks install)
  • hybrid retriever mode (lexical + embedding blend → graph expand)
  • LongMemEval session + turn KPI baselines and full-corpus reproduce recipe
  • domain packs for contract / runbook / research documents

What is out of scope for now

  • Multi-backend vector zoo / hosted embedding services
  • LLM-backed answer synthesis (path reasoning stays deterministic)
  • multi-modal memory encoding
  • full UI

Planned next steps

See docs/ROADMAP.md. Product M1–M11 are done (memory-path-engine is on PyPI); next levers:

  • Organization-side Layer C private gold labels
  • Optional HotpotQA full distractor KPI / turn+fastembed tables
  • Optional path explainability / agent UX polish

For suggested GitHub topic tags (About section), see docs/github-topics.md.

License

MIT. See LICENSE.

Releases

Packages

Contributors

Languages