Skip to content

About

Search private data inside your app. A fast local SDK for vector, BM25, hybrid, and graph-aware retrieval—no server, account, or API key.

Topics

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

343 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RetrievalKit — fast, private retrieval inside your app

A local retrieval SDK for exact vector search, BM25, hybrid ranking, graph traversal, and graph-scoped context—all backed by one Rust core.

Website · Live demo · Documentation · Quickstart · Install · See validated benchmarks

Release CI Apache-2.0 license

1K–<50K chunks · in-process · SwiftPM · PyPI · npm · Maven Central

RetrievalKit searches app-owned data without requiring a retrieval server, account, or API key. Database construction, indexing, filtering, ranking, graph traversal, and persistence run locally. Embeddings remain explicit: bring your own vectors or use an optional local provider.

A recorded Swift hybrid query returns a ranked Apollo 11 passage with its vector and keyword trace.

This recorded hybrid run shows the retrieval contract: one query can combine meaning and keywords, return a deterministic ranking, and expose why the top passage won. The live browser demo uses the same local browser embedding, RetrievalKit WASM search, and browser answer pipeline for both suggested and free-form questions.

Quickstart

The Python base package is the shortest path to a complete local retrieval query.

python -m pip install retrievalkit==0.1.0
from retrievalkit import Document, RetrievalDatabaseBuilder

builder = RetrievalDatabaseBuilder(
    corpus_id="project-notes",
    metric="dot_product",
    encoding="f32",
)
builder.upsert(
    Document(
        id="decision-swift",
        text="We chose Swift for Project Apollo's Apple platform client.",
        metadata={"project": "apollo", "status": "approved"},
    ),
    embedding=[1.0, 0.0],
)
builder.upsert(
    Document(
        id="launch-checklist",
        text="Project Apollo launch checklist and release owners.",
        metadata={"project": "apollo", "status": "draft"},
    ),
    embedding=[0.0, 1.0],
)

database = builder.build()
hits = database.retrieval.hybrid_search(
    "Why did we choose Swift?",
    [1.0, 0.0],
    where={"project": "apollo"},
    limit=1,
)

print(hits[0]["document_id"])
decision-swift

The two-dimensional vectors keep the example deterministic. Production documents and queries should use embeddings from the same model. A checked-in version lives at database_quickstart.py.

Prefer another language? Start with the Swift, Python, TypeScript, or Kotlin guide.

Choose an API

Choose the smallest database that matches how your application finds context.

If you need to… Use Query behavior
Search a flat collection RetrievalDatabase Exact vector, BM25 text, or hybrid ranking
Follow declared relationships without embeddings GraphDatabase Match, traverse, and project related records
Select related records, then rank within that scope GraphRetrievalDatabase Graph-scoped exact vector, BM25, or hybrid ranking

For retrieval-capable databases, query inputs select the mode:

embedding only           → exact vector search
text only                → BM25 search
text + embedding + alpha → hybrid search

Metadata filters are hard constraints in every retrieval mode. Graph scope is different: it answers “which records are related?” before ranking. Your application supplies those relationships; RetrievalKit does not infer a graph.

How it works

One canonical corpus feeds retrieval, graph, and graph-scoped retrieval paths in the shared Rust core.

The canonical corpus owns records, chunks, metadata, stable identities, and generations. Retrieval and graph indexes are derived capabilities over that state. The Rust core owns indexing, traversal, filtering, ranking, traces, and persistence; wrappers provide idiomatic APIs without reimplementing retrieval.

Native databases use transactional, checksummed snapshots. Browser databases are in-memory and owned by a dedicated Worker in v0.1.0.

Privacy boundary

RetrievalKit database operations do not make network calls. If text must stay on-device, use a local embedding provider. If your application sends text to a remote embedding API, that embedding step is remote even though RetrievalKit search remains local.

Install

RetrievalKit v0.1.0 is a published preview. Graph aggregates include base retrieval; optional embedding packages remain independent.

SDK Graph-enabled install Qualified preview target
Swift .package(url: "https://github.com/gungorbasa/RetrievalKit.git", from: "0.1.0") macOS 14+ arm64; iOS 15+ arm64 device and simulator
Python python -m pip install retrievalkit-graph==0.1.0 macOS arm64; CPython 3.10–3.14
Node.js npm install @gungorbasa/retrievalkit-graph@0.1.0 macOS arm64; Node.js 22.13+ or 24 LTS
Browser npm install @gungorbasa/retrievalkit-browser@0.1.0 Dedicated Worker; portable and SIMD128 WASM tiers
Kotlin/JVM implementation("io.github.gungorbasa:retrievalkit-graph:0.1.0") macOS arm64 native library; JDK 17 build, Java 11+ runtime
Android implementation("io.github.gungorbasa:retrievalkit-graph-android:0.1.0") API 24+ arm64-v8a packaging; live-device behavior unqualified

Base packages are retrievalkit, @gungorbasa/retrievalkit, and io.github.gungorbasa:retrievalkit. Optional local embedding integrations are published separately.

Important

Python, Node, and Kotlin base and graph native aggregates are mutually exclusive within one process. Install exactly one retrieval aggregate; the independent embedding package can be used alongside either one.

Published package inventory
SDK Capability Status
Swift RetrievalKit Base corpus and retrieval Published preview
Swift RetrievalKitGraph Graph aggregate with retrieval Published preview
Swift EmbeddingKit Local Core ML embedding integration Published preview
Swift RetrievalKitPipeline Chunk → embed → index → search orchestration Published preview
Python retrievalkit Base corpus and retrieval Published preview
Python retrievalkit-graph Graph aggregate with retrieval Published preview
Python retrievalkit-embedding Local FP32 MiniLM embedding integration Published preview
TypeScript @gungorbasa/retrievalkit Base corpus and retrieval Published preview
TypeScript @gungorbasa/retrievalkit-graph Graph aggregate with retrieval Published preview
TypeScript @gungorbasa/retrievalkit-embedding Local FP32 MiniLM embedding integration Published preview
Browser @gungorbasa/retrievalkit-browser Worker-owned base, graph, and graph-scoped WASM retrieval Published preview
Browser @gungorbasa/retrievalkit-browser-embedding Worker-owned local FP32 MiniLM embedding Published preview
Kotlin/JVM io.github.gungorbasa:retrievalkit Base corpus and retrieval Published preview
Kotlin/JVM io.github.gungorbasa:retrievalkit-graph Graph aggregate with retrieval Published preview
Kotlin/JVM io.github.gungorbasa:retrievalkit-embedding Local FP32 MiniLM embedding integration Published preview
Android io.github.gungorbasa:retrievalkit-android Base AAR for arm64-v8a Published preview; live-device unqualified
Android io.github.gungorbasa:retrievalkit-graph-android Graph aggregate AAR for arm64-v8a Published preview; live-device unqualified
Android io.github.gungorbasa:retrievalkit-embedding-android Local FP32 MiniLM embedding AAR for arm64-v8a Published preview; live-device inference unqualified

Public SwiftPM, PyPI, npm, and Maven publication completed from the signed release revision and authorized artifacts. Android qualification covers cross-compilation, packaging, inventory, ABI/JNI contracts, and fresh consumer resolution and compilation—not physical-device inference, compatibility, memory, thermal behavior, or performance.

Swift ships one graph-capable native aggregate so RetrievalKit and RetrievalKitGraph can coexist in one app. Selecting only RetrievalKit keeps graph APIs out of the Swift target, although SwiftPM still downloads the shared binary.

Run from source

Clone the repository and run one checked quickstart from the repository root. Each script verifies its required toolchain before building.

Python

Requires Rust and CPython 3.10–3.14.

PYTHON_BIN=python3 scripts/check-python-graph-wrapper.sh
target/python-graph-wrapper-check-venv-py*/bin/python \
  wrappers/python-graph/examples/graph_retrieval_quickstart.py

Expected output: graph-hybrid=decision-swift.

TypeScript / Node.js

Requires Rust and Node.js 22.13+ or 24 LTS.

cd wrappers/typescript
npm ci
npm run preflight
npm run build
node graph/examples/graph-retrieval.mjs

The result contains documentId: 'local'.

Kotlin / JVM

Requires Rust and JDK 17 on macOS arm64. Published bytecode runs on Java 11+.

export JAVA_HOME=$(/usr/libexec/java_home -v 17)
export PATH="$JAVA_HOME/bin:$PATH"
cd wrappers/kotlin
./scripts/preflight.sh jvm
./scripts/build-native.sh jvm
./gradlew :example-retrieval:run

Expected output includes kotlin: Kotlin calls the local Rust retrieval core. (1.0).

Swift

Requires Rust, Swift 6.2, and macOS 14+ on Apple silicon.

scripts/build-xcframework.sh --macos-only --graph
scripts/run-swift-quickstart.sh graph-retrieval

Expected output: graph-hybrid=decision-swift.

Validated evidence

The public claims below are historical observations from frozen workloads, not measurements of the current checkout. They apply to RetrievalKit revision 9c784d2f11b91bb907150aa1b6046880ff89fde6, were reported on 2026-07-21, and expire on 2027-07-21. Retrieval timings exclude embedding generation.

Exact retrieval on Apple M1 Max

On the frozen exact F32, 384-dimensional, top-10 benchmark, RetrievalKit revision 9c784d2 delivered the following P50 unfiltered retrieval ratios versus sqlite-vec 0.1.9 on an Apple M1 Max running macOS 26.5.2. Each lane used 100 measured queries after 20 warmups; embedding was excluded.

Corpus sqlite-vec / RetrievalKit P50 Observation
10K 7.17× RetrievalKit lower latency
25K 7.60× RetrievalKit lower latency
50K 7.29× RetrievalKit lower latency

With the same frozen filter enabled, the P50 retrieval ratios were 10.38× at 10K, 9.08× at 25K, and 8.43× at 50K versus sqlite-vec 0.1.9. This was the same Apple M1 Max exact F32 workload at revision 9c784d2; embedding was excluded.

RetrievalKit exact F32 and sqlite-vec 0.1.9 both passed the frozen Phase 5 identity, filtering, deletion, determinism, and reload gates at 10K, 25K, and 50K. This result is scoped to the frozen workload and is not proof for every possible input.

These ratios describe one exact-search workload, not universal competitor superiority. See the methodology and Mac evidence report.

Graph-scoped quality on HotpotQA

Across the frozen 296-query HotpotQA linked-abstracts test comparison, graph-scoped weighted-I8 retrieval increased NDCG@10 from 0.858036 to 0.927909 versus whole-corpus weighted-I8 retrieval. There were 121 wins, 157 ties, and 18 losses. This is a scoped quality result, not a universal graph winner or a latency claim.

On the same frozen 296-query weighted-I8 comparison, Recall@10 increased from 0.871622 to 0.957770 and complete-evidence recall@10 increased from 0.743243 to 0.922297. Sixteen queries lost on each recall measure; those losses are part of the result.

The frozen candidate stage reduced the mean per-query candidate set by 972.65× while retaining 96.79% candidate recall and 94.26% candidate complete evidence across 296 valid graph queries, with zero empty scopes. Candidate reduction is not a retrieval-latency speedup and retention was not perfect.

The workload contains 12,670 chunks. See the full retrieval-quality evidence.

Physical-device qualification

On a physical iPhone 17 Pro Max (iPhone18,2, V54AP), the supported 10K, 25K, and 50K F32/I8 product workflows passed. All six graph-free candidate-to-baseline median-session P95 ratios were at or below the frozen 1.03 gate. Query/prepare evidence used iOS 26.5.1 (23F81); remaining lifecycle evidence used iOS 26.5.2 (23F84). This is supported-workload qualification for that device, with embedding excluded, not a claim about other hardware.

V1 targets fewer than 50K chunks. The 100K Phase 4b stress workload remains not_run_device_safety, produced zero accepted stress artifacts, and is not eligible for support, performance, latency, quality, product, or marketing claims.

See the physical-device evidence report and Phase 6 validation result.

Scope and limitations

  • V1 is optimized for exact retrieval over local indexes with fewer than 50K chunks. HNSW and other ANN indexes are intentionally deferred.
  • Native persistence uses transactional, checksummed snapshots. Browser v0.1.0 databases are Worker-owned and in-memory.
  • Initial native binaries focus on arm64 Apple platforms. Node.js, Python, and Kotlin/JVM packages initially target macOS arm64.
  • Android API 24+ arm64-v8a is a packaging-qualified preview. Live-device inference, compatibility, lifecycle, memory, thermal behavior, and performance remain unqualified.
  • Embedding latency is separate from retrieval latency. RetrievalKit never hides a hosted inference call inside a database operation.
  • Benchmark evidence supports only the named frozen workloads.

Documentation

Start building Architecture and policy
Swift guide Product specification
Python guide Capability-separated architecture
TypeScript and browser guide Compatibility policy
Kotlin/JVM and Android guide Release process

API references: Swift base · Swift graph · Python base · Python graph · TypeScript / Node.js · Browser / WebAssembly · Kotlin/JVM and Android

The complete public documentation is available at retrievalkit.com/documentation.

Contributing, support, and license

Focused bug reports, documentation corrections, reproduction cases, and changes within the V1 product scope are welcome. Read CONTRIBUTING.md before opening a pull request and use the issue templates for bugs or feature requests.

Report vulnerabilities privately through the security policy. Release history is recorded in the changelog.

RetrievalKit is licensed under the Apache License 2.0. Copyright and distribution notices are in NOTICE.

About

Search private data inside your app. A fast local SDK for vector, BM25, hybrid, and graph-aware retrieval—no server, account, or API key.

Topics

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages