A local retrieval SDK for exact vector search, BM25, hybrid ranking, graph traversal, and graph-scoped context—all backed by one Rust core.
Website · Live demo · Documentation · Quickstart · Install · See validated benchmarks
1K–<50K chunks · in-process ·
SwiftPM · PyPI · npm · Maven Central
RetrievalKit searches app-owned data without requiring a retrieval server, account, or API key. Database construction, indexing, filtering, ranking, graph traversal, and persistence run locally. Embeddings remain explicit: bring your own vectors or use an optional local provider.
This recorded hybrid run shows the retrieval contract: one query can combine meaning and keywords, return a deterministic ranking, and expose why the top passage won. The live browser demo uses the same local browser embedding, RetrievalKit WASM search, and browser answer pipeline for both suggested and free-form questions.
The Python base package is the shortest path to a complete local retrieval query.
python -m pip install retrievalkit==0.1.0from retrievalkit import Document, RetrievalDatabaseBuilder
builder = RetrievalDatabaseBuilder(
corpus_id="project-notes",
metric="dot_product",
encoding="f32",
)
builder.upsert(
Document(
id="decision-swift",
text="We chose Swift for Project Apollo's Apple platform client.",
metadata={"project": "apollo", "status": "approved"},
),
embedding=[1.0, 0.0],
)
builder.upsert(
Document(
id="launch-checklist",
text="Project Apollo launch checklist and release owners.",
metadata={"project": "apollo", "status": "draft"},
),
embedding=[0.0, 1.0],
)
database = builder.build()
hits = database.retrieval.hybrid_search(
"Why did we choose Swift?",
[1.0, 0.0],
where={"project": "apollo"},
limit=1,
)
print(hits[0]["document_id"])decision-swift
The two-dimensional vectors keep the example deterministic. Production
documents and queries should use embeddings from the same model. A checked-in
version lives at
database_quickstart.py.
Prefer another language? Start with the Swift, Python, TypeScript, or Kotlin guide.
Choose the smallest database that matches how your application finds context.
| If you need to… | Use | Query behavior |
|---|---|---|
| Search a flat collection | RetrievalDatabase |
Exact vector, BM25 text, or hybrid ranking |
| Follow declared relationships without embeddings | GraphDatabase |
Match, traverse, and project related records |
| Select related records, then rank within that scope | GraphRetrievalDatabase |
Graph-scoped exact vector, BM25, or hybrid ranking |
For retrieval-capable databases, query inputs select the mode:
embedding only → exact vector search
text only → BM25 search
text + embedding + alpha → hybrid search
Metadata filters are hard constraints in every retrieval mode. Graph scope is different: it answers “which records are related?” before ranking. Your application supplies those relationships; RetrievalKit does not infer a graph.
The canonical corpus owns records, chunks, metadata, stable identities, and generations. Retrieval and graph indexes are derived capabilities over that state. The Rust core owns indexing, traversal, filtering, ranking, traces, and persistence; wrappers provide idiomatic APIs without reimplementing retrieval.
Native databases use transactional, checksummed snapshots. Browser databases are in-memory and owned by a dedicated Worker in v0.1.0.
RetrievalKit database operations do not make network calls. If text must stay on-device, use a local embedding provider. If your application sends text to a remote embedding API, that embedding step is remote even though RetrievalKit search remains local.
RetrievalKit v0.1.0 is a published preview. Graph aggregates include base retrieval; optional embedding packages remain independent.
| SDK | Graph-enabled install | Qualified preview target |
|---|---|---|
| Swift | .package(url: "https://github.com/gungorbasa/RetrievalKit.git", from: "0.1.0") |
macOS 14+ arm64; iOS 15+ arm64 device and simulator |
| Python | python -m pip install retrievalkit-graph==0.1.0 |
macOS arm64; CPython 3.10–3.14 |
| Node.js | npm install @gungorbasa/retrievalkit-graph@0.1.0 |
macOS arm64; Node.js 22.13+ or 24 LTS |
| Browser | npm install @gungorbasa/retrievalkit-browser@0.1.0 |
Dedicated Worker; portable and SIMD128 WASM tiers |
| Kotlin/JVM | implementation("io.github.gungorbasa:retrievalkit-graph:0.1.0") |
macOS arm64 native library; JDK 17 build, Java 11+ runtime |
| Android | implementation("io.github.gungorbasa:retrievalkit-graph-android:0.1.0") |
API 24+ arm64-v8a packaging; live-device behavior unqualified |
Base packages are retrievalkit, @gungorbasa/retrievalkit, and
io.github.gungorbasa:retrievalkit. Optional local embedding integrations are
published separately.
Important
Python, Node, and Kotlin base and graph native aggregates are mutually exclusive within one process. Install exactly one retrieval aggregate; the independent embedding package can be used alongside either one.
Published package inventory
| SDK | Capability | Status |
|---|---|---|
Swift RetrievalKit |
Base corpus and retrieval | Published preview |
Swift RetrievalKitGraph |
Graph aggregate with retrieval | Published preview |
Swift EmbeddingKit |
Local Core ML embedding integration | Published preview |
Swift RetrievalKitPipeline |
Chunk → embed → index → search orchestration | Published preview |
Python retrievalkit |
Base corpus and retrieval | Published preview |
Python retrievalkit-graph |
Graph aggregate with retrieval | Published preview |
Python retrievalkit-embedding |
Local FP32 MiniLM embedding integration | Published preview |
TypeScript @gungorbasa/retrievalkit |
Base corpus and retrieval | Published preview |
TypeScript @gungorbasa/retrievalkit-graph |
Graph aggregate with retrieval | Published preview |
TypeScript @gungorbasa/retrievalkit-embedding |
Local FP32 MiniLM embedding integration | Published preview |
Browser @gungorbasa/retrievalkit-browser |
Worker-owned base, graph, and graph-scoped WASM retrieval | Published preview |
Browser @gungorbasa/retrievalkit-browser-embedding |
Worker-owned local FP32 MiniLM embedding | Published preview |
Kotlin/JVM io.github.gungorbasa:retrievalkit |
Base corpus and retrieval | Published preview |
Kotlin/JVM io.github.gungorbasa:retrievalkit-graph |
Graph aggregate with retrieval | Published preview |
Kotlin/JVM io.github.gungorbasa:retrievalkit-embedding |
Local FP32 MiniLM embedding integration | Published preview |
Android io.github.gungorbasa:retrievalkit-android |
Base AAR for arm64-v8a | Published preview; live-device unqualified |
Android io.github.gungorbasa:retrievalkit-graph-android |
Graph aggregate AAR for arm64-v8a | Published preview; live-device unqualified |
Android io.github.gungorbasa:retrievalkit-embedding-android |
Local FP32 MiniLM embedding AAR for arm64-v8a | Published preview; live-device inference unqualified |
Public SwiftPM, PyPI, npm, and Maven publication completed from the signed release revision and authorized artifacts. Android qualification covers cross-compilation, packaging, inventory, ABI/JNI contracts, and fresh consumer resolution and compilation—not physical-device inference, compatibility, memory, thermal behavior, or performance.
Swift ships one graph-capable native aggregate so RetrievalKit and
RetrievalKitGraph can coexist in one app. Selecting only RetrievalKit keeps
graph APIs out of the Swift target, although SwiftPM still downloads the shared
binary.
Clone the repository and run one checked quickstart from the repository root. Each script verifies its required toolchain before building.
Python
Requires Rust and CPython 3.10–3.14.
PYTHON_BIN=python3 scripts/check-python-graph-wrapper.sh
target/python-graph-wrapper-check-venv-py*/bin/python \
wrappers/python-graph/examples/graph_retrieval_quickstart.pyExpected output: graph-hybrid=decision-swift.
TypeScript / Node.js
Requires Rust and Node.js 22.13+ or 24 LTS.
cd wrappers/typescript
npm ci
npm run preflight
npm run build
node graph/examples/graph-retrieval.mjsThe result contains documentId: 'local'.
Kotlin / JVM
Requires Rust and JDK 17 on macOS arm64. Published bytecode runs on Java 11+.
export JAVA_HOME=$(/usr/libexec/java_home -v 17)
export PATH="$JAVA_HOME/bin:$PATH"
cd wrappers/kotlin
./scripts/preflight.sh jvm
./scripts/build-native.sh jvm
./gradlew :example-retrieval:runExpected output includes
kotlin: Kotlin calls the local Rust retrieval core. (1.0).
Swift
Requires Rust, Swift 6.2, and macOS 14+ on Apple silicon.
scripts/build-xcframework.sh --macos-only --graph
scripts/run-swift-quickstart.sh graph-retrievalExpected output: graph-hybrid=decision-swift.
The public claims below are historical observations from frozen workloads, not measurements of the current checkout. They apply to RetrievalKit revision
9c784d2f11b91bb907150aa1b6046880ff89fde6, were reported on 2026-07-21,
and expire on 2027-07-21. Retrieval timings exclude embedding generation.
Exact retrieval on Apple M1 Max
On the frozen exact F32, 384-dimensional, top-10 benchmark, RetrievalKit revision
9c784d2 delivered the following P50 unfiltered retrieval ratios versus
sqlite-vec 0.1.9 on an Apple M1 Max running macOS 26.5.2. Each lane used 100
measured queries after 20 warmups; embedding was excluded.
| Corpus | sqlite-vec / RetrievalKit P50 | Observation |
|---|---|---|
| 10K | 7.17× | RetrievalKit lower latency |
| 25K | 7.60× | RetrievalKit lower latency |
| 50K | 7.29× | RetrievalKit lower latency |
With the same frozen filter enabled, the P50 retrieval ratios were 10.38× at
10K, 9.08× at 25K, and 8.43× at 50K versus sqlite-vec 0.1.9. This was the
same Apple M1 Max exact F32 workload at revision 9c784d2; embedding was
excluded.
RetrievalKit exact F32 and sqlite-vec 0.1.9 both passed the frozen Phase 5
identity, filtering, deletion, determinism, and reload gates at 10K, 25K, and
50K. This result is scoped to the frozen workload and is not proof for every
possible input.
These ratios describe one exact-search workload, not universal competitor superiority. See the methodology and Mac evidence report.
Graph-scoped quality on HotpotQA
Across the frozen 296-query HotpotQA linked-abstracts test comparison, graph-scoped weighted-I8 retrieval increased NDCG@10 from 0.858036 to 0.927909 versus whole-corpus weighted-I8 retrieval. There were 121 wins, 157 ties, and 18 losses. This is a scoped quality result, not a universal graph winner or a latency claim.
On the same frozen 296-query weighted-I8 comparison, Recall@10 increased from 0.871622 to 0.957770 and complete-evidence recall@10 increased from 0.743243 to 0.922297. Sixteen queries lost on each recall measure; those losses are part of the result.
The frozen candidate stage reduced the mean per-query candidate set by 972.65× while retaining 96.79% candidate recall and 94.26% candidate complete evidence across 296 valid graph queries, with zero empty scopes. Candidate reduction is not a retrieval-latency speedup and retention was not perfect.
The workload contains 12,670 chunks. See the full retrieval-quality evidence.
Physical-device qualification
On a physical iPhone 17 Pro Max (iPhone18,2, V54AP), the supported 10K,
25K, and 50K F32/I8 product workflows passed. All six graph-free
candidate-to-baseline median-session P95 ratios were at or below the frozen
1.03 gate. Query/prepare evidence used iOS 26.5.1 (23F81); remaining lifecycle
evidence used iOS 26.5.2 (23F84). This is supported-workload qualification for
that device, with embedding excluded, not a claim about other hardware.
V1 targets fewer than 50K chunks. The 100K Phase 4b stress workload remains
not_run_device_safety, produced zero accepted stress artifacts, and is not
eligible for support, performance, latency, quality, product, or marketing
claims.
See the physical-device evidence report and Phase 6 validation result.
- V1 is optimized for exact retrieval over local indexes with fewer than 50K chunks. HNSW and other ANN indexes are intentionally deferred.
- Native persistence uses transactional, checksummed snapshots. Browser v0.1.0 databases are Worker-owned and in-memory.
- Initial native binaries focus on arm64 Apple platforms. Node.js, Python, and Kotlin/JVM packages initially target macOS arm64.
- Android API 24+ arm64-v8a is a packaging-qualified preview. Live-device inference, compatibility, lifecycle, memory, thermal behavior, and performance remain unqualified.
- Embedding latency is separate from retrieval latency. RetrievalKit never hides a hosted inference call inside a database operation.
- Benchmark evidence supports only the named frozen workloads.
| Start building | Architecture and policy |
|---|---|
| Swift guide | Product specification |
| Python guide | Capability-separated architecture |
| TypeScript and browser guide | Compatibility policy |
| Kotlin/JVM and Android guide | Release process |
API references: Swift base · Swift graph · Python base · Python graph · TypeScript / Node.js · Browser / WebAssembly · Kotlin/JVM and Android
The complete public documentation is available at retrievalkit.com/documentation.
Focused bug reports, documentation corrections, reproduction cases, and changes within the V1 product scope are welcome. Read CONTRIBUTING.md before opening a pull request and use the issue templates for bugs or feature requests.
Report vulnerabilities privately through the security policy. Release history is recorded in the changelog.
RetrievalKit is licensed under the Apache License 2.0. Copyright and distribution notices are in NOTICE.